Skip to content
ittechwire

Technology news, clearly sourced

  1. Home
  2. AI

AI

Google releases EmbeddingGemma 2, an open embedding model for text, images, audio and video

The 740-million-parameter model places several kinds of media in one shared vector space, is designed to run on phones and is published under the Apache 2.0 license, Google says.

ittechwire Editorial3 min readSources: 2

Key points

  1. 1EmbeddingGemma 2 maps text, code, images, audio and video into one shared embedding space.
  2. 2It has 740 million parameters, is built on Gemma 4 and is released under the Apache 2.0 license.
  3. 3Text-only use needs about 270 million parameters; the vision and audio encoders are optional.
  4. 4Google reports about 191MB of active RAM for text-only use on a Pixel 11 Pro after quantization.
  5. 5Its context window is four times larger than before, and Google says the MTEB Code score rose from 68.76 to 78.68.

Full story

Google has released EmbeddingGemma 2, the successor to the text-only embedding model it introduced last year. Embedding models turn content into lists of numbers, called vectors, so that software can find related items by meaning rather than by matching keywords. The new version handles code, images, video and audio as well as text and maps all of them into a single shared space, so that a text query can, for example, surface a matching audio clip or a moment in a video. The announcement was written by two research engineers at Google DeepMind.

According to Google, the model has 740 million parameters, is built on the Gemma 4 architecture and draws on the same technology as the company's Gemini Embedding models. It is modular: text-only use needs about 270 million parameters, while a vision encoder (170 million) and an audio encoder (300 million) can be added for full multimodal support. The context window, the amount of input the model can take at once, is four times larger than in the first version; Google says it is enough for up to 5.5 minutes of audio, 29 images or 58 video frames in one input.

Google says that with quantization the text-only weights need roughly 191MB of active memory on a Pixel 11 Pro, and the full multimodal model about 567MB. The output uses Matryoshka Representation Learning, so developers can shorten vectors from 768 dimensions to 512, 256 or 128; the company says this cuts storage by up to six times. On benchmarks, Google reports leading scores among multimodal embedding models below one billion parameters on MTEB Code and on MAEB, an audio benchmark. Its MTEB Code score rose from 68.76 to 78.68, while text performance is said to stay at the level of the first version. These are Google's own figures; the full results are in the model card.

The weights are available on Hugging Face and Kaggle, with versions tuned for on-device use in the LiteRT Community on Hugging Face; availability in the Gemini Enterprise Agent Platform Model Garden is listed as coming soon. Google names support in tools such as transformers, sentence-transformers, MLX, vLLM, llama.cpp, Ollama and LMStudio, plus Unsloth guidance for fine-tuning and Qdrant for storing vectors. It also points to demo apps in Google AI Edge Gallery for searching a media library and finding moments in videos.

Because the model uses the same text-splitting component and audio encoder as Gemma 4, Google says the two can run side by side in one pipeline with a smaller combined memory footprint, for example to build retrieval-augmented generation (RAG) systems that work offline on a device. The first EmbeddingGemma passed 20 million downloads, according to the company, and developers used it for on-device search tools and privacy-focused RAG pipelines.

Why it matters

Embedding models are the search layer behind semantic search and retrieval-augmented generation. A small open model that covers several media types and fits in phone memory lets developers build that kind of search without sending personal photos, recordings or files to a server, and the Apache 2.0 license allows commercial use. The quality figures so far come only from Google.

Timeline

  1. · Published

Topics#Google#Gemma#Embeddings#On-device AI#Open models

Sources

This story draws on the following sources. Read them for full context.

  1. 1Google · Primary sourceEmbeddingGemma 2: an open, lightweight multimodal embedding modelblog.google
  2. 2Google DeepMind · Primary sourceEmbeddingGemma 2: an open, lightweight multimodal embedding modeldeepmind.google