Google's EmbeddingGemma 2 Packs Text, Image, Video and Audio Into 740M Parameters Under Apache 2.0
A.I. / news
Google's EmbeddingGemma 2 Packs Text, Image, Video and Audio Into 740M Parameters Under Apache 2.0
The Oct. 6 release is ungated and fits in 567MB when quantised, but every score in it is Google's own and none has been reproduced outside.

Google DeepMind released EmbeddingGemma 2 on Oct. 6, a 740-million-parameter embedding model that takes text, images, video, audio and code in one vector space. The weights are on Hugging Face and Kaggle under Apache 2.0, and the Hugging Face repository is not gated.
The announcement was written by Sahil Dua and Henrique Schechter Vera, both research engineers at Google DeepMind, in Google's developer blog. The model card is on Hugging Face. As of Wednesday morning the card showed 731 likes and 7,562 downloads.
An embedding model turns content into a list of numbers so a search system can find similar items. It does not write text. The first EmbeddingGemma, Google said, passed 20 million downloads.
What EmbeddingGemma 2 is, and what it needs
The card breaks the 740M into three parts: a 270M text backbone, 170M for vision and 300M for audio. Google says it is built on the Gemma 4 architecture. MindStudio reports that Gemma 4 itself uses Apache 2.0, so the licence matches its parent.
Loading only the text encoder gives the 270M version, which is the one aimed at phones. Google's blog says it needs about 191MB of RAM on a Pixel 11 Pro, and about 567MB for the whole multimodal model, both with quantisation.
Output vectors are 768 numbers long and can be cut to 512, 256 or 128. Google says that truncation, called Matryoshka representation learning, gives up to a 6x cut in vector storage cost.
| Item | Limit on the model card |
|---|---|
| Context window | 8,192 tokens, shared across all inputs |
| Audio | 25 tokens a second, about 327 seconds, mono 16 kHz |
| Video | 140 tokens a frame at 1 frame a second, about 58 frames |
| Languages | 100+ for text |
The benchmark numbers are all Google's
Google says EmbeddingGemma 2 has leading scores among embedders under 1 billion parameters that handle several media types. The blog names MTEB Code and the Massive Audio Embedding Benchmark. The one gain it quantifies is MTEB Code, from 68.76 to 78.68 NDCG@10, a rise of 9.92 points.
The card adds scores for four more tests. These are vendor-supplied. No independent group is named as having run any of them, and The Terminal found no outside evaluation.
- MTEB Code (NDCG@10)78.68 points
- MSEB audio retrieval (MRR@10)69.54 points
- MIEB lite images (mean)64.64 points
- MTEB multilingual (mean)61.36 points
- MMEB v2 video (Hit@1)50.67 points
Source: Hugging Face model card for google/embeddinggemma-2, accessed 2026-10-07
The bars use different metrics, so they say nothing about which medium the model handles best. They are listed so a reader can check them against a rerun.
Limits Google lists
The card says performance varies across languages despite the 100-plus claim. It also warns that leaving out the task instruction prefixes lowers embedding quality, and that float16 precision causes NaN errors, so the model needs bfloat16 or float32.
Training data has a January 2025 cutoff and includes web documents in more than 140 languages, code, images, video and audio. The card says the data was filtered for child sexual abuse material and sensitive personal data. It does not give dataset sizes or sources.
Where it runs

Google lists integrations with MediaPipe, LiteRT, sentence-transformers, Ollama and Qdrant. LiteRT builds sit in Hugging Face's litert-community. A Model Garden listing is listed by Google as not yet live.
The model sits beside other open-weight releases that this site has tracked, including Reflection's Beam and the Clef and Jev decision models. Neither is a multimodal embedder.
Google has not said when the Model Garden listing will open or whether a larger variant is planned.
Sources
More in A.I.
- 01LTX-2.5's Free Commercial Licence Stops at $10 Million in Group Revenue, and Its Card Asks for Contact DetailsLightricks' open-weight video and audio model is free for production use below that line, but the revenue test counts parent companies and affiliates, and the card publishes no benchmark scores.
- 02Gemini 4 Argon Costs $1.99 a Task at Promo Price Against $0.72 for GPT-6.1 Sol, and Is Not Yet on SaleGoogle's introductory $2 and $10 rates match OpenAI's Sol per token, but Artificial Analysis figures cited by eesel show Argon writing 62,000 output tokens a task where GPT-6 Astra writes 27,000.
- 03EmbeddingGemma 2 Embeds Text, Images, Audio and Video in 567MB of RAM on a Pixel 11 ProGoogle DeepMind's Apache 2.0 embedding model has 740M parameters in three modular pieces, and the only benchmark number its launch post prints is a 9.92-point gain on MTEB Code.
- 04Kolibri-1 Is Apache 2.0 and Fits on One B200, but Trails Qwen3.8 27B by 9.1 Points in GermanAleph Alpha's 78B-parameter mixture-of-experts model activates 3.46B parameters per token, and its own model card shows a larger dense Qwen ahead on every headline benchmark.