Google's EmbeddingGemma 2 Adds Images, Video and Audio to a 740M Embedder
A.I. / news
Google's EmbeddingGemma 2 Adds Images, Video and Audio to a 740M Embedder
The Apache 2.0 weights lift the code-retrieval score by 9.9 points but move the multilingual text score by only 0.21.

Google released EmbeddingGemma 2 on Oct. 6, an open embedding model of about 740 million parameters that maps text, code, images, video and audio into one shared space. The weights are downloadable without a gate, and the model card lists Apache 2.0, with a link to the Gemma terms and a note that the Gemma Prohibited Use Policy applies to deployments.
EmbeddingGemma 1 handled text only, OfficeChai reported on Oct. 7, and Google says it was downloaded more than 20 million times. The card shows 29,185 downloads for the new model over the last month. All benchmark figures below are vendor-supplied.

Which parts you load decide the memory bill
The model is modular. A 270-million-parameter text backbone is the core, a 170-million-parameter vision encoder and a 300-million-parameter audio encoder are optional, and the card says to load only the encoders a job needs.
| Configuration | Effective parameters |
|---|---|
| Text only | About 270M |
| Text and image | About 440M |
| Text and audio | About 570M |
| Full model | About 740M |
Google's quantised memory figures, measured on a Pixel 11 Pro, are about 191 MB for text only and about 567 MB for the full model, OfficeChai reported. The photograph above shows the older Pixel 9 Pro XL, not the 11 Pro Google tested. The card gives no RAM or VRAM figures of its own.
The context window is 8,192 tokens shared across every modality. An image costs about 280 tokens by default, video about 140 tokens per frame and audio about 25 tokens per second. A one-minute clip of audio therefore takes about 1,500 of the 8,192 tokens.
The card warns against float16, which it says produces NaN or degraded embeddings. It recommends bfloat16 or float32.
The gain is in code retrieval, not text
The card compares the new model with EmbeddingGemma 1 at 768 dimensions on two text benchmarks.
- EmbeddingGemma 278.68 points
- EmbeddingGemma 168.76 points
Source: Hugging Face model card for google/embeddinggemma-2, accessed 2026-10-09 (vendor-supplied)
On MTEB multilingual v2 the scores are 61.36 and 61.15, a gap of 0.21 points. The card reports no earlier-model comparison for its image, video or audio benchmarks, so the multimodal scores have no baseline. They include 64.64 on MIEB lite and 49.39 on MAEB.
Google also supports shrinking the vectors with Matryoshka learning, which lets a developer truncate the 768-number output. The card shows the cost in multilingual MTEB.
| Dimensions | MTEB multilingual v2 |
|---|---|
| 768 | 61.36 |
| 256 | 60.41 |
| 128 | 57.89 |
The card says 128 dimensions suits text-only work. OfficeChai reported that Google says the shorter vectors cut local vector storage by up to six times.
Training data and tooling
The training mix is web documents in more than 140 languages, code, images, video, audio and paired cross-modal data, with a cutoff of January 2025, according to the card. It lists filtering for child sexual abuse material, sensitive data, quality and safety.
Google did not publish the data itself. Weights are on Hugging Face and Kaggle, and the model runs in transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio, OfficeChai reported. A listing in Google's Model Garden is marked as upcoming.
Other open-weight coverage on the site includes the Qwen3.8-27B report and the report on Nous Research's $90 million round.
Google has not said when the Model Garden listing will go live.
Sources
More in A.I.
- 01OpenAI Fires Three Safety Staff Nine Days After Publishing Outside-Audit PrinciplesThe company confirmed the dismissals on Oct. 1 but has not said what information moved, who received it, or whether the three first raised concerns internally.
- 02Cloudflare's Clef Beats TypeSafe's Jev on Three Tests and Loses on OneThe Apache 2.0 decision model costs nearly six times as much per million tokens as Jev, and its smaller sibling trails badly on one intent benchmark.
- 03Qwen3.8-27B Has 6.78M Downloads and an Apache 2.0 Licence, but Its Card Gives Training Data One LineAlibaba's 27B dense model scores 61.7 on SWE-bench Pro by its own table. The card does not say what it was trained on or what hardware it needs.
- 04ChatGPT's Intelligent UI Skips the Pro Thinking Level and Older Desktop Apps, and Publishes No Usage FiguresOpenAI put GPT-6 into the Chat tab on October 7 with generated buttons, forms and charts. The announcement gives one speed number and no accuracy figure.