EmbeddingGemma 2 Embeds Text, Images, Audio and Video in 567MB of RAM on a Pixel 11 Pro
A.I. / news
EmbeddingGemma 2 Embeds Text, Images, Audio and Video in 567MB of RAM on a Pixel 11 Pro
Google DeepMind's Apache 2.0 embedding model has 740M parameters in three modular pieces, and the only benchmark number its launch post prints is a 9.92-point gain on MTEB Code.

Google DeepMind released EmbeddingGemma 2 on Oct. 6, an Apache 2.0 embedding model that maps text, code, images, audio and video into one 768-dimension space. Google's launch post says the quantized full model uses about 567MB of active RAM on a Pixel 11 Pro, and the text-only part about 191MB.
An embedding model turns a piece of content into a list of numbers so a search system can find similar items by distance. The weights are on Hugging Face as google/embeddinggemma-2 and on Kaggle. Availability in Google's Gemini Enterprise Agent Platform Model Garden is listed as "coming soon".
What the 740M parameters are made of
The model is built on the Gemma 4 architecture and has 740 million parameters in total. Per the post and the Hugging Face model card, that splits into a 270M-parameter text model, a 170M vision encoder and a 300M audio encoder. A developer who only needs text can load the 270M piece alone, which is where the 191MB figure comes from.
| Component | Parameters | Quantized RAM on Pixel 11 Pro |
|---|---|---|
| Text only | 270M | about 191MB |
| Full multimodal | 740M | about 567MB |
Output vectors are 768 dimensions and can be cut to 512, 256 or 128 using Matryoshka representation learning, a training method that makes the first numbers in the vector carry the most information. Google says truncation saves up to 6x in storage. The card says truncated vectors must be re-normalized.
What changed from EmbeddingGemma 1
The first EmbeddingGemma, which Hugging Face covered on Sept. 4, 2025, had 308M parameters, took text only and had a 2,048-token window. EmbeddingGemma 2 raises the window to 8,192 tokens, shared across all modalities. Google says that holds up to 5.5 minutes of audio, 29 images or 58 video frames, or a mix.
- EmbeddingGemma 168.76 points
- EmbeddingGemma 278.68 points
Source: Google launch post, vendor-supplied, accessed 2026-10-08
That code-retrieval gain, from 68.76 to 78.68, is the only score the launch post prints. The charts for MTEB, the audio benchmark MAEB and the image benchmark MIEB Lite carry no numbers in the text. The model card has them: 61.36 on multilingual MTEB v2, 68.46 on English MTEB v2, 64.64 on MIEB Lite and 49.39 on MAEB. All are Google's figures, and the card lists no independent reruns.
The claims Google makes and does not support
The post calls EmbeddingGemma 2 "the most capable model for on-device multimodal embeddings" and says it "even outperforms some specialist models more than twice its size". It does not name those models. It also says the model is "built from the same technology as Gemini Embedding models" but gives no head-to-head with Gemini Embedding 2, Google's hosted multimodal embedder.
The Hugging Face card lists text support for more than 100 languages in one place and more than 140 in its training-data section, so the language count is not settled by the document itself.

Where it runs
Google lists support in transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio. On-device paths include LiteRT, MediaPipe and transformers.js with WebGPU. Qdrant is named for vector storage. The post names two Google DeepMind research engineers, Sahil Dua and Henrique Schechter Vera.
The scale is the point of comparison. Microsoft's Surface Laptop Ultra, covered on the hardware desk, needs its 128GB tier for a 120B-parameter local model. EmbeddingGemma 2 fits in a phone's spare memory, at the cost of being an embedder rather than a chat model; for an open chat model under the same Apache 2.0 licence, see Qwen3.8-27B.
The release date for the Model Garden listing is the open item. Google gave none.
Sources
More in A.I.
- 01LTX-2.5's Free Commercial Licence Stops at $10 Million in Group Revenue, and Its Card Asks for Contact DetailsLightricks' open-weight video and audio model is free for production use below that line, but the revenue test counts parent companies and affiliates, and the card publishes no benchmark scores.
- 02Gemini 4 Argon Costs $1.99 a Task at Promo Price Against $0.72 for GPT-6.1 Sol, and Is Not Yet on SaleGoogle's introductory $2 and $10 rates match OpenAI's Sol per token, but Artificial Analysis figures cited by eesel show Argon writing 62,000 output tokens a task where GPT-6 Astra writes 27,000.
- 03Kolibri-1 Is Apache 2.0 and Fits on One B200, but Trails Qwen3.8 27B by 9.1 Points in GermanAleph Alpha's 78B-parameter mixture-of-experts model activates 3.46B parameters per token, and its own model card shows a larger dense Qwen ahead on every headline benchmark.
- 04Qwen3.8-27B Is Apache 2.0 With a 262,144-Token Window and Already Has 482 Finetunes, Including Cloudflare's ClefAlibaba's dense 27B model claims 61.7 on SWE-bench Pro against 53.4 for Opus 4.6 Max, and the model card says nothing about training data.