Google’s EmbeddingGemma 2 brings text, code, images, video and audio into one embedding space

Google released 740M-parameter EmbeddingGemma 2 under Apache 2.0, with modular multimodal encoders and publisher-reported gains on code retrieval benchmarks.

2 min read

Google DeepMind announced EmbeddingGemma 2 on October 6, a compact model for placing text, code, images, video and audio into one shared 768-dimensional embedding space. Unlike a generative chatbot, it produces vectors for retrieval, semantic search, clustering and classification—useful when an app needs to search across mixed media without sending source files to a cloud service.

The full model has 740 million parameters: a 270M text-and-code backbone, plus optional 170M vision and 300M audio encoders. Developers can load only the modalities they need. Google lists an 8,192-token shared context window and says the model is available as downloadable weights on Hugging Face and Kaggle. The model card lists Apache 2.0, so this is an actual permissive-weight release rather than API-only access.

Google reports 78.68 on MTEB Code (v1), versus 68.76 for EmbeddingGemma 1, using NDCG@10; multilingual MTEB v2 is 61.36 versus 61.15. These are publisher-reported results in the model card, not independent confirmation. The comparison is specifically against its predecessor on those stated benchmark versions; it does not establish that EmbeddingGemma 2 is best across embedding workloads. For other modalities the card reports scores on MIEB, MMEB, MSEB and MAEB, but cross-benchmark scores should not be compared as if they shared one scale.

The model also supports Matryoshka truncation to 512, 256 or 128 dimensions, reducing vector storage at some quality cost. Google says 128 dimensions fit text-focused use better; multimodal deployments should validate retrieval quality on their own data. The 8K budget is shared across media, so adding video frames, images or audio reduces room for other inputs.

Sources: Google announcement; Google model card and benchmark table; Hugging Face weights.

multimodal embeddingsGoogleMTEBApache 2.0EmbeddingGemma 2