EmbeddingGemma 2: An open, lightweight multimodal embedding model
Google DeepMind's new model is built on Gemma 4 and uses modular parts. Text-only workloads need as little as 270M parameters, and optional vision (170M) and audio (300M) encoders add full multimodal support. It outputs 768-dimension vectors that can be truncated to 128 through Matryoshka representation learning, and it has an 8,000 token context window. Weights have been on Hugging Face and Kaggle since October 6, 2026. The first EmbeddingGemma handled only text, and the new context window is 4x larger than that version's. For hosted use, Google also offers Gemini Embedding 2, which produces embeddings with a maximum dimensionality of 3,072 through its API. Google says the new open model reduces the latency and memory overhead of chaining separate image captioning, speech-to-text, and text-embedding models. For founders, this means cross-modal search, retrieval-augmented generation (RAG) and routing can run offline on phones without per-call API costs, and the license permits commercial use. Startups that sell multimodal embedding APIs now compete with a free model they will have to beat on quality or scale.