AI KNOWLEDGE DESK

Models · entities · concepts · comparisons · practical tools

GETLLMS.ORG
← All entities
Multimodal embedding modelsChecked

EmbeddingGemma 2

EmbeddingGemma 2 is Google’s October 6, 2026 open multimodal embedding model, released as google/embeddinggemma-2 under Apache 2.0. It maps text/code, images, video and audio into numerical vectors rather than generating chat answers or images.

Why it matters

A shared vector space lets local applications retrieve across media without first converting every input into a text caption. The practical choices are which encoders to load, how to spend the shared input budget and how many vector dimensions your index can store.

Modular weights and vector output

The full model has 740M parameters: 270M for text/code, 170M for vision and 300M for audio. Load text-only, text-plus-vision, text-plus-audio or all encoders according to the task. Output is a 768-dimensional embedding, with supported 512, 256 or 128 dimensional truncation.

One shared input budget

Inputs share an 8,192-token context. Default accounting is about 280 tokens per image, 140 per video frame and 25 per audio second; mixed prompts spend the same budget. Audio should be 16 kHz mono, and default video sampling is one frame per second. These limits are input-processing contracts, not generated-media capabilities.

Normalize and match index dimensions

After truncation, L2-normalize the shorter vector again. Queries and indexed documents must use the same dimensions. A mismatched dimension or unnormalized truncated vector can quietly degrade retrieval even when the API call succeeds. Keep task prefixes and modality preprocessing consistent between indexing and querying.

Local deployment and official demonstrations

Use the official weights and Google’s Sentence Transformers guide to test local search, RAG, classification or clustering. Google’s AI Edge Gallery demonstrates Instant Media Search and Video Moments Finder; ML Kit support announced for future weeks is not current GA. Test memory, latency and retrieval quality on the actual device and corpus.

Useful first jobs

  • Local cross-modal semantic search.
  • Retrieve images, video moments or audio using a query.
  • Embed a local corpus for RAG.
  • Zero-shot classification and clustering with reviewed labels.

Frequently asked questions

Is EmbeddingGemma 2 a chat or image generator?

No. It outputs numerical embeddings for retrieval, similarity, classification and clustering.

Can I use 128-dimensional vectors?

Yes. The model supports 128, 256, 512 and 768 dimensions. Normalize after truncation and keep query and document dimensions identical.

Sources and evidence

Google’s model card and official Hugging Face repository establish released weights, license, dimensions and limits. The developer guide explains modular loading and retrieval use. Vendor performance and device-memory measurements are not independent proof on your hardware.

Continue reading