Google released EmbeddingGemma 2 on October 6, 2026, with downloadable weights under Apache 2.0. The model turns text, code, images, video and audio into vectors that can be compared in one 768-dimensional space. For a developer with a local collection, the practical question is which encoders to load, how to represent queries and files, and whether the resulting matches are useful on that collection. This guide follows Google's published instructions and model card; BIG CHANGE has not run the model or measured retrieval quality.
The big change
- What changed: EmbeddingGemma's earlier model handled text. Version 2 adds vision and audio encoders, allowing a text query to be compared with supported media in a shared vector space without first generating captions or transcripts. Google has released the weights; ML Kit service access on Android is planned for the coming weeks.
- Why it matters: A developer can load the encoders a local collection needs and build one retrieval index across its supported file types. Google's AI Edge Gallery demonstrates local media vectors in SQLite with cosine-similarity ranking. An application still needs file ingestion, vector storage, result display and checks for misleading matches.
- What to watch: The immediate decision is whether representative files rank well within a device's memory and latency budget. The later Android service could change the integration route; its announced timing does not establish availability today.
Choose the inputs before loading the model
The developer guide documents Sentence Transformers 6.1.0 or later and this installation command for all supported media:
pip install -U sentence-transformers[image,audio,video] transformersThe exact checkpoint is google/embeddinggemma-2. Its default Sentence Transformers load includes all encoders. For a collection of text and code only, the guide says to leave out vision and audio at load time:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"vision_config": None, "audio_config": None},
)Keeping vision while omitting audio requires config_kwargs={"audio_config": None}; keeping audio while omitting vision requires config_kwargs={"vision_config": None}. These options change which weights load, while the resulting vectors remain in the model's shared space. A text-only query can therefore be compared with an image encoded by a configuration that includes vision, according to the guide. Google describes the model as suited to consumer devices, but the published Python example is a library quickstart, not a measured memory or latency guarantee for a reader's device.
Encode queries and collection items differently
For text retrieval, the model card specifies separate task instructions for queries and corpus items. Sentence Transformers supplies them through prompt_name:
query_vector = model.encode("northern lights", prompt_name="SearchQuery")
document_vector = model.encode(
"Charged particles from the sun cause the northern lights.",
prompt_name="Document",
)
score = model.similarity(query_vector, document_vector)This documented snippet produces vectors and a similarity score; it is not an observed result from our environment. Document inserts title: none. If an item has a real title, Google instructs developers to format it explicitly as title: {title} | text: {content}. Code search uses the CodeRetrieval query instruction and a title or filename with the indexed code. Classification, clustering and sentence similarity use their own symmetric instructions. The model card says omitted text instructions still run but reduce precision.
The earlier model has no vision or audio encoder. Before encoding an image, load the guide's text-and-vision configuration from the same checkpoint, then use that model for the media and query. This sequence follows the documentation; we have not run it:
media_model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"audio_config": None},
)
image_vector = media_model.encode({"image": "photo.jpg"})
query_vector = media_model.encode("a photo query", prompt_name="SearchQuery")
score = media_model.similarity(query_vector, image_vector)For audio, reload with its encoder enabled by omitting audio_config: None; the full default load includes vision and audio. Google's guide passes a dictionary such as {"audio": "recording.wav"} without a text task prefix. An input can also combine text and media by placing <|image|>, <|video|> or <|audio|> markers in the text and supplying the matching files in the dictionary. That call produces one vector for the combined input. The model card specifies 16 kHz mono audio and says video is sampled at one frame per second by default. Images, video frames, audio and text all spend from the same 8,192-token context budget, so a long clip or mixed item needs input planning rather than an assumption that every file fits.
Pick a vector size, then check the trade-off
The native output is 768 dimensions. Google's model card supports 512, 256 and 128 dimensions by truncating the vector. The guide documents truncate_dim=256 with normalize_embeddings=True in model.encode(). Query and indexed-item vectors must use the same dimension; shortened vectors need normalization before cosine comparison. The card reports that 128 dimensions loses substantially more multimodal quality than 256, and recommends 128 mainly for text workloads. Those figures are Google's full-precision benchmark results, not a prediction of recall on a particular collection.
Numerical precision is another setup choice. The card says to use bfloat16 on hardware that supports it or float32 elsewhere, including most CPUs. It warns that float16 can yield NaNs or silently degraded embeddings. Its selective-loading parameter counts describe model components, while actual device memory and search speed depend on execution path, precision, index and input size.
What to evaluate before using the results
A prototype can index a small permitted sample, hold back representative queries, and inspect whether the right items appear near the top. Include text-to-text, text-to-image or text-to-audio cases only for media actually in the collection. Compare 768 and 256 dimensions on the same queries, check failures across relevant languages and document types, and measure memory, indexing time and query latency on the intended device. These are evaluation steps for the reader; BIG CHANGE has not performed them.
The model card's published results put EmbeddingGemma 2 at 78.68 versus 68.76 for the earlier model on MTEB code v1, while multilingual MTEB v2 scores are 61.36 versus 61.15. The card also reports image, video and audio results without a version 1 comparator for those modalities. All cited benchmark results use the full-precision checkpoint. They establish Google's evaluation conditions, not accuracy on a local archive. The card cautions that quality can vary by language and task, and that retrieval applications need their own filtering and fairness checks.
The weights are released under Apache 2.0 and the guide links to their Hugging Face repository. Google's materials give a local execution path but no universal deployment price: device purchase, compute, storage and engineering remain application costs. For this reader task, the next decision is whether the model ranks the right items on representative files within the device budget.
Sources & further reading
- Google AI Edge announcement, October 6, 2026: release stage, AI Edge Gallery's local media-search implementation and the announced ML Kit timing. The device measurements in this post are Google's measurements.
- EmbeddingGemma 2 developer guide, October 6, 2026: library version, loading options, encode examples and vector truncation instructions.
- EmbeddingGemma 2 model card: model architecture, input limits, task prompts, precision warning, benchmark tables and limitations. Its results are Google-reported.
- Google model repository: the named checkpoint and published model files. Repository availability does not validate performance in an application.



