# EmbeddingGemma 2: a local search prototype across text and media
> Google's EmbeddingGemma 2 brings text, image, video and audio embeddings into one local retrieval model. This documentation-based guide maps the setup and evaluation decisions.
By BIG CHANGE Editorial
Published: 2026-10-06T20:48:02.995Z
Updated: 2026-10-06T20:48:02.995Z
Canonical: https://bigchange.ai/blog/embeddinggemma-2-local-multimodal-search-guide

A person considers a small personal collection of photos and records beside a phone. Conceptual AI-generated illustration by BIG CHANGE; it does not show an EmbeddingGemma 2 interface or tested search result.
Google released EmbeddingGemma 2 on October 6, 2026, with downloadable weights under Apache 2.0. The model turns text, code, images, video and audio into vectors that can be compared in one 768-dimensional space. For a developer with a local collection, the practical question is which encoders to load, how to represent queries and files, and whether the resulting matches are useful on that collection. This guide follows Google's published instructions and model card; BIG CHANGE has not run the model or measured retrieval quality.
## The big change
- **What changed:** EmbeddingGemma's earlier model handled text. Version 2 adds vision and audio encoders, allowing a text query to be compared with supported media in a shared vector space without first generating captions or transcripts. Google has released the weights; [ML Kit service access on Android is planned for the coming weeks](https://developers.googleblog.com/google-ai-edge-with-embeddinggemma-2/).
- **Why it matters:** A developer can load the encoders a local collection needs and build one retrieval index across its supported file types. [Google's AI Edge Gallery](https://developers.googleblog.com/google-ai-edge-with-embeddinggemma-2/) demonstrates local media vectors in SQLite with cosine-similarity ranking. An application still needs file ingestion, vector storage, result display and checks for misleading matches.
- **What to watch:** The immediate decision is whether representative files rank well within a device's memory and latency budget. The later Android service could change the integration route; its announced timing does not establish availability today.
## Choose the inputs before loading the model
The [developer guide](https://developers.googleblog.com/en/embeddinggemma-2-the-developer-guide/) documents Sentence Transformers 6.1.0 or later and this installation command for all supported media:
```bash
pip install -U sentence-transformers[image,audio,video] transformers
```
The exact checkpoint is `google/embeddinggemma-2`. Its default Sentence Transformers load includes all encoders. For a collection of text and code only, the guide says to leave out vision and audio at load time:
```python
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"vision_config": None, "audio_config": None},
)
```
Keeping vision while omitting audio requires `config_kwargs={"audio_config": None}`; keeping audio while omitting vision requires `config_kwargs={"vision_config": None}`. These options change which weights load, while the resulting vectors remain in the model's shared space. A text-only query can therefore be compared with an image encoded by a configuration that includes vision, according to the guide. Google describes the model as suited to consumer devices, but the published Python example is a library quickstart, not a measured memory or latency guarantee for a reader's device.
## Encode queries and collection items differently
For text retrieval, the [model card](https://ai.google.dev/gemma/docs/embeddinggemma/model_card_2) specifies separate task instructions for queries and corpus items. Sentence Transformers supplies them through `prompt_name`:
```python
query_vector = model.encode("northern lights", prompt_name="SearchQuery")
document_vector = model.encode(
"Charged particles from the sun cause the northern lights.",
prompt_name="Document",
)
score = model.similarity(query_vector, document_vector)
```
This documented snippet produces vectors and a similarity score; it is not an observed result from our environment. `Document` inserts `title: none`. If an item has a real title, Google instructs developers to format it explicitly as `title: {title} | text: {content}`. Code search uses the `CodeRetrieval` query instruction and a title or filename with the indexed code. Classification, clustering and sentence similarity use their own symmetric instructions. The model card says omitted text instructions still run but reduce precision.
The earlier `model` has no vision or audio encoder. Before encoding an image, load the [guide's text-and-vision configuration](https://developers.googleblog.com/en/embeddinggemma-2-the-developer-guide/) from the same checkpoint, then use that model for the media and query. This sequence follows the documentation; we have not run it:
```python
media_model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"audio_config": None},
)
image_vector = media_model.encode({"image": "photo.jpg"})
query_vector = media_model.encode("a photo query", prompt_name="SearchQuery")
score = media_model.similarity(query_vector, image_vector)
```
For audio, reload with its encoder enabled by omitting `audio_config: None`; the full default load includes vision and audio. Google's guide passes a dictionary such as `{"audio": "recording.wav"}` without a text task prefix. An input can also combine text and media by placing `<|image|>`, `<|video|>` or `<|audio|>` markers in the text and supplying the matching files in the dictionary. That call produces one vector for the combined input. The model card specifies 16 kHz mono audio and says video is sampled at one frame per second by default. Images, video frames, audio and text all spend from the same 8,192-token context budget, so a long clip or mixed item needs input planning rather than an assumption that every file fits.
## Pick a vector size, then check the trade-off
The native output is 768 dimensions. Google's model card supports 512, 256 and 128 dimensions by truncating the vector. The guide documents `truncate_dim=256` with `normalize_embeddings=True` in `model.encode()`. Query and indexed-item vectors must use the same dimension; shortened vectors need normalization before cosine comparison. The card reports that 128 dimensions loses substantially more multimodal quality than 256, and recommends 128 mainly for text workloads. Those figures are Google's full-precision benchmark results, not a prediction of recall on a particular collection.
Numerical precision is another setup choice. The card says to use `bfloat16` on hardware that supports it or `float32` elsewhere, including most CPUs. It warns that `float16` can yield NaNs or silently degraded embeddings. Its selective-loading parameter counts describe model components, while actual device memory and search speed depend on execution path, precision, index and input size.
## What to evaluate before using the results
A prototype can index a small permitted sample, hold back representative queries, and inspect whether the right items appear near the top. Include text-to-text, text-to-image or text-to-audio cases only for media actually in the collection. Compare 768 and 256 dimensions on the same queries, check failures across relevant languages and document types, and measure memory, indexing time and query latency on the intended device. These are evaluation steps for the reader; BIG CHANGE has not performed them.
The [model card's published results](https://ai.google.dev/gemma/docs/embeddinggemma/model_card_2) put EmbeddingGemma 2 at 78.68 versus 68.76 for the earlier model on MTEB code v1, while multilingual MTEB v2 scores are 61.36 versus 61.15. The card also reports image, video and audio results without a version 1 comparator for those modalities. All cited benchmark results use the full-precision checkpoint. They establish Google's evaluation conditions, not accuracy on a local archive. The card cautions that quality can vary by language and task, and that retrieval applications need their own filtering and fairness checks.
The weights are released under Apache 2.0 and the guide links to their [Hugging Face repository](https://huggingface.co/google/embeddinggemma-2). Google's materials give a local execution path but no universal deployment price: device purchase, compute, storage and engineering remain application costs. For this reader task, the next decision is whether the model ranks the right items on representative files within the device budget.
## Sources & further reading
- [Google AI Edge announcement](https://developers.googleblog.com/google-ai-edge-with-embeddinggemma-2/), October 6, 2026: release stage, AI Edge Gallery's local media-search implementation and the announced ML Kit timing. The device measurements in this post are Google's measurements.
- [EmbeddingGemma 2 developer guide](https://developers.googleblog.com/en/embeddinggemma-2-the-developer-guide/), October 6, 2026: library version, loading options, encode examples and vector truncation instructions.
- [EmbeddingGemma 2 model card](https://ai.google.dev/gemma/docs/embeddinggemma/model_card_2): model architecture, input limits, task prompts, precision warning, benchmark tables and limitations. Its results are Google-reported.
- [Google model repository](https://huggingface.co/google/embeddinggemma-2): the named checkpoint and published model files. Repository availability does not validate performance in an application.
## Sources
- [Bring multimodal semantic search to the edge with EmbeddingGemma 2](https://developers.googleblog.com/google-ai-edge-with-embeddinggemma-2/) — Launch stage, Google AI Edge Gallery implementation and upcoming ML Kit access; device figures are vendor measurements.
- [EmbeddingGemma 2: The Developer Guide](https://developers.googleblog.com/en/embeddinggemma-2-the-developer-guide/) — Sentence Transformers 6.1.0+, load configurations, task and media encode examples and vector truncation.
- [EmbeddingGemma 2 model card](https://ai.google.dev/gemma/docs/embeddinggemma/model_card_2) — Architecture, input limits, prompts, precision, Google-reported benchmark scores, risks and limitations.
- [google/embeddinggemma-2 checkpoint](https://huggingface.co/google/embeddinggemma-2) — Named model repository and published model files/license.
The BIG CHANGE newsletter
The big picture. At your pace.
Recent stories on AI and robotics, the shifts worth watching and practical ideas to use. Choose a daily briefing, weekly digest or monthly perspective.
Sent at 09:00 Belgrade time: daily, Mondays or the first of the month. Your first edition arrives at the next scheduled send after you confirm.
Your privacy, your choice.
Necessary storage supports site security and remembers your choices. Optional Google Analytics stays off until you allow it. You can read every story with necessary storage only. Privacy details