Google released EmbeddingGemma 2 on Tuesday, putting text, code, image, video, and audio retrieval into a 740-million-parameter open model that used about 567MB of active RAM with quantization in the company’s testing on a Pixel 11 Pro.Built on Gemma 4, the model maps all five input types into the same 768-dimensional vector space. Images no longer have to be captioned and audio doesn’t have to be transcribed before either can be searched alongside text. The company released the weights under Apache 2.0, and on-device deployment through LiteRT and MediaPipe Tasks is available now. An Android ML Kit integration, with NPU acceleration on devices that support it, is due in the coming weeks.Modular encoders, one vector spaceThe full model tops out at 740 million parameters, but developers only load the encoders their data requires. Text and code run on a 270-million-parameter base, which Google measured at about 191MB of active RAM on the same phone. With the vision encoder for images and video, the model reaches 440 million parameters; adding the audio encoder instead brings it to 570 million, while loading both takes it to the full 740 million. Every configuration projects into the same embedding space, so a team that starts with a text-only index can add image or audio search later without re-embedding anything it has already stored.Google also quadrupled the context window from 2,048 to 8,192 tokens, which the company says covers as much as 5.5 minutes of audio, 29 images or 58 video frames in a single input. Google also quadrupled the context window from 2,048 to 8,192 tokens, which the company says covers as much as 5.5 minutes of audio, 29 images or 58 video frames in a single input. Video is sampled at one frame per second by default, so those 58 frames amount to just under a minute of footage.In Google’s Video Moments Finder demo, video frames and audio chunks are indexed locally and then searched with plain text to land on a specific moment, with no captions or transcripts generated along the way. Instant Media Search applies the same approach to the photos and videos on a phone, storing the embeddings in SQLite and updating results as the user types.Matryoshka shrinks the indexOn a phone, the index can compete with the model itself for space, since a million 768-dimensional bfloat16 vectors take up roughly 1.5GB. Google trained EmbeddingGemma 2 with Matryoshka Representation Learning, which lets developers truncate those embeddings to 512, 256 or 128 dimensions without retraining anything.At 256 dimensions, that million-vector index drops to about 500MB. Google says the shorter embeddings keep most of their full-size quality on text and code and roughly 95% on image, video and speech retrieval. Google says the shorter embeddings keep most of their full-size quality on text and code and roughly 95% on image, video and speech retrieval. The trade gets steeper at 128 dimensions, where Google puts text and code at around 90% but multimodal retrieval at about 75%, and the company advises testing that setting on real data before relying on it for multimodal queries. Giving up a sliver of retrieval quality for efficiency has become a familiar pitch in embedding releases this fall, as when Cohere’s faster query model barely dented retrieval quality in its own tests.On-device code searchCode runs on the same 270-million-parameter base as text, and Google reports an MTEB Code score of 78.68 for EmbeddingGemma 2, up from 68.76 for the original EmbeddingGemma.To show how that holds up in an agent workflow, Google embedded the Hugging Face Transformers repository with the text-only setup and paired the index with Gemma 4 26B A4B running in Pi. That is the same agent harness behind a workaround for an MCP server that used 18,000 tokens before doing anything. In Google’s setup, EmbeddingGemma 2 handled retrieval across the repository while the larger model drove the agent.The same embedding model can also handle classification through MediaPipe Decision, which compares an incoming embedding against candidate descriptions instead of generating a response. In an on-device chess demo, Google used it to evaluate 500 options per turn in under 100 milliseconds. If you’ve watched an agent burn tokens on a decision that never needed a generated answer in the first place, this gives it a much cheaper way to make the call.Code runs on the same 270-million-parameter base as text, and Google reports an MTEB Code score of 78.68 for EmbeddingGemma 2, up from 68.76 for the original EmbeddingGemma.Leaner local RAG pipelinesEmbeddingGemma 2 borrows Gemma 4’s text tokenizer and audio encoder architecture, meaning, the two models need less memory when they’re running together on a device. So far, Google has shown multimodal retrieval working on its own flagship phone. We’ll just have to wait and see what happens with bigger indexes, different hardware, and apps that aren’t Google demos.The post Your phone’s vector index might be bigger than the AI model running it appeared first on The New Stack.