Google DeepMind wants your phone to understand your videos, voice notes, and photo library without sending everything to the cloud.
The company launched EmbeddingGemma 2 on Tuesday, a 740 million-parameter open-weight model that maps text, code, images, video, and audio into a shared embedding space. The model is designed to let developers build multimodal search and retrieval features that run directly on consumer devices.
Google said the first EmbeddingGemma has passed 20 million downloads since its 2025 launch, with developers using it for local search and retrieval-augmented generation (RAG). The second generation moves beyond text while keeping the model small enough for edge hardware.
Its modular design lets developers use a 270 million-parameter text component independently, while optional vision and audio encoders add 170 million and 300 million parameters, respectively.
The broader significance lies in its memory efficiency
Google says a quantized EmbeddingGemma 2 build uses about 191MB of active RAM for text-only processing on a Pixel 11 Pro and about 567MB with the full multimodal model.
That matters because embedding models typically sit underneath search systems, meaning their memory requirements can affect whether sophisticated retrieval can run locally at all. EmbeddingGemma 2 also uses Matryoshka Representation Learning to reduce its 768-dimensional vectors to 512, 256 or 128 dimensions, cutting local storage requirements by up to six times, according to Google.
The model also has an 8,192-token context window, four times that of the previous generation. Google says that is enough for roughly 5.5 minutes of audio, 29 images or 58 video frames, including combinations of those formats.
Performance has improved too. EmbeddingGemma 2 scored 78.68 on the MTEB Code benchmark, up from 68.76 for the first model.
Google is turning retrieval into an edge feature
Google is already showing how that could work through Google AI Edge Gallery and its new AI Edge Foresight Mac app.
Video Moments Finder can index local video and audio, then surface timestamps matching searches such as “dog catching a frisbee” or “person blowing out birthday candles.” Foresight takes the idea into productivity, using local meeting audio and files to enrich notes and retrieve information without sending the underlying data to the cloud.
The developer stack is equally important. MediaPipe Tasks can handle multimodal preprocessing and retrieval, while LiteRT provides more control over deployment and acceleration across CPUs, GPUs and NPUs. Google says its vision embedding implementation can run in as little as 37.3 milliseconds on a MacBook M5 Pro GPU.
What this means for users
The biggest change is not that phones suddenly have a smarter search box. It is that personal media can become searchable without first being uploaded somewhere.
A user could search years of photos and videos using ordinary language, find a moment in a recording or retrieve information from private files while offline. That could reduce latency and make privacy-sensitive applications easier to build.
There are still limits. The model has had no safety tuning, and Google says performance is not equal across all of the more than 100 languages it supports.
More Google coverage
- New Google Search AI Mode is ‘Total Reimagining,’ Says CEO Sundar Pichai
- In Major Ruling, Judge Finds Google ‘Willfully Acquired and Maintained Monopoly Power’ Over Digital Ad Market
- Google’s Big Bet on Nuclear Energy: ‘The Race to Power AI-Driven Data Centers is Accelerating’
- Computer History Museum Releases Original AlexNet Code: Why It Matters
The bigger shift: retrieval moves closer to the data
EmbeddingGemma 2 matters less as a standalone AI model than as another sign that sophisticated retrieval is moving from cloud services onto everyday hardware.
For developers, that creates another architecture choice: send personal data to a remote model, or keep more of the search and indexing process on the device itself. As phones and laptops gain more capable AI hardware, Google is betting that local multimodal retrieval will increasingly become the default rather than the exception.
Other news: Google’s rumored Fitbit Edge could bring a larger display, altimeter, fast charging, and a user-replaceable battery, but Charge 6 owners may want to wait for confirmed pricing and features before upgrading.