Google DeepMind Launches EmbeddingGemma 2 for On-Device Multimodal Search

Google DeepMind’s EmbeddingGemma 2 brings text, image, video and audio search to consumer devices with lower memory and less reliance on cloud AI.

Escrito por
Aminu Abdullahi
Aminu Abdullahi
Oct 7, 2026
Google DeepMind Launches EmbeddingGemma 2 for On-Device Multimodal Search

Google debuts on-device EmbeddingGemma 2 model. Image: Google

We may earn from vendors via affiliate links or sponsorships. This might affect product placement on our site, but not the content of our reviews. See our Terms of Use for details.

Google DeepMind wants your phone to understand your videos, voice notes, and photo library without sending everything to the cloud.

The company launched EmbeddingGemma 2 on Tuesday, a 740 million-parameter open-weight model that maps text, code, images, video, and audio into a shared embedding space. The model is designed to let developers build multimodal search and retrieval features that run directly on consumer devices.

Google said the first EmbeddingGemma has passed 20 million downloads since its 2025 launch, with developers using it for local search and retrieval-augmented generation (RAG). The second generation moves beyond text while keeping the model small enough for edge hardware.

Its modular design lets developers use a 270 million-parameter text component independently, while optional vision and audio encoders add 170 million and 300 million parameters, respectively.

The broader significance lies in its memory efficiency

Google says a quantized EmbeddingGemma 2 build uses about 191MB of active RAM for text-only processing on a Pixel 11 Pro and about 567MB with the full multimodal model.

That matters because embedding models typically sit underneath search systems, meaning their memory requirements can affect whether sophisticated retrieval can run locally at all. EmbeddingGemma 2 also uses Matryoshka Representation Learning to reduce its 768-dimensional vectors to 512, 256 or 128 dimensions, cutting local storage requirements by up to six times, according to Google.

The model also has an 8,192-token context window, four times that of the previous generation. Google says that is enough for roughly 5.5 minutes of audio, 29 images or 58 video frames, including combinations of those formats.

Performance has improved too. EmbeddingGemma 2 scored 78.68 on the MTEB Code benchmark, up from 68.76 for the first model.

Advertisement

Google is turning retrieval into an edge feature

Google is already showing how that could work through Google AI Edge Gallery and its new AI Edge Foresight Mac app.

Video Moments Finder can index local video and audio, then surface timestamps matching searches such as “dog catching a frisbee” or “person blowing out birthday candles.” Foresight takes the idea into productivity, using local meeting audio and files to enrich notes and retrieve information without sending the underlying data to the cloud.

The developer stack is equally important. MediaPipe Tasks can handle multimodal preprocessing and retrieval, while LiteRT provides more control over deployment and acceleration across CPUs, GPUs and NPUs. Google says its vision embedding implementation can run in as little as 37.3 milliseconds on a MacBook M5 Pro GPU.

What this means for users

The biggest change is not that phones suddenly have a smarter search box. It is that personal media can become searchable without first being uploaded somewhere.

A user could search years of photos and videos using ordinary language, find a moment in a recording or retrieve information from private files while offline. That could reduce latency and make privacy-sensitive applications easier to build.

There are still limits. The model has had no safety tuning, and Google says performance is not equal across all of the more than 100 languages it supports.

More Google coverage

Advertisement

The bigger shift: retrieval moves closer to the data

EmbeddingGemma 2 matters less as a standalone AI model than as another sign that sophisticated retrieval is moving from cloud services onto everyday hardware.

For developers, that creates another architecture choice: send personal data to a remote model, or keep more of the search and indexing process on the device itself. As phones and laptops gain more capable AI hardware, Google is betting that local multimodal retrieval will increasingly become the default rather than the exception.

Other news: Google’s rumored Fitbit Edge could bring a larger display, altimeter, fast charging, and a user-replaceable battery, but Charge 6 owners may want to wait for confirmed pricing and features before upgrading.

Aminu Abdullahi

Aminu Abdullahi is a B2C and B2B technology and finance writer with more than six years of experience covering enterprise IT, cybersecurity, cloud computing, artificial intelligence, fintech, business software, and emerging technologies. He has written for a wide range of technical and business audiences, from IT professionals and cybersecurity leaders to small business owners, executives, and technology buyers. His work has appeared in publications including: TechRepublic eWEEK Channel Insider Geekflare Enterprise Networking Planet eSecurity Planet CIO Insight Webopedia With a background in computer science, Aminu specializes in translating complex technical subjects into clear, practical, and accessible content. His writing helps readers understand emerging technologies, evaluate business software, strengthen cybersecurity strategies, and make more informed decisions about technology investments. Across his work, Aminu focuses on the real-world impact of technology, connecting technical innovation with business value, operational efficiency, security, and long-term digital transformation.