← Back to blog
AiAbout 6 min read

EmbeddingGemma 2 Puts Photo, Audio and Video Search on Your Phone

Published Oct 11, 2026
EmbeddingGemma 2 Puts Photo, Audio and Video Search on Your Phone

On October 6, Google DeepMind released EmbeddingGemma 2, a 740-million-parameter model that takes text, code, images, audio and video and places them all in one shared mathematical space. The number that matters is not the size. It is that a model this small can run on a phone, which means search over your own media can happen without anything leaving the device.

Embeddings are the quiet machinery behind a lot of things people use every day. An embedding model turns a piece of content into a list of numbers, arranged so that similar things end up near each other. That is how a search engine knows that "how do I fix a leaky faucet" and "repair a dripping tap" are about the same subject. It is also how recommendation systems decide what to show you next. Until now, doing this well across images, audio and video usually meant sending data to a server and paying for the round trip.

One space for everything

What makes EmbeddingGemma 2 interesting is the word "shared." Many models handle one modality well: text here, images there. A shared space means a single model can compare a sentence to a photo, a photo to a snippet of audio, or a video frame to a written description. The practical example Google offers is finding a specific video clip by searching with a voice memo. You describe what you remember, in audio, and the system finds the matching footage.

That kind of cross-modal search has existed in cloud services for a while. What is new is the size. At 740 million parameters, the model is small enough for a phone, which flips the privacy equation. If search runs locally, your photos, audio and video never leave the device. For anyone who has hesitated to upload personal media to a cloud service just to make it searchable, that changes the calculation.

Four media tiles (text, photo, audio wave, video) converging into a single glowing point

Why on-device keeps winning small battles

There is a pattern here that goes beyond one model. Over the past year, a steady stream of small, capable models has moved tasks that used to require a data center onto devices people already own. Google's own Gemma line has pushed in this direction, and other labs have followed with small models tuned for specific jobs. The appeal is not only privacy. It is latency, offline availability, and cost. A search that runs on the phone answers instantly and works on a plane.

That shift also changes who gets to build. When a capability lives in the cloud, a developer needs an account, a budget and a network to use it. When it runs on the device, a developer needs a model file and some cleverness. Certain kinds of products become possible that were awkward before: a personal photo organizer that never phones home, a field tool that works without a signal, a note-taking app that indexes your audio and images locally. None of these are flashy, and each one depends on the same quiet advance, which is that a model small enough for a phone is now good enough to be trusted for real work.

For developers, this opens a specific door: retrieval-augmented generation that never leaves the machine. RAG, the technique of having a model answer questions using your own documents, normally depends on an embedding model to find the relevant passages first. If that embedding step can run locally, then a local assistant can answer questions about your files and media without a network call. For regulated industries, or for anyone handling sensitive material, that is the difference between a tool that is allowed and one that is not.

The trade-off nobody should ignore

A small model is not a big model, and the gap shows up in accuracy on hard cases. A hosted embedding model with far more parameters will generally retrieve better results, especially on messy, ambiguous or unusual content. If your application depends on getting the top result right nearly every time, a 740-million-parameter model running on a phone may not be enough, and you should test it against your actual data before committing.

The second limit is what the model does not do. EmbeddingGemma 2 is a search and organization tool, not a generator. It cannot make an image or a video. It can help you find the ones you already have. That distinction matters when the industry's attention is fixed on generation, because the unglamorous search-and-retrieve layer is often what makes a generative feature usable in the first place.

There are practical constraints too. Indexing a large personal library takes storage and energy, and phones are not infinite on either. A 740M model is small for an AI model and large for a phone app, so developers will need to be thoughtful about when it runs and how much it indexes. None of these are deal-breakers. They are the details that decide whether a feature feels instant or feels like a battery drain.

The multilinguality is worth a note as well. A shared space that covers many languages and many media types is more valuable the more varied your library is. Someone with photos from three countries, voice notes in two languages and documents in a third benefits far more from cross-modal search than someone with a tidy, single-language collection. For a global product, that is the whole point. For an individual, it is the difference between a feature that works on your best-organized folder and one that works on the messy reality of how people actually save things.

What this says about the road ahead

The most useful way to read this release is as a marker of where multimodal AI is heading. Generation gets the headlines, but the models that will shape daily experience are often the smaller ones that organize, retrieve and connect. A model that lets you search your own photos, audio and video by describing what you remember, without uploading anything, is a modest-sounding capability with a wide reach.

It also fills a gap that generation alone cannot. A generative feature is only half a product; the other half is finding the thing you want to change. Anyone who has scrolled through thousands of photos looking for one has felt that gap. A shared embedding space turns an impossible search into a sentence: describe it, and the matching items surface. That is a small convenience on paper and a large one in practice, because it is the difference between owning a library and being able to use it.

If the pattern holds, the next year will bring more tasks back onto the device. Each one moves a little more of what currently requires a server into something that works offline, in your hand, on your terms. EmbeddingGemma 2 is one step in that direction, and a quiet one. The quiet ones are often the ones that stick.

Related articles