Local AI

EmbeddingGemma 2 Is Capable of Running in 567MB of RAM

EmbeddingGemma 2 logo and name on a blue modern background.
Google DeepMind’s EmbeddingGemma 2 is built for lightweight multimodal search on phones and laptops.

Google DeepMind’s new EmbeddingGemma 2 is a small 740 million parameter model made for running directly on everyday devices like phones and laptops.

Instead of focusing only on text, it can work with text, code, images, video, and audio, which makes it useful for building local search across different types of content.

However, EmbeddingGemma 2’s memory requirement is the part worth focusing on. Google says the full multimodal model can run using around 567MB of active RAM on a Pixel 11 Pro, while the text-only version drops to roughly 191MB. That puts it in a range where developers could realistically build it into normal apps instead of relying on a cloud server for every search.

What it does (and doesn’t)

EmbeddingGemma 2 works more like a search layer than a chatbot. Instead of replying with text, it converts whatever you give it into a mathematical representation of what that content means.

Because text, code, images, video, and audio are all placed into the same 768-dimensional space, the model can compare completely different types of content directly. A normal text query could be used to find a matching photo, a voice recording, or even a moment inside a video.

That is basically its whole job, finding the right content. A separate language model can then take that content and generate the answer.

The size

EmbeddingGemma 2 is modular, so an app does not need to load the full 740 million parameter model every time. Developers can load only the parts they actually need. A notes app searching through documents can stick to text and code, while a photo gallery can add vision without keeping the audio part in memory.

Components loadedParameters
Text and code270M
Text, code and vision440M
Text, code and audio570M
All in one740M

Google measured both figures on a Pixel 11 Pro. For comparison, even a small quantized chat model usually needs several gigabytes.

At this size it can sit next to a normal app without hogging the phone’s memory. The download is bigger than that, though. The full model files come to about 1.53GB on Hugging Face, with the main weights file at roughly 1.49GB, so 567MB is the runtime figure, not the download size. Google also ships pre-quantized INT4 and INT8 builds through the Hugging Face LiteRT Community.

Smaller vectors

Once an app starts indexing photos, documents, or other content, each item also needs its own vector, and that can start eating into storage pretty quickly on a phone.

EmbeddingGemma 2 creates 768-dimensional vectors by default, but it can also use 512, 256, or 128. Google lists 256 as 3x smaller and 128 as 6x smaller than the full vector. Smaller vectors make the index lighter and can speed up comparisons, but going too small hurts search quality.

Google’s model card says quality stays close to the full vector down to 256 dimensions. 128 is best suited to text-only search. On image and video retrieval it drops noticeably, so for photo search on a phone, stay at 256 or above.

Using standard float32 values, where each number takes 4 bytes, a library with 100,000 indexed photos would need around 307MB at 768 dimensions. Dropping that to 256 dimensions cuts the index to roughly 102MB.

That is a pretty big saving for a phone app. Smaller vectors also mean less data to compare during each search, which can help with speed too. For mobile use, 256 dimensions looks like a sensible starting point, then developers can test whether the search quality is still good enough for their own data.

Context window

The context window is 8,192 tokens. For text, that fits large documents in one pass. Audio costs 25 tokens per second, so about 327 seconds fits in one input. Video is sampled at 1 frame per second by default, and about 58 frames fit, so roughly a minute of video per input. It won’t process a two-hour movie. You get searchable chunks, not a full replay.

What it’s for

EmbeddingGemma 2 is mainly useful when an app needs to find relevant content fast. That could mean pulling the right document from a large folder, finding a specific moment in a video, matching a text query to a photo, or searching through voice recordings without relying on exact keywords.

It can also sit behind a local RAG system, where it finds the most relevant files first and another model handles the actual answer. For agents, it can help classify a command and decide which tool or action should handle it.

So the model is less about generating content and more about helping apps quickly figure out what information, file, or action is the closest match.

Conclusion

EmbeddingGemma 2 is available under the Apache 2.0 license, so developers can modify it, redistribute it, and use it commercially under the license terms. Google has published the model on Hugging Face and Kaggle.

The main thing to keep in mind is that the RAM numbers come from Google’s own testing on a Pixel 11 Pro, and active RAM is not the same as total memory use. Your app, the vector index, and any language model running alongside it will all add extra memory on top. The same goes for the smaller vector sizes, so developers should still test performance and retrieval quality on their own data.

EmbeddingGemma 2 won’t write a word for you, and it doesn’t need to. Its job is finding things, and at this size it can do that inside a normal phone app without your files leaving the device.

Sources

EmbeddingGemma 2: The Developer Guide

Bring Multimodal Semantic Search to the Edge with EmbeddingGemma 2

EmbeddingGemma 2 on Hugging Face

EmbeddingGemma 2 model files

    Updated Oct 7, 2026