Skip to Content
Explainer

One Map

EmbeddingGemma 2 puts text, pictures, sound and video into one vector space, and it runs without a graphics card. How it works, and what we measured on four processor cores.

Share this case file← All articles

Monday. A pipe bursts, a bedroom floods, and by evening the job has left its trail on the shared drive: a report, a voice memo, photographs, a scanned delivery note, a video walking through the flat.

Three weeks later somebody asks one thing: show me everything about the bedroom ceiling. Search finds the report, because the report has the words. The photograph of the ceiling has no words in it at all.

A black table with a hard drive, an inspection report stamped FOUND, a tape recorder, a projector, a scanned delivery note and a photograph of a stained ceiling stamped NOT FOUND, with a tag: no words in it

On 6 October 2026 Google DeepMind released a model that is meant to close that gap: EmbeddingGemma 2. This post explains what it is and how it works, and then reports what happened when we ran it on four processor cores.

What an embedding is

An embedding is a list of numbers that stands for what something means. The easiest way to think about it is as a place on a map. A model reads a sentence and puts it somewhere; sentences that mean the same land close together, and it does not matter whether they share a word. “A burst pipe” and the German Rohrbruch are neighbours.

Searching is then simple. Put the question on the same map and look at what lies nearby.

A sheet of squared paper as a map: "a burst pipe", "Rohrbruch", "water through the ceiling" and "the joint under the sink failed" close together inside a lime circle round the question; "holiday rota" and "van service record" far away

How pictures and sound were searched before

For years that map was for text only, so everything else was turned into text first. A picture went through a reader that looks for letters (OCR). A recording went through a transcriber. Then the words were embedded.

It works, and it is what most document search does today. It also has two costs. It is three machines in a row, each with its own errors. And you can only find what one of them wrote down. A stain on a ceiling writes nothing down.

Three machines in a row: a photograph goes to a magnifying glass labelled "a reader looks for letters" and comes out as an empty sheet stamped NOTHING; a tape recorder goes to a typewriter and comes out as a transcript; only the transcript reaches the map

There were steps in between. CLIP (2021) trained two models side by side, one for pictures and one for captions, so that a photograph and its caption land together. That made text-to-image search good, for one pair of formats. ImageBind (2023) aligned six formats to images. The newest approach is one model that reads every format itself. Google’s Gemini Embedding 2 works that way, but only as an API. EmbeddingGemma 2 is the same idea with open weights, in a size you can run yourself.

EmbeddingGemma 2

Released6 October 2026, Google DeepMind
LicenceApache 2.0, open weights
Size740 million parameters: 270M for text, 170M for the vision encoder, 300M for the audio encoder
ReadsText and code, images, video, audio, and mixtures of them
ReturnsOne vector of 768 numbers, whatever came in
Window8,192 tokens: about 29 images, 58 video frames or five and a half minutes of audio

A model that reads text cannot read pixels. So two small encoders stand in front of it. One turns a picture into tokens, the other does the same for sound. From there on everything is tokens in one line, through one shared backbone. By default a photograph costs 280 tokens, a video frame 140, and a second of audio 25.

At the end the model averages the lot into 768 numbers: one place on the map.

A diagram on a table: a photograph goes into a box "vision encoder 170M", a tape recorder into "audio encoder 300M", a strip of text goes straight on; all three arrive in a box "one shared backbone 270M" as rows of tokens, and a strip of 768 numbers comes out and lands on a small map

We found no technical report for this model itself. Its larger sibling behind the API, Gemini Embedding 2, has a paper, and that describes large-scale contrastive training: show the model pairs that belong together, a photograph and its caption, a clip and its transcript, and train it to pull each pair together and push everything else away.

Two design choices that matter once you run it

Smaller vectors. The model is trained so that the first 512, 256 or 128 numbers of a vector work as a vector on their own (Matryoshka representation learning, after the Russian dolls). At 256 numbers you store a third. The catch is in Google’s own table:

Numbers keptText (MTEB multilingual)Pictures, video, documents (MMEB v2)
76861.3659.01
25660.4156.24
12857.8945.65

Cut down to 128, text loses three and a half points; pictures and video lose thirteen.

Six nesting dolls above a strip of numbers whose first third is inked lime, and a sheet with Google's table: at 128 numbers text scores 57.9 and pictures and video 45.7, both marked

Optional encoders. The vision and audio encoders can be left out. Text alone is 270 million parameters, and it lands on the same map. So you can index your photographs once with the full model and search them with the small one.

Three blocks: vision, optional, plus 170M; audio, optional, plus 300M; text alone, 270M, in lime. Threads lead to one map: "index once, full model" and "search, small model"

Running it on a processor

Google says the model is built for laptops and phones. We wanted to know what that means on a plain server processor, with no graphics card. The set-up: a single-node Kubernetes cluster in Docker on an Apple M1 Pro, the model pod limited to four cores, the 8-bit quantised weights.

Where it runs today

The first surprise had nothing to do with speed. Ollama is the usual way to run such models, and its library has embeddinggemma-2. In a Linux container every tag answered the same:

Error: this model requires MLX support, but the MLX runtime is not available

On 11 October 2026 (Ollama 0.40.3) the library’s builds are for Apple’s MLX runtime only. A maintainer wrote on the issue that Linux and Windows support will be added soon. Pulling the GGUF file through Ollama did not help either: its bundled llama.cpp did not know the architecture yet.

llama.cpp’s own server did. One container, the official GGUF from ggml-org/embeddinggemma-2-GGUF, and an OpenAI-style endpoint:

containers: - name: llama-server image: ghcr.io/ggml-org/llama.cpp:server args: - --hf-repo - ggml-org/embeddinggemma-2-GGUF:Q8_0 - --no-mmproj # text tower only; leave it out to load the encoders - --embedding - --ctx-size - "8192" - --parallel - "4" - --threads - "4"

Two tabs on a table. "Ollama 0.40.3, Linux" with the error text and a stamp: Apple chips only. "llama.cpp, official GGUF" with "model loaded, POST /v1/embeddings, 768 numbers" and a stamp: works

How fast

EmbeddingGemma 2, 4 coresEmbeddingGemma 2, 8 coresMiniLM-L6, 2 cores
One question0.12 s0.11 s0.002 s
Passages of 1,200 characters, a second0.40.6 to 0.75.9

MiniLM-L6 is the small English model we had used until then (22 million parameters). A passage of 1,200 characters is about 300 tokens.

So a search is fast enough that nobody notices. Building the index is the slow part: at 0.4 passages a second, a hundred thousand passages take about three days on four cores. For a drive of a few thousand files that is an evening. For an archive it is a reason to borrow a graphics card for the indexing and keep the CPU for searching.

A stopwatch, the figure 0.12 s for one question, and two bars for indexing in passages a second: EmbeddingGemma 2 on four cores 0.4, MiniLM on two cores 5.9

Does it find things

We wrote ten questions about the drive, each with the files a person would want back, and asked them through the same search with three models.

ModelRight file firstMean reciprocal rank
MiniLM-L6 (384 numbers, trained on English)5 of 100.64
EmbeddingGemma 2 on four cores6 of 100.78
Gemini Embedding 2 over the API (768 numbers)7 of 100.83

Ten questions are a demonstration, not a benchmark. The clearest difference was not in the totals. “Flooded basement after heavy rain”, asked in English, returned the German note Kellerüberflutung first with both Google models. MiniLM did not return it at all.

A scoreboard of three rows of ten squares: MiniLM 5 of 10, EmbeddingGemma 2 6 of 10, Gemini Embedding 2 7 of 10; under it the English question "flooded basement after heavy rain" and the German note it found, stamped FIRST

The prompts that are easy to miss

One line in the model card decides whether you get those numbers. The model wants to be told which side of a search a text is on:

task: search result | query: water coming through the bedroom ceiling title: none | text: Water has come through the ceiling over the bed…

A library such as sentence-transformers adds these for you. A bare /v1/embeddings endpoint embeds whatever string arrives. Without the prompts the same model, on the same chunks, put the right file first five times instead of six (mean reciprocal rank 0.70 instead of 0.78): short strings that merely looked like the question outranked the passage that answered it.

Two paper strips: "task: search result | query:" in front of a question, and "title: none | text:" in front of a document; under them two rows of squares, with prompts six of ten, without five

The photograph itself

Then we loaded the full model, encoders included, on the same four cores, and gave it the pictures themselves. No OCR, no caption.

First a sanity check: six captions against six pictures. Every caption scored highest on its own picture.

A six by six grid of similarity scores between six captions and six pictures; the diagonal, each caption's own picture, is lime, from 0.67 to 0.81

Then one index of 43 things: the 35 files that have text, the six pictures as pixels, and the two voice memos as sound.

  • “Water coming through the bedroom ceiling”: the photograph came third, the spoken memo fourth.
  • “The pipe joint that failed under a sink”: the photograph of the pipe came first.
  • Schimmel im Schlafzimmer hinter dem Schrank, asked in German: the photograph of the mouldy corner came fourth.
  • Each voice memo, embedded as sound, was closest to its own transcript (0.83 against 0.72 for the other memo’s).

The ceiling photograph next to a list headed "43 things on one map, nearest first": the voice memo's transcript, the moisture readings, then "the photograph, pixels" marked in lime at 0.763, then "voice memo, sound"

The price is time. A photograph took about a minute on four cores (53 to 60 seconds each), the A4 scan nearly four minutes, a 33-second memo 12 seconds. That is fine for a few hundred pictures and hopeless for a million.

Where we ran it

We ran the text part of this inside Classifyre, which scans a source, reads what it can into text, and indexes the text. The model plugs in as an OpenAI-compatible endpoint: an address inside the cluster, no key, and the two prompts.

The Classifyre provider page: embedding model switched on, 768 dimensions, a document prompt "title: none | text:" and a query prompt "task: search result | query:", and an empty API key field with the note that a server like Ollama or llama.cpp takes none

Asked about the bedroom ceiling, the workspace returns the voice memo first (Whisper had transcribed it), then the moisture readings, then the report.

The Classifyre assets list searched in semantic mode for "water coming through the bedroom ceiling": the voice memo HR-2231-site-memo.m4a first, then HR-2231-readings.csv, then HR-2231-inspection-report.docx

What it does not do: it uses the model’s text side only. Pictures are read for their words, and the three photographs with no writing in them are not in any result. Embedding the picture itself, as in the section above, is not in the product.

Testing a new model also turned up things that had nothing to do with it. Four of them are fixed on the development branch as of today:

  • The reader had stopped reading. A release of the OCR library on 8 October dropped a name another library imports. That library caught the error, logged “no OCR engine found” and carried on, so every image and scan came back as an empty document. It now counts as a failed extraction with the reason named, and the version is pinned.
  • The transcriber crashed on every audio and video file, after an update of the media library removed an argument the speech library still passes. Pinned as well.
  • An endpoint without an API key was refused, which ruled out every self-hosted server. A key is now optional for those.
  • There was nowhere to put the two prompts. They are now fields on the provider, and changing the document prompt re-embeds the workspace.

What it will not do

  • Its scores sit close together. In our grid the right picture scored 0.67 to 0.81 and a wrong one 0.51 to 0.71. Rank the results; do not draw a line at a number.
  • 128 numbers are for text. Google’s table says so, and so does the model card.
  • It is not fast on a processor when indexing, and pictures cost far more than text.
  • It reads 8,192 tokens at a time. A long recording or video has to be cut up first.
  • It was five days old when we tested it. The tools are still catching up, as the Ollama error shows.

A rule of thumb

One map beats three machines in a row. Index once with the big model, search with the small one. And test it on your own files, with questions whose answers you know.

Sources

More from the blog

Last updated on