Pull to refresh
Logo
Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model

Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model

New Capabilities

740M-parameter model maps text, code, images, video, and audio into one vector space for on-device search

Today: Google DeepMind launches EmbeddingGemma 2

Overview

Updated 1 hour ago

Google DeepMind released EmbeddingGemma 2 on October 6, a 740-million-parameter model that maps text, code, images, video, and audio into a single 768-dimensional vector space. It runs on a phone or laptop, not a data center.

The model is open under the Apache 2.0 license, so developers can use it commercially. It can retrieve a specific video clip with a voice memo, or search hours of audio with a text query — all processed locally.

Why it matters

If on-device embeddings catch on, personal photos, voice notes, and documents become searchable without ever leaving the device.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

740M
Total model parameters
Combines a 270M text model with modular vision and audio encoders.
768
Embedding dimensions
All five modalities project into a single 768-dimensional vector space.
8K
Context window (tokens)
4x larger than EmbeddingGemma 1; processes up to 5.5 minutes of audio or 29 images.
9.92
MTEB Code score gain
Code benchmark score rose from 68.76 to 78.68 versus EmbeddingGemma 1.
~567MB
RAM for full multimodal model
On a Google Pixel 11 Pro with quantization; text-only needs ~191MB.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

People Involved

Organizations Involved

Timeline

1 event Latest: Today
  1. Google DeepMind launches EmbeddingGemma 2

    Today Product Launch

    Open 740M-parameter model maps text, code, images, video, and audio into one 768-dimensional space. Weights released on Hugging Face and Kaggle under Apache 2.0.

Scenarios

1

EmbeddingGemma 2 becomes the standard on-device embedder

Likely Resolves by Apr 6, 2027

Discussed by: Google's developer blog and the model's benchmark results

If the model's benchmarks hold in practice, developers building privacy-sensitive apps adopt it. The shared tokenizer with Gemma 4 lowers the barrier for on-device RAG pipelines, and the Apache 2.0 license removes licensing friction. The 9.92-point code benchmark gain makes it attractive for local codebase indexing.

2

Competitors ship comparable open multimodal embedders

Possible Resolves by Oct 6, 2027

Discussed by: Industry analysts tracking the open-model race

Meta, Mistral, or another lab releases a sub-1B open multimodal embedding model to match. The Apache 2.0 license and the code benchmark gain set a high bar for open alternatives, but the open-model field has moved fast since Llama's release.

3

Adoption stays limited to flagship hardware

Unlikely Resolves by Oct 6, 2027

Discussed by: Skeptics noting the RAM and hardware requirements

The full multimodal model needs ~567MB of RAM, which limits it to flagship phones. If developers find the quality-per-parameter tradeoff insufficient for their use cases, adoption stays niche and the model's reach is confined to high-end devices.

Historical Context

2 moments from history that rhyme with this story — and how they unfolded.

January 2021

CLIP (2021)

OpenAI released CLIP, a model trained on 400 million image-text pairs to map images and text into a shared embedding space. It enabled zero-shot image classification and text-to-image retrieval without task-specific training.

Then

CLIP became the backbone for a wave of multimodal systems, including image generation models like DALL-E 2.

Now

It established contrastive learning on paired data as the standard way to build multimodal embeddings.

Why this matters now

EmbeddingGemma 2 extends CLIP's core idea — a shared embedding space for different modalities — from two modalities to five, and shrinks it to run on a phone.

February 2023

Llama (2023)

Meta released Llama, an open-weight large language model, under a permissive license. It showed open models could approach the performance of closed models like GPT-3.

Then

Llama sparked a wave of local AI development, fine-tunes, and on-device deployments.

Now

It established open-weight models as a credible alternative to API-only models, a trend that continued with Gemma and others.

Why this matters now

EmbeddingGemma 2 follows the same playbook — open weights, permissive license, on-device focus — applied to embeddings rather than generation.

Sources

(7)