Launch Directory/AI Agents/Google launches EmbeddingGemma 2 open multimodal embedding model
Google launches EmbeddingGemma 2 open multimodal embedding model

Google launches EmbeddingGemma 2 open multimodal embedding model · Launch Video Breakdown: Hook, Pacing & Motion Design

We’re releasing EmbeddingGemma 2, our first natively multimodal open model engineered for on-device embeddings. Built on the Gemma 4 architecture and released under an Apache 2.0 license, it goes bey

AI AgentsSeries A / GrowthOctober 10, 2026@Google
0:00 · The Hook · Introducing EmbeddingGemma 2: Multimodal Embeddings
0:00 / 0:00

Scene-by-scene timeline & spoken transcript

  1. The Hook

    Introducing EmbeddingGemma 2: Multimodal Embeddings

    “We're releasing Embedding Gemma 2, our first natively multimodal open model engineered for on-device embeddings. Previously, if you wanted to embed different data types, you'd have to use separate models, one for text, one for images, one for video, and audio, or any combination of them. Embedding Gemma 2 unifies all of these into a single embedding space. This means that a picture of a cat, the word cat, and the sound of a cat are all mapped close to each other in the embedding space. So with this one model, you can build search, retrieval, classification, and clustering applications across all of these modalities. And because Embedding Gemma 2 is built on the Gemma 4 architecture, it's incredibly efficient, even outperforming much larger models.”

    On screen
    Embedding Gemma 2 a lightweight, open model Sebastian Russo Product Manager Ivan Llanos Senior Product Manager Sahil Dua Research Lead separate models, one for text, video, and audio, or any combination of them Embedding Gemma 2 cat and the sound of a cat are all mapped close to each other Search Retrieval So with this one model, you can build search, retrieval, MIEB(lite) Score by Model Size EmbeddingGemma 2 LCO-Embedding-Omni-3B jina-embeddings-v5-omni-small Mean (TaskType) BirdirLM Omni-2.5B-Embedding ebind-full siglip-so400m-patch14-384 siglip-base-patch16-512 jina-embeddings-v5-omni-nano Model Size (Parameters, Billions) VL M2Vec-LoRA even outperforming much larger models
    Camera
    Alternates between a medium shot of three presenters seated at a table and dynamic, illustrative motion graphics. The camera is static during presenter shots, then transitions to a top-down isometric perspective for data visualization.
    Motion
    Seamless transitions between live-action and 3D isometric data visualizations. Text and icons animate in with subtle scaling and fading. Data points on graphs highlight and connect dynamically. Grid lines provide a sense of depth and structure.
  2. Feature Teaser

    Efficiency and On-Device Capabilities

    “Embedding Gemma 2 comes in three sizes: a 270 million parameter text-only model, a 440 million parameter text plus vision model, and a 570 million parameter text plus audio model. And because it's so efficient, you can truncate outputs down to as low as 128 dimensions. This means you can embed long documents or extended voice recordings right on device, which is perfect for edge hardware. For example, you can use Embedding Gemma 2 to power instant media search with the Google AI Edge Gallery app. I can also search within videos using Video Moment Finder.”

    On screen
    270M Text 440M Text + Vision 570M Text + Audio of 740 million parameters. you can truncate outputs down to as low as 128 dimensions. PDF PDF PDF or extended voice recordings right on device with the Google AI Edge Gallery app. I can also search within videos using Video Moment Finder. Blowing out candles
    Camera
    Mix of static medium shots of presenters and animated isometric views. The camera in the animated scenes maintains a consistent, slightly elevated isometric perspective, allowing for clear data flow visualization.
    Motion
    Animated bar charts and data flow diagrams. Text boxes and icons slide in and out. A simulated phone screen demonstrates on-device search, with a video playing and search results appearing dynamically. Elements animate along a grid, maintaining a clean, technical aesthetic.
  3. Product Reveal

    Privacy-Preserving AI Pipelines

    “And because all of this happens on device, no external API calls are required. This means you can use Embedding Gemma 2 to power even more complex pipelines, like retrieval augmented generation with Gemma 4, where sensitive user data never leaves the hardware. For example, let's say I'm using a personal assistant app like Foresight and I ask, "Hey, how does Foresight actually leverage Embedding Gemma 2?" Foresight can then use Embedding Gemma 2 to retrieve relevant personal context chunks from a local knowledge base, which then feeds into a multimodal answer generation. And because all of this happens on device, sensitive information stays strictly private.”

    On screen
    and no external API calls are required. even more complex pipelines. Retrieval Query Local Knowledge Base EmbeddingGemma 2 Gemma 4 Output Generation where sensitive user data never leaves the hardware. "Hey, how does Foresight actually leverage EmbeddingGemma 2?" Foresight leverages EmbeddingGemma 2 as part of its on-device architecture to perform Cross-Modal Vector Embeddings 1. This model is used in the process flow where an "Extracted Audio Query" is processed by the EmbeddingGemma v2 to retrieve relevant personal context chunks (text & images) from a knowledge base, which then feeds into the multimodal answer generation 2. Foresight EG2 Architect... Sources: Foresight: Gemma features and pipelines Foresight EG2 Architecture.png alongside the system architecture diagram sensitive information stays strictly private.
    Camera
    Continues alternating between presenter shots and isometric motion graphics. The animated scenes feature a consistent, slightly elevated isometric perspective, focusing on data flow and architectural diagrams.
    Motion
    Complex data flow diagrams animate, showing inputs, processing steps (EmbeddingGemma 2, Gemma 4), and outputs. Text blocks and diagrams slide into view, highlighting key information. A simulated document view appears, showcasing detailed architectural explanations and source links, emphasizing transparency and technical depth.
  4. Call to Action

    Fine-Tuning and Availability

    “And because Embedding Gemma 2 is open, you can fine-tune it with your own data and specialized assets, whether that's to improve retrieval precision without having to retrain the entire model, or to adapt it to a specific domain. Embedding Gemma 2 is available today on Kaggle, Hugging Face, and through the Google AI Edge Gallery app. We're excited to see what you build with it.”

    On screen
    and specialized assets, whether that's EmbeddingGemma 2 This helps improve retrieval precision without Gemma
    Camera
    Mix of presenter shots and animated graphics. The final shot is a static, centered view of the Gemma logo.
    Motion
    Animated text and bounding boxes highlight the model name. Data points and lines animate to illustrate fine-tuning concepts. The video concludes with a clean, full-screen display of the Gemma logo, featuring a subtle animated starburst icon.

Related AI Agents Product Launches

Explore all AI Agents launches →
OpenAI
Hook 9.2137.9M
OpenAIAI Agents

To ensure that artificial general intelligence benefits all of humanity

@OpenAI
Anthropic
Hook 9.257.6M
AnthropicAI Agents

Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available. https://t.co/2AvmEjHIX8

@claudeai
xAI
Hook 9.256.9M
xAIAI Agents

To understand the true nature of the universe.

@bot