We’re releasing EmbeddingGemma 2, our first natively multimodal open model engineered for on-device embeddings. Built on the Gemma 4 architecture and released under an Apache 2.0 license, it goes bey
The Hook
“We're releasing Embedding Gemma 2, our first natively multimodal open model engineered for on-device embeddings. Previously, if you wanted to embed different data types, you'd have to use separate models, one for text, one for images, one for video, and audio, or any combination of them. Embedding Gemma 2 unifies all of these into a single embedding space. This means that a picture of a cat, the word cat, and the sound of a cat are all mapped close to each other in the embedding space. So with this one model, you can build search, retrieval, classification, and clustering applications across all of these modalities. And because Embedding Gemma 2 is built on the Gemma 4 architecture, it's incredibly efficient, even outperforming much larger models.”
Feature Teaser
“Embedding Gemma 2 comes in three sizes: a 270 million parameter text-only model, a 440 million parameter text plus vision model, and a 570 million parameter text plus audio model. And because it's so efficient, you can truncate outputs down to as low as 128 dimensions. This means you can embed long documents or extended voice recordings right on device, which is perfect for edge hardware. For example, you can use Embedding Gemma 2 to power instant media search with the Google AI Edge Gallery app. I can also search within videos using Video Moment Finder.”
Product Reveal
“And because all of this happens on device, no external API calls are required. This means you can use Embedding Gemma 2 to power even more complex pipelines, like retrieval augmented generation with Gemma 4, where sensitive user data never leaves the hardware. For example, let's say I'm using a personal assistant app like Foresight and I ask, "Hey, how does Foresight actually leverage Embedding Gemma 2?" Foresight can then use Embedding Gemma 2 to retrieve relevant personal context chunks from a local knowledge base, which then feeds into a multimodal answer generation. And because all of this happens on device, sensitive information stays strictly private.”
Call to Action
“And because Embedding Gemma 2 is open, you can fine-tune it with your own data and specialized assets, whether that's to improve retrieval precision without having to retrain the entire model, or to adapt it to a specific domain. Embedding Gemma 2 is available today on Kaggle, Hugging Face, and through the Google AI Edge Gallery app. We're excited to see what you build with it.”
Synthesizable in Remotion · key components of the gemma-multimodal-launch template.
A dynamic 3D isometric grid serves as a background, with data points, icons, and text elements animating along its lines, suggesting data movement and connectivity. Elements often slide or scale into position with a slight bounce.
Utilize Remotion's `AbsoluteFill` for the grid background. Implement `interpolate` for position and scale animations, using `spring` for a subtle bounce effect. SVG paths can define the grid lines, animated with `stroke-dasharray` and `stroke-dashoffset`. Use `perspective` and `rotateX`/`rotateY` for the isometric view.
Abstract representations of different data types (text, image, audio, video) are depicted as glowing, interconnected nodes that flow into a central processing unit (EmbeddingGemma 2) and then map into a unified 3D embedding space.
Create SVG icons for each data type. Animate their `opacity` and `transform` (translate, scale) using `spring` for entry. Use `Path` components to draw connecting lines that animate with `stroke-dasharray` and `stroke-dashoffset`. The 3D embedding space can be a `Canvas` or `WebGL` component rendering simple spheres or cubes with `interpolate` for position.
Line graphs and bar charts dynamically draw themselves, highlighting specific data points or bars to illustrate performance metrics and model sizes. Presenter's face often appears in a small, animated inset.
For line graphs, use `Path` components with `stroke-dasharray` and `stroke-dashoffset` to animate the line drawing. Data points can be animated circles with `spring` for their appearance. Bar charts can use `Rect` components with `height` animated via `interpolate` and `spring`. The presenter inset can be a `Video` component with a `clipPath` and `scale` animation for entry.
Purpose. Introduce the core problem of disparate models and present EmbeddingGemma 2 as the elegant, multimodal solution.
Execution. Starts with presenters, quickly transitions to abstract icons representing different data types, then shows them converging into a single 'EmbeddingGemma 2' node, finally mapping into a unified 3D embedding space.
Purpose. Detail the model's technical advantages: efficiency, smaller sizes, and on-device capabilities.
Execution. Features dynamic line graphs comparing performance, animated bar charts showing model sizes, and a simulated phone screen demonstrating on-device search, emphasizing speed and local processing.
Purpose. Showcase practical use cases, particularly for complex AI pipelines and highlight the privacy benefits of on-device processing.
Execution. Animated data flow diagrams illustrating Retrieval Augmented Generation (RAG) with Gemma 4, followed by a simulated document view explaining a real-world application (Foresight) and its privacy implications.
Purpose. Encourage developers to use and fine-tune the model, providing clear access points.
Execution. Presenter explains fine-tuning, accompanied by abstract animations of data flowing into and out of the 'EmbeddingGemma 2' node, concluding with a clear display of the Gemma logo and platform availability.
Objective
Launch 'SynapseAI', a new on-device federated learning platform for smart home devices.How the recipe adapts
The video would open with a founder explaining the privacy challenges of current smart home AI. This transitions to an IsometricGridFlow showing data from various home devices (represented by MultimodalDataNodes like a thermostat icon, a speaker icon, a camera icon) flowing into a central 'SynapseAI' node, but crucially, the data remains within a local network boundary. DynamicMetricGraphs would illustrate the efficiency and low latency of on-device processing compared to cloud solutions. The narrative would then demonstrate how SynapseAI powers personalized routines without sending sensitive data externally, using a simulated smart home dashboard. The video would conclude with a call to action for developers to integrate SynapseAI, featuring the SynapseAI logo with a subtle, interconnected network animation.Scene durations (4 scenes, 189.5s)
Key Remotion components
To ensure that artificial general intelligence benefits all of humanity

Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available. https://t.co/2AvmEjHIX8
To understand the true nature of the universe.