EmbeddingGemma 2 adds multimodal search to Google creative ecosystem
Google DeepMind launched EmbeddingGemma 2 to improve how AI systems understand and organize visual data. The model allows filmmakers to build more efficient search tools for managing large media archives.
Google DeepMind released EmbeddingGemma 2, an open-weights multimodal embedding model designed to improve how AI systems categorize and retrieve visual and textual data. The release provides a more efficient way for developers to build search tools for creative assets, directly impacting the retrieval-augmented generation (RAG) pipelines used in media production. By mapping images and text into a shared vector space, the model allows for more nuanced understanding of visual content in tools like Google Veo 3.
What's new
EmbeddingGemma 2 is available in two sizes: a 9-billion parameter model and a 27-billion parameter version. Unlike previous text-only models, this version is natively multimodal, meaning it can map images and text into the same mathematical space. This allows for "find a video that looks like this" or "find clips matching this description" queries without needing manual tagging or metadata entry.
The model outperforms larger competitors on the MTEB (Massive Text Embedding Benchmark) while maintaining a smaller footprint. As of February 2025, the model is available for download and integration via Hugging Face and Google’s Vertex AI platform. It supports a context window that handles complex descriptions, making it suitable for indexing long-form video metadata or detailed scene descriptions.
How it fits your workflow
For editors and VFX houses managing massive amounts of footage, EmbeddingGemma 2 serves as the engine for semantic search. Instead of relying on filename-based searches or manual logs, users can query their local databases using natural language. This functions similarly to how the search feature in Google Photos works, but with the flexibility of an open-weights model that can be hosted on private servers to protect intellectual property.
In the context of AI video generation, EmbeddingGemma 2 acts as a bridge between a user's prompt and the visual training data. It competes directly with OpenAI’s CLIP and Cohere’s Embed models. By using EmbeddingGemma 2, developers can create more precise style reference features where a model like Google Veo 3 or Runway Gen-3 Alpha can better understand the nuances of a reference image provided by the user. This leads to better consistency when generating video from a specific art style or character design.
What it costs / how to try it
EmbeddingGemma 2 is an open-weights model, meaning it is free to download for research and commercial use under the Gemma Terms of Use. Developers can access the weights on Hugging Face or deploy it through Google Cloud Vertex AI to begin building custom media search tools.
Read the original announcement on Google Veo 3 ↗