Skip to content

UPDATED 18:18 EDT / OCTOBER 06 2026

AI

Google expands EmbeddingGemma beyond text to images, audio and video

Google LLC today released EmbeddingGemma 2, an open multimodal embedding model small enough to run on a smartphone.

The release takes the EmbeddingGemma line beyond text, which was all the first version handled when Google introduced it in September 2025. Images, audio and video now share one embedding space with text. An app built on the new model could take a voice memo, for example, and find the matching moment in a video without the data ever leaving the phone.

The response to the original “blew past our expectations,” Google DeepMind research engineers Sahil Dua and Henrique Schechter Vera wrote in the announcement. By their count, developers have downloaded it more than 20 million times.

Built on the Gemma 4 architecture Google released in April, the new version is more than twice the size of the original at 740 million parameters. Most of the growth is in the vision and audio encoders, which apps that only work with text can leave off. On its own, the 270 million-parameter text core used about 191 megabytes of memory when Google tested a quantized build on a Google Pixel 11 Pro.

An app also has to store what the model produces. Each embedding is a list of 768 numbers, and every photo, clip or document it indexes adds one more to a local vector database. A training technique called Matryoshka Representation Learning lets developers cut those lists to as few as 128 numbers, which reduces the space they take up by as much as six times. At 256, Google’s developer guide says, image, video and speech retrieval keep about 95% of their full quality.

Code showed the biggest benchmark gain. EmbeddingGemma 2 scored 78.68 on the code section of the Massive Text Embedding Benchmark, almost 10 points above the first version. Google is pitching the result at developers who build retrieval for coding agents. Multilingual text scores barely moved.

The company also claims leading results among multimodal embedding models under 1 billion parameters, and it said the model outperforms some specialist models more than twice its size on image, video and audio tasks.

Because EmbeddingGemma 2 shares a text tokenizer and an audio encoder with Gemma 4, an on-device retrieval-augmented generation setup running both needs less memory than two unrelated models would. Google’s AI Edge Foresight meeting app for Mac already runs the pair together. In Google’s AI Edge Gallery demo app, a Video Moments Finder feature locates a scene inside a video from a typed or spoken query.

Model weights are available now from Hugging Face Inc. and Google’s Kaggle under an Apache 2.0 license that permits commercial use. Google said the model will reach the Model Garden in Gemini Enterprise Agent Platform soon, and the weights already work with open-source serving tools such as vLLM, llama.cpp and Ollama.

Image: Google

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

 

About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

Send us a news tip

Send us a News Tip

  • This field is for validation purposes and should be left unchanged.
  • Max. file size: 244 MB.

Sign in

SIGN IN

Bio

Ethics statement

Extract the signal from the noise

Get SiliconANGLE updates and analysis.

Contact us

Partner with us

Contact us

Guest inquiry