AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Google announced EmbeddingGemma 2 on Oct. 6, 2026, a 740-million-parameter embedding model designed to map text, code, images, audio and video into a shared vector space. Google says it is available under the Apache 2.0 license and can run on consumer hardware, with memory needs varying by configuration and device.

Google DeepMind announced EmbeddingGemma 2 on Oct. 6, a 740-million-parameter embedding model that maps text, code, images, audio and video into a shared representation for search and retrieval. Released under the commercially permissive Apache 2.0 license, it is designed to run on local devices, potentially letting apps search across different media without sending their contents to a remote service.

The model extends Google’s earlier EmbeddingGemma, which focused on text. Google says that first version has been downloaded more than 20 million times. EmbeddingGemma 2 is built on the Gemma 4 architecture and is intended to support tasks such as finding a video clip using a voice memo or searching audio recordings with a text query. Those are examples of intended uses, not independently verified product results.

Google describes the model as modular: text-only workloads can use a configuration with as little as 270 million parameters, while optional vision and audio encoders are listed at 170 million and 300 million parameters, respectively. Its 8,192-token context window is four times the size of the first EmbeddingGemma’s, according to Google. The company says it can cover up to 5.5 minutes of audio, 29 images or 58 video frames, or combinations of those inputs, subject to the model’s processing limits.

For storage, Google says developers can shorten output vectors from 768 dimensions to 512, 256 or 128 using Matryoshka Representation Learning. It estimates this can cut local vector-database storage and memory use by up to six times. On a Google Pixel 11 Pro, Google reports that quantized weights require about 191 MB of active RAM for text-only use and about 567 MB for the full multimodal model. The figures are vendor-reported and tied to that stated device and setup.

At a glance
announcementWhen: Announced Oct. 6, 2026; some distributi…
The developmentGoogle DeepMind announced EmbeddingGemma 2, a commercially permissive multimodal embedding model intended for on-device search and retrieval.

Local Search Across Media Types

Embeddings let software represent content as vectors so that related items can be found by meaning, rather than only by matching exact words. A shared multimodal embedding space could allow developers to build search that connects a query in one format to content in another—for example, text to video or voice to image—without maintaining separate systems for each media type.

Running the embedding process on a device may also reduce the need to upload private recordings, images or code to a service, and can support use when a network connection is unavailable. Those are potential advantages of local processing, not guarantees: privacy depends on the full application design, and the source does not provide independent measurements of latency, accuracy or energy use across devices.

Google positions the model for on-device retrieval-augmented generation, in which a retrieval system supplies relevant material to a generative model. It says EmbeddingGemma 2 shares a text tokenizer and audio encoder with Gemma 4, which may help developers run them together with a lower combined memory footprint. The announcement does not quantify that footprint for a complete application.

Amazon

on-device multimodal embedding model

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Text Embeddings to Multimodal

Google introduced the original EmbeddingGemma last year as a lightweight option for text embeddings on consumer hardware. The company says developers used it for local search and privacy-focused retrieval-augmented generation. The new release carries that on-device focus into code, images, audio and video while retaining a text-only configuration for projects that do not need every modality.

Google says EmbeddingGemma 2 achieved leading results among multimodal embedding models with fewer than one billion parameters on benchmarks including MTEB Code and the Massive Audio Embedding Benchmark. It also reports a 9.92-point improvement over the original model on MTEB Code, from 68.76 to 78.68. These are company-reported evaluations; the announcement points readers to the model card for full metrics, and benchmark scores do not establish performance in every application.

The weights are listed on Hugging Face and Kaggle. Google says developers can deploy through Google AI Edge MediaPipe or LiteRT, and lists support or serving options involving tools such as Transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio. The model is also intended for browser use through transformers.js or WebGPU. Availability through Google’s Gemini Enterprise Agent Platform Model Garden was described as coming soon.

““EmbeddingGemma 2” expands beyond text “to unify code, images, video, and audio in a shared embedding space.””

— Sahil Dua and Henrique Schechter Vera, Google DeepMind research engineers

Amazon

multimodal search engine software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Performance and Deployment Questions

Google’s announcement does not include independent evaluations of the model or comparative testing across a broad range of hardware. The reported memory figures apply to a Pixel 11 Pro with quantization; actual requirements may differ with device, runtime, configuration and workload. It is also not yet clear how performance changes when the model processes mixed media near its stated context limits.

The company characterizes benchmark performance as leading for its size, but the announcement does not spell out all test conditions or establish how scores translate to real-world search quality. Developers will need to review the model card and test their own data to assess retrieval accuracy, language coverage, speed and resource use. Google also has not provided a date for Model Garden availability.

Amazon

local media search device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Model Access and Developer Testing

Developers can access the weights through Hugging Face and Kaggle, and Google points to its developer documentation and guides for inference and fine-tuning. The company also directs developers to MediaPipe and LiteRT for deployment, with browser and third-party serving options listed in its announcement. Model Garden access is expected later, but Google gave no launch date.

The next practical test will be how the model performs in applications that combine media types on real devices, including whether its storage and memory trade-offs suit local databases and retrieval pipelines. Google has published the headline benchmark claims and device-specific memory estimates; broader independent testing and further availability details remain pending.

Amazon

AI embedding model for multimedia

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is EmbeddingGemma 2?

It is a 740-million-parameter embedding model from Google DeepMind that represents text, code, images, audio and video in a shared vector space for search and retrieval.

Is EmbeddingGemma 2 open source?

Google released it under the Apache 2.0 license, which the company describes as commercially permissive. The model weights are listed on Hugging Face and Kaggle.

Can it run entirely on a device?

Google designed it for on-device use and reports quantized memory requirements of about 191 MB for text-only and 567 MB for full multimodal use on a Pixel 11 Pro. Requirements can vary across devices and configurations.

What is its context limit?

The model has an 8,192-token context window. Google says this can cover up to 5.5 minutes of audio, 29 images or 58 video frames, or combinations of those inputs.

Where can developers get it?

Google lists the weights on Hugging Face and Kaggle. Availability through the Gemini Enterprise Agent Platform Model Garden was described as coming soon, with no date given.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can AI Design Circuit Boards Yet?

Exploring whether AI can now design circuit boards, what is confirmed, claims made, and what remains uncertain in this emerging field.

DeepSeek V4 Pro 0813

DeepSeek has announced the release of V4 Pro 0813, a new version promising improved search performance and expanded features for enterprise data management.

Inside My September 2026 AI Stack: Opus Builds, Sol Digs, Jev Decides

Thorsten Meyer’s September 2026 AI stack: Opus 5.5 builds, GPT-6.1 Sol reviews at a fraction of the cost, and Jev routes high-volume decisions.

Candor as a Moat: A Critical Reading of Dario Amodei and Anthropic

Examining how Dario Amodei’s transparency and policy proposals serve as a strategic barrier for Anthropic amid AI advancements.