Skip to content
AI News HubLIVE
Source content · Analysis pending6 min read

EmbeddingGemma 2: Text, Code, Images, Video and Audio in One Vector Space

Summary

EmbeddingGemma 2 launched on October 6, 2026 under Apache 2.0. It is a sub-1B model built on Gemma 4 that maps text, code, images, video and audio into one 768-dimensional space. This article covers the architecture, the benchmarks, and runnable scripts to provide measured results. Specifications Specification EmbeddingGemma 2 Base model Gemma 4 License Apache […] The post EmbeddingGemma 2: Text, Code, Images, Video and Audio in One Vector Space appeared first on Analytics Vidhya.

SourceAnalytics VidhyaAuthor: Sree Vamsi
EmbeddingGemma 2: Text, Code, Images, Video and Audio in One Vector Space
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

EmbeddingGemma 2: Google's Open Multimodal Model India's Most Futuristic AI Conference Is Back – Bigger, Sharper, Bolder d : h : m : s Career GenAI Prompt Engg ChatGPT LLM Langchain RAG AI Agents Machine Learning Deep Learning GenAI Tools LLMOps Python NLP SQL AIML Projects Reading list How to Become a Data Analyst in 2025: A Complete RoadMap A Comprehensive Learning Path to Tableau in 2025 A Comprehensive NLP Learning Path 2025 Learning Path to Become a Data Scientist in 2025 Step-by-Step Roadmap to Become a Data Engineer in 2025 A Comprehensive MLOps Learning Path: 2025 Edition Roadmap to Become an AI Engineer in 2025 A Comprehensive Learning Path to Master Computer Vision in 2025 Best Roadmap to Learn Generative AI in 2025 GenAI Roadmap for Enterprises Large Language Models Demystified: A Beginner’s Roadmap Learning Path to Become a Prompt Engineering Specialist EmbeddingGemma 2: Text, Code, Images, Video and Audio in One Vector Space Sree Vamsi Last Updated : 09 Oct, 2026 11 min read EmbeddingGemma 2 launched on October 6, 2026 under Apache 2.0. It is a sub-1B model built on Gemma 4 that maps text, code, images, video and audio into one 768-dimensional space. This article covers the architecture, the benchmarks, and runnable scripts to provide measured results. Table of contents Specifications How it works under the hood Benchmarks and evaluation Getting started Step 0: Check your setup Step 1: Text and code retrieval Step 2: Cross-modal search Step 3: Matryoshka truncation Step 4: Code search Choosing your configuration Conclusion Specifications Specification EmbeddingGemma 2 Base model Gemma 4 License Apache 2.0 Output dimension 768, truncatable to 512, 256, 128 Context window 8,192 tokens, shared across modalities Modalities Text, code, images, video, audio Parameters 270M text-only to 740M full multimodal Released October 6, 2026 How it works under the hood The usual way to search across mixed content is to run one model per content type and stitch the results together afterwards. That means separate indexes, separate score ranges, and no reliable way to compare a photo against a paragraph. EmbeddingGemma 2 removes that problem by sending every content type through one backbone. Each modality gets its own front-end encoder, but the output always lands in the same 768-dimensional space. Architecture diagram, Google Developers Blog. The three encoders Encoder Size Handles Built on Text and code 270M, the base Plain text and source code, up to 8,192 tokens An adapted Gemma 4 decoder Vision +170M Photos, charts, slides, PDF pages, and video frames A dedicated vision encoder Audio +300M Speech and ambient sound, fed in as raw audio A dedicated audio encoder The text encoder is the base and is always present. Vision and audio are additions on top of it, which is where the four parameter counts come from. Why modular loading matters One checkpoint, four ways to load it. You are not choosing between four different models. There is one set of weights, and you decide at load time which encoders get read into memory. Skipped encoders cost nothing, on disk or in RAM. Because the output space is identical in all four cases, a query embedded with the 270M text setup can be matched against documents embedded with the full 740M model. Two practical consequences: Start text-only and add images later without re-embedding anything you already indexed. Run a small config on an edge device and a large one on a server, against the same index. Pairing with Gemma 4 If you run Gemma 4 as the generator in an on-device RAG stack, the two models share a text tokenizer and the same audio encoder design. That overlap is loaded once rather than twice, which matters when the constraint is phone memory rather than server memory. Benchmarks and evaluation What Google published Benchmark Result MTEB (Code) 14% higher than EmbeddingGemma 1 Multilingual text Retains EmbeddingGemma 1 accuracy Image, video, audio retrieval New capability, not present in version 1 Google published the MTEB (Code) gain as a percentage improvement rather than absolute scores, so there is no single number to compare against other models directly. Quality retention under truncation Dimension Text and code Image, video, speech Storage per 1M vectors 768d baseline baseline 1,465 MB 512d no stated loss no stated loss 977 MB 256d near baseline roughly 95% 488 MB 128d roughly 90% roughly 75% 244 MB Two things to read off that table. Storage figures assume bfloat16, and the media columns fall away faster than text does. Google flags 128d multimodal as the case to test before you ship it, and the 75% figure explains why. What this run measured Results from the five scripts below, on a small test corpus. These are measurements on one corpus, not benchmark scores. Measurement Result 256d versus 768d ranking Identical top-1 and MRR 128d ranking Top-1 held, MRR fell from 0.667 to 0.656 Mean similarity at 128d Rose 12.6% versus 768d despite identical ranking Image cross-modal margin 0.1154 over the nearest wrong query Audio cross-modal margin 0.0137 over the nearest wrong query Default output precision float32, not bfloat16, so indexes are double the quoted size Getting started First install sentence transformer and other relvant libraries using: pip install -U sentence-transformers[image,audio,video] transformers Requires sentence-transformers 6.1.0 or later. Step 0: Check your setup Reports Python version, library version, available modalities and hardware before you download 740M parameters. File: 00_check_setup.py import sys import importlib import shutil def ok(b): return "OK " if b else "-- " print("=" * 66) print("EmbeddingGemma 2 setup check") print("=" * 66) print(f"\nPython {sys.version.split()[0]} (3.9+ needed)") # --- core library ------------------------------------------------- try: import sentence_transformers as st ver = st.version major, minor = (int(x) for x in ver.split(".")[:2]) good = (major, minor) >= (6, 1) print(f"{ok(good)}sentence-transformers {ver} (need 6.1.0+)") if not good: print(" pip install -U sentence-transformers") except ImportError: print("-- sentence-transformers NOT INSTALLED") print( " pip install -U sentence-transformers[image,audio,video] " "transformers" ) sys.exit(1) # --- which modalities can actually run --------------------------- print("\nModality support:") mods = { "text / code": [], # always available "images": ["PIL"], "audio": ["soundfile", "librosa"], "video": ["decord"], } for name, deps in mods.items(): missing = [ d for d in deps if importlib.util.find_spec(d) is None ] print(f" {ok(not missing)}{name:5} {label: {doc_emb.shape}\n") for q in QUERIES: q_emb = model.encode_query(q) sims = model.similarity(q_emb, doc_emb)[0] best = int(sims.argmax()) print(f"Q: {q}") print(f" score {sims[best]:.4f} -> {DOCS[best][:72]}...") # Show the runner-up so you can see the margin. order = sims.argsort(descending=True) second = int(order[1]) print( f" runner-up {sims[second]:.4f} " f"margin {sims[best] - sims[second]:.4f}\n" ) print( "If the margin is small, your corpus has near-duplicates " "or the query is ambiguous." ) print( "That number is more useful than the raw score " "for debugging retrieval." ) Output: All three queries retrieved the correct document. Margins: Query Top score Margin what causes the northern lights 0.8525 0.1299 how do I reverse a linked list 0.8485 0.0599 why does my database table keep growing 0.6673 0.0407 The database query has the lowest score and thinnest margin because the corpus holds two competing database documents. Margin matters more than raw score. A 0.04 margin means different phrasing could flip the result. Thin margins across a corpus point to chunking or deduplication problems, not the model. Step 2: Cross-modal search Media is passed as a dictionary keyed by modality with no prompt. Only the text query gets a task prompt. File: 02_multimodal.py from sentence_transformers import SentenceTransformer MODEL_ID = "google/embeddinggemma-2" print("Loading full multimodal model (740M)...") model = SentenceTransformer(MODEL_ID) # Media is passed as a dict keyed by modality, with NO prompt. # Only the text query gets a task prompt. image_emb = model.encode({"image": "data/sunset_beach.jpg"}) audio_emb = model.encode({"audio": "data/ocean_waves.wav"}) for query in [ "ocean waves at sunset", "a busy city street", "someone playing piano", ]: q = model.encode_query(query) s_img = float(model.similarity(q, image_emb)[0][0]) s_aud = float(model.similarity(q, audio_emb)[0][0]) print(f"\n{query!r}") print(f" vs image : {s_img:+.4f}") print(f" vs audio : {s_aud:+.4f}") # Interleaved: ONE embedding covering text, photo and video together. # The markers say where each media item sits inside the text. listing_emb = model.encode( { "text": ( "Waterproof trail shoe. " "Grip test on wet rock: " ), "image": "data/trail_shoe.jpg", "video": "data/grip_test.mp4", } ) q = model.encode_query("waterproof trail shoes") print( f"\ninterleaved product listing : " f"{float(model.similarity(q, listing_emb)[0][0]):+.4f}" ) print( "\nThe control to watch: scores for unrelated queries " "should sit clearly lower." ) print( "If 'a busy city street' scores close to 'ocean waves at sunset' " "on the same" ) print( "image, the embedding is not discriminating and retrieval " "will be noisy." ) Output: Measured. Image separates the correct query cleanly. Audio does not. Image: match 0.7026, nearest wrong 0.5872. Margin 0.1154. Thresholdable. Audio: match 0.6466, nearest wrong 0.6329. Margin 0.0137. All three queries inside a 0.045 band. No usable threshold exists on that audio result. Possible causes: the audio encoder is weaker at cross-modal matching, or ocean waves and traffic sit close together because both are broadband noise. Validate audio retrieval on your own clips. Do not assume parity with image. Step 3: Matryoshka truncation Measures what truncation costs on your corpus rather than on Google’s. File: 03_matryoshka.py import numpy as np from sentence_transformers import SentenceTransformer MODEL_ID = "google/embeddinggemma-2" DIMS = [768, 512, 256, 128] model = SentenceTransformer( MODEL_ID, config_kwargs={ "vision_config": None, "audio_config": None, }, ) # Replace these with YOUR corpus and YOUR queries plus known-correct answers. DOCS = [ "Postgres VACUUM reclaims storage from dead tuples after updates and deletes.", "A B-tree index stores keys sorted and supports efficient range scans.", "Connection pooling reuses database connections to avoid per-request handshakes.", "Write-ahead logging records changes before they are applied to data files.", "Table partitioning splits one large table into smaller physical pieces.", "The northern lights are caused by charged particles from the sun.", "Linked list reversal rewires each node's next pointer while walking the list.", "A hash index supports equality lookups but not range queries.", ] # (query, index of the correct document) GOLD = [ ("why does my table keep growing after deletes", 0), ("speed up queries over a date range", 1), ("too many database connections being opened", 2), ("how is durability guaranteed on crash", 3), ] print( f"{'dim':>5} {'storage/1M':>12} {'top-1':>7} " f"{'MRR':>7} {'mean score':>11}" ) print("-" * 52) results = {} for d in DIMS: doc_emb = model.encode_document( DOCS, truncate_dim=d, normalize_embeddings=True, ) hits, rr, scores = 0, [], [] for q, gold_idx in GOLD: q_emb = model.encode_query( q, truncate_dim=d, normalize_embeddings=True, ) sims = model.similarity(q_emb, doc_emb)[0].cpu().numpy() order = np.argsort(-sims) rank = int(np.where(order == gold_idx)[0][0]) + 1 hits += rank == 1 rr.append(1.0 / rank) scores.append(float(sims[gold_idx])) mb = 1_000_000 * d * 2 / 1024**2 # bfloat16 result [truncated for AI cost control]

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • EmbeddingGemma 2 launched on October 6, 2026 under Apache 2.0. It is a sub-1B model built on Gemma 4 that maps text, code, images, video and audio into one 768-dimensional space.…

Highlights and analysis are generated automatically and may contain errors. Check the original source.