Skip to content
AI News HubLIVE
Source content · Analysis pending6 min read

Cohere Embed 5: Frontier Embedding Models for Enterprise

Summary

Key takeaways State-of-the-art enterprise retrieval: Embed 5 Pro achieves the highest average score of any model we tested - particularly across financial datasets, parsed PDFs, and visually rich documents. A new Fast t…

Cohere Embed 5: Frontier Embedding Models for Enterprise
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

Key takeaways State-of-the-art enterprise retrieval: Embed 5 Pro achieves the highest average score of any model we tested - particularly across financial datasets, parsed PDFs, and visually rich documents. A new Fast tier: Embed 5 Fast brings strong retrieval quality to latency - and cost-sensitive workloads, at $0.08 per million tokens. One index, two models: Pro and Fast share an embedding space, so teams can index with Pro and query with either model without re-indexing. Built for complex enterprise data: Embed 5 supports multimodal inputs and retrieval, 100+ languages, and a 128K-token context window for longer documents. More efficient at scale: Matryoshka representations and lower-precision outputs reduce vector storage and search costs, while quantized weights lower serving requirements for private deployments. Today, we're releasing Embed 5, a new family of embeddings models at the frontier of high-quality enterprise retrieval. Embed 5 delivers stronger retrieval across complex enterprise data while giving teams more control over latency, cost, and deployment. Embed 5 Pro is optimized for maximum quality across multimodal, multilingual, financial, code, and parsed-document retrieval. Embed 5 Fast brings highly competitive performance to latency- and cost-sensitive workloads. Both tiers share a single embedding space, so teams can index with Pro and query with either model without rebuilding the index. Embed 5 establishes the retrieval foundation for search, RAG, and agentic workflows, surfacing more relevant context while filtering out noise before it reaches expensive generative models. Use Embed to improve answer quality and user experience while helping keep downstream inference costs under control. Embed 5 is generally available today on the Cohere API and Model Vault, Microsoft Foundry, and Amazon SageMaker. Pricing is $0.12 per million tokens for Pro and $0.08 per million tokens for Fast. Snapshot Capability Embed 5 Pro Embed 5 Fast Best for Maximum retrieval quality; offline indexing; complex enterprise corpora Interactive search; high-volume RAG; agentic retrieval Context length 128K tokens 128K tokens Inputs Text, images, fused text + image Text, images, fused text + image Languages 100+ 100+ Output dimensions 2048, 1536, 1024, 768, 512, 256 2048, 1536, 1024, 768, 512, 256 Embedding formats float, int8, binary float, int8, binary Matryoshka Embeddings Yes Yes Shared embedding space Yes Yes Supports self-hosting Yes Yes Pricing $0.12 / 1M tokens $0.08 / 1M tokens Performance Embed 5 Pro delivers our strongest retrieval performance to date. It achieves the highest average score of any model we tested across ViDoRe V3, financial documents, parsed PDFs, image retrieval, and across key business languages. Embed 5 is also the first model family evaluated with RCP-nDCG@10, our latest retrieval methodology. Instead of scoring only against a limited set of fixed labels, it evaluates retrieved documents against query-specific relevance criteria, capturing relevant results and giving a fuller view of performance on your own corpus1. Read more about RCP-nDCG@10. Enterprise documents Embed 5 excels with visually rich documents where meaning lives in tables, charts, diagrams, and layout - not just text. On ViDoRe V3, which features documents sampled across key enterprise domains including financial filings, technical manuals, regulatory material, government reports, textbooks, and lectures, Embed 5 Pro averages 86.1 - an impressive 8.8-point gain from Embed 42. That puts it ahead of Voyage 4 Large (83.7), Gemini Embedding 2 (83.2), and OpenAI text-embedding-3-large (75.3). Pro leads five of the eight domains outright and ties Voyage 4 Large on energy, with its largest gains over Embed 4 on HR (+11.4) and industrial (+10.3). Embed 5 Fast averages 84.7, ahead of both Gemini Embedding 2 and Voyage 4 Large. See the full results here. Retrieval quality (RCP-nDCG@10) across eight visually rich document domains. Evaluations consisted of parsed text outputs, curated by the authors of ViDoRe. Finance Embed 5 Pro establishes itself as the leading embeddings model for financial document retrieval. Pro ranks first on three leading public financial benchmarks, with Fast second on each despite being considerably smaller than its peers: FinanceBench (80.1 Pro, 80.0 Fast), FinQA (90.0, 88.8), and ViDoRe V3 Finance (85.0, 83.9). Across these, Pro averages 3.3 points higher than the next non-Cohere competitor, Gemini Embedding 2. Compared with OpenAI text-embedding-3-large, the lead grows to 21.4 points on FinanceBench. Retrieval quality (RCP-nDCG@10) across eight visually rich document domains. Evaluations consisted of parsed text outputs, curated by the authors of ViDoRe. Multimodal Parsed PDFs Most enterprise search pipelines still convert PDFs to text before embedding them, but that process can strip away structure. Tables lose row and column relationships, multi-column layouts can scramble reading order, repeated headers add noise, and charts often disappear entirely. That makes parsed-document retrieval a harder test than clean-text benchmarks suggest. Our parsed-document suite spans service documentation, corporate reports, SEC filings, product manuals, and privacy policies. Embed 5 Pro achieves the highest average across the suite at 84.8, ahead of Voyage 4 Large at 83.6, Embed 5 Fast at 83.4, Gemini Embedding 2 at 80.8, and Embed 4 at 78.6. The figure below highlights a subset of familiar public benchmarks, with Embed 5 Pro especially strong on financial documents represented by FinanceBench and CoFiF. Most enterprise search pipelines still convert PDFs to text before embedding them, but that process can strip away structure. Tables lose row and column relationships, multi-column layouts can scramble reading order, repeated headers add noise, and charts often disappear entirely. That makes parsed-document retrieval a harder test than clean-text benchmarks suggest. Our parsed-document suite spans service documentation, corporate reports, SEC filings, product manuals, and privacy policies. Embed 5 Pro achieves the highest average across the suite at 84.8, ahead of Voyage 4 Large at 83.6, Embed 5 Fast at 83.4, Gemini Embedding 2 at 80.8, and Embed 4 at 78.6. The figure below highlights a subset of familiar public benchmarks, with Embed 5 Pro especially strong on financial documents represented by FinanceBench and CoFiF. Parsed-document retrieval (RCP-nDCG@10) across representative public benchmarks. Documents were parsed using Gemini 1.5 Flash. ‘Full suite average’ includes additional datasets not shown. Page-image and fused text-image documents Some documents are better represented visually. Scanned pages, slide decks, schematics, and charts contain information that text extraction may miss. Embed 5 can embed page images directly (page-image), or combine an image with its metadata into a single vector (fused text-image). On fused text-image corpora, Embed 5 Pro averages 82.3 across five datasets, ahead of Embed 5 Fast at 81.2 and Gemini Embedding 2 at 61.3. Pro outperforms Gemini Embedding 2 on every dataset in the suite. Page image retrieval is also robust: Embed 5 Pro continues to lead on financial datasets, averaging 77.0 from five datasets, ahead of Embed 5 Fast (73.2), Embed 4 (71.1), Voyage Multimodal 3.5 (70.1), and Gemini Embedding 2 (56.7). Multimodal document retrieval (nDCG@10) with text queries. Fused text-image retrieval combines the page image and document metadata in a single embedding. RepairBench reports the average across multiple multilingual query subsets. High Finance is an internal, Cohere-annotated set of questions that ask models to retrieve the relevant investment banking and hedge fund presentation material. Multimodal document retrieval (nDCG@10) with text queries. Image-only retrieval uses the page image alone to answer a prompt. AR = Arabic; JA = Japanese; KO = Korean. Multilingual Embed 5 is trained on more than 100 languages, with particular focus on the languages most used by our global customer base. Across German, French, Spanish, Italian, and Russian, Embed 5 Pro achieves the highest average of the models we tested: 77, compared with 76 for Voyage 4 Large, and 73 for Gemini Embedding 2. It improves on Embed 4 by around 7 points on average, with the largest gains in Russian (+9) and Italian (+7). Retrieval quality across key European languages. Scores represent an average across a number of composite benchmarks. Benchmarks were measured in either nDCG@10 or RCP-nDCG@10. The table below covers ten further languages where Embed 5 has made important strides against Embed 4. Pro’s largest gains are in middle eastern and subcontinent languages, notably Farsi (+12.8), Telugu (+12.3), and Hindi (+11.5). For the full list of multilingual evaluation results, click here. Language Cohere Embed 5 Pro Cohere Embed 5 Fast Gemini Embedding 2 Voyage 4 Large Zembed-1 (4B) Cohere Embed 4 Jina Embeddings v5 Text Small OpenAI text-embedding-3-large Japanese 86.6 84.8 90 87.1 84.7 83.1 82.7 79.9 Chinese 82.4 80.3 80.7 82 84.6 78.9 78 72.9 Korean 85 83.2 87.3 85.6 81.8 78.9 79.2 69.7 Arabic 82.8 79.3 86.5 85.9 79.5 72.2 70.9 67 Farsi 80.6 78 82.7 78.7 73.8 67.8 70.2 59.9 Hindi 79.6 77.4 84.3 82.9 79.7 68.1 72.8 59.3 Bengali 83.1 80.7 89.1 85.1 80.2 72.9 79.1 61.1 Telugu 80.3 75.5 91.1 88.6 67.3 68 82 63.1 Indonesian 85.2 82.7 87.8 84.5 83.9 78.6 79.1 81.2 Thai 82 74.6 87.7 84 78.7 74.5 77.8 66.9 Meet Embed 5 Fast Embed 5 Fast is a lighter weight model built for latency-sensitive, high-volume retrieval. It costs a third less than Pro while retaining the same 128K-token context, multimodal inputs, multilingual coverage, and multiple compressed output formats. That matters most on the query path, where embedding latency is paid on every search - and multiplied in agentic workflows that may issue dozens of searches per task. Fast’s smaller footprint also lowers serving costs in private deployments and speeds large ingestion and re-indexing jobs. For document throughput - a closer proxy for indexing efficiency - Fast is consistently more efficient, delivering an average of 2.4× higher throughput than Pro across context sizes. Average document throughput for Fast and Pro across ~200-token and ~1K-token context lengths, measured in documents processed per second. Higher is better. Performance Fast raises the bar for compact embedding models. On ViDoRe V3, it leads Voyage 4 Nano by more than seven points and Jina Embeddings v5 Text Small, Perplexity, and Microsoft’s Harrier 0.6B by ten or more. It outperforms Qwen3-VL-Embedding-2B, despite being roughly half the size, by about 20 points. As seen above, its average also exceeds Gemini Embedding 2 and Voyage 4 Large on ViDoRe V3 and on financial retrieval. On parsed PDFs it exceeds Gemini Embedding 2 (83.4 vs 80.8) and trails Voyage 4 Large (83.6). Retrieval quality (RCP-nDCG@10) across eight visually rich document domains. Evaluations consisted of parsed text outputs, curated by the authors of ViDoRe. Pro Fast Use when… Use for offline indexing and quality-critical retrieval — especially across complex documents, multimodal content, or nuanced queries. Use on the live request path, especially for interactive search, agent loops, and other high-volume query workloads. Financial services Bulk indexing of 10-Ks, earnings reports, tables, and footnotes for equity research; compliance or risk search over dense financial records. Customer-service search, advisor copilots, transaction-support workflows, and agents issuing repeated retrieval calls. Retail + Commerce Product discovery across large multimodal catalogs, including nuanced attribute matching and image-plus-text retrieval. Site search, shopping assistants, recommendations, and conversational product lookup serving large numbers of live queries. Legal Digitiz [truncated for AI cost control]

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • Key takeaways State-of-the-art enterprise retrieval: Embed 5 Pro achieves the highest average score of any model we tested - particularly across financial datasets, parsed PDFs, a…

Highlights and analysis are generated automatically and may contain errors. Check the original source.