Cohere released Embed 5 Pro and Fast on September 30, 2026, giving teams building search and retrieval-augmented generation (RAG) a new choice: encode their documents with one tier and their queries with the other. The release documentation confirms that both models share an embedding space, allowing a Pro-built index to serve Fast queries without rebuilding it.
The practical recommendation: evaluate Pro indexing with Fast queries, but compare it with all-Fast, all-Pro and your existing system. Re-index when an identified retrieval problem or operating target justifies it. A cheaper query encoder alone is not a reason to replace a working corpus index.
Shared space changes query routing, not backward compatibility
The model IDs are embed-v5.0-pro and embed-v5.0-fast. Cohere recommends Pro for quality-focused indexing and Fast for the live query path. Its compatibility note requires the document and query vectors to have the same output dimension; it also covers reduced dimensions and int8 quantization.
Both tiers accept text, images and mixed text/image inputs, support more than 100 languages, and offer a 128K-token context. The model table lists six dimensions: 256, 512, 768, 1024, 1536 and 2048. It also shows that Embed 4 already supported mixed inputs and 128K context; those capabilities are not new with this release.
The documented compatibility is between the two v5 tiers. It does not establish that v5 queries can search an Embed 4 or another vendor’s index. Plan a separate v5 index for migration unless the provider explicitly supports another arrangement. Matching vector lengths is not evidence of matching embedding spaces.
What the Pro/Fast price difference actually saves
Cohere’s API list prices are $0.12 per million text tokens for Pro and $0.08 for Fast. The launch rate card lists image tokens at $0.40 per million for either tier. The one-third discount therefore applies to text embedding, not every input or the whole RAG bill.
Illustrative calculation, not an observed bill: assume one initial ingestion of 1 million text chunks averaging 500 billable tokens each, plus 1 million queries per month averaging 50 billable tokens. That is 500 million corpus tokens and 50 million monthly query tokens. Cost equals billable tokens divided by 1 million, multiplied by the relevant rate.
Document / query tier | Initial indexing | Monthly queries | First-month total |
|---|---|---|---|
Pro / Pro | $60 | $6 | $66 |
Pro / Fast | $60 | $4 | $64 |
Fast / Fast | $40 | $4 | $44 |
In this scenario, switching only the queries to Fast saves $2 per month: 33.3% of query-embedding spend, but about 3.0% of the first-month all-Pro embedding total. The mixed route still pays a $20 indexing premium over all-Fast. That premium needs a retrieval-quality benefit; the query discount does not pay it back against an all-Fast baseline.
These USD estimates exclude corpus refreshes, parsing, vector-database service, reranking, generation, networking, retries, taxes and discounts. Image-heavy workloads need a separate image-token estimate: the posted rate is not a price per PDF page.
Choose vector size and document representation separately
Encoder tier is only one cost decision. The model documentation defaults v5 to 2048 dimensions, versus 1536 for v4. Calculated at the same float32 precision, the new default uses 33.3% more raw vector storage. An upgrade does not automatically shrink an index.
The embedding guide describes reduced-dimension and compressed outputs. For 10 million vectors, the following are calculated payload sizes, not measured database footprints. The formula is vector count × dimensions × bytes per value; packed binary uses one bit per dimension.
Representation | Bytes per vector | Raw values for 10 million vectors |
|---|---|---|
2048-dimensional float32 | 8192 | 81.92 GB |
1024-dimensional int8 | 1024 | 10.24 GB |
256-dimensional packed binary | 32 | 0.32 GB |
GB here means decimal gigabytes. Index structures, IDs, metadata, replicas and any retained higher-precision vectors add storage. These layouts are not quality-equivalent: evaluate dimension and precision changes separately from the Pro/Fast comparison.
For PDFs, also decide what each vector represents. Compare parsed text with page images where the answer depends on charts, table geometry or scanned content. The mixed-input guide describes text and image components, not a generic raw-PDF upload. Preserve page identifiers and source originals so retrieved evidence remains citable.
For the upstream document-reading decision, RohitAI’s North Micro Vision analysis discusses page fidelity and local visual processing. That model reads documents; Embed 5 produces vectors for finding them.
What Cohere’s benchmarks do—and do not—establish
Cohere’s cross-model evaluation reports mean nDCG@10 across 40 development datasets, normalized so Pro documents with Pro queries equal 100. Pro documents with Fast queries score 98.4; Fast documents with Fast queries score 96.6. These are relative ranking scores, not 98.4% or 96.6% answer accuracy, and not a guarantee for your corpus.
Its ViDoRe V3 RCP-nDCG@10 results are 85.8 for Pro, 84.5 for Fast and 77.0 for Embed 4. Cohere identifies that comparison as parsed-text evaluation. The launch footnote says RCP scores measure reordering a fixed candidate set, and the published evaluation protocol restricts ViDoRe rankings to the judged pool.
The implication: better ranking within that pool does not prove your production search will find more relevant pages across the full corpus. Measure first-stage Recall@k—how much relevant evidence appears among the retrieved candidates—alongside ranking and answer quality. Keep parsed-text and page-image results separate. These figures are Cohere’s measurements, not a RohitAI test.
A migration decision you can reverse
Define the reason to change. Identify missed evidence, expensive operations or unacceptable latency. Build a representative query set with human-reviewed relevance labels and no-answer cases, separated by language and query type. Set acceptance criteria before examining results.
Compare the four candidates. Evaluate the incumbent, all-Pro, Pro-document/Fast-query and all-Fast. Hold chunking, filters, candidate counts and reranking constant; match dimensions and precision where supported. Record unavoidable differences instead of attributing every gain to the encoder.
Measure the complete retrieval path. Track Recall@k, ranking quality, evidence-supported answer success, false positives on no-answer queries, p50/p95 latency and total cost. Recheck any similarity threshold used to reject weak matches when changing query encoders; shared-space support does not guarantee identical threshold behavior.
Keep the old index until the new one passes. Use separate versioned indexes, preserve source-to-chunk identities, and record model IDs, preprocessing, dimensions and precision. Replay or shadow queries before switching traffic. If Fast meets the targets, use it; retain Pro for query classes only where its measured gain justifies the cost.
For text retrieval, the Embed API reference distinguishes search_document from search_query. Set the roles and output dimension explicitly. Its default truncation discards the end of overlength input; truncate=NONE instead returns an error, which can help catch ingestion mistakes during evaluation.
Check the deployment and ingestion constraints
Cohere’s release note lists API, Model Vault, Microsoft Foundry and Amazon SageMaker access. However, the Microsoft catalogs for Pro and Fast label both models Preview. Verify the chosen region, lifecycle and account entitlement rather than treating the launch’s broad availability language as a universal GA guarantee.
Dedicated pricing also differs from API economics: Model Vault’s rate card lists the same instance prices for Pro and Fast at each listed size. Fast’s lower text-token rate does not establish a cheaper dedicated deployment without throughput and capacity evidence.
The default Embed quota is 2,000 inputs per minute. Calculated at that ceiling, the example’s 1 million text chunks need at least 500 minutes—8 hours 20 minutes—even before failures or processing delays. The limit counts inputs, not HTTP requests; batching does not remove it. Other deployments or negotiated quotas can differ.
Finally, the launch advertises batch embedding, but its linked Embed Jobs guide still explicitly restricts compatibility to v3. Confirm the supported v5 asynchronous ingestion path before building around that endpoint.
Methodology: This AI-assisted guide draws on Cohere’s published launch, documentation, pricing and evaluation protocol, plus Microsoft’s model catalogs, checked September 30, 2026. Cost, storage and ingestion-time examples are arithmetic under stated assumptions. No models were run or benchmarks independently reproduced.
