provider-capability

Cohere Embeddings, retrieval, and reranking

Direct answerUse this guide to understand the Cohere Embeddings, retrieval, and reranking capability, then verify the exact callable model ID, endpoint, and price in the model marketplace before integration.

Updated · Reviewed

Capability boundary

The Cohere Embeddings, retrieval, and reranking catalog describes the jobs this class of model is intended to solve. It does not mean that every listed model shares request fields, input limits, or output formats. Read each model card together with its lifecycle status, origin, and official source. A platform provider may host models from several original vendors, so both the displayed ID and origin matter.

Selection criteria

Embeddings, retrieval, and reranking occupy different stages of a RAG pipeline. Embedding models map queries and documents into a comparable space, while rerankers rescore an existing candidate set; their scores are not interchangeable. Use a domain query set to measure recall, nDCG, Top-K hits, truncation on long documents, multilingual behavior, and terminology. Pin chunking, normalization, distance metric, and index version so that vector dimension or one attractive example does not become the selection criterion.

Integration and protocol boundary

Confirm the exact model ID, authentication method, endpoint, and request schema in the Cohere official documentation, then look for the same exact ID and supported endpoint in this site's model marketplace. Changing only the base URL does not guarantee an identical response contract. Streaming events, asynchronous jobs, file uploads, tool calls, and error payloads may require separate adapters.

Production checklist

Build a repeatable evaluation from representative inputs and record the model ID, parameters, input version, and result. Measure success rate, tail latency, throttling, timeout behavior, cost per completed task, and safety outcomes. Prepare a replacement path for preview or legacy models. The official catalog explains product boundaries; callable availability, channel status, and current pricing remain the responsibility of the model marketplace.

Reviewed provider catalog

Models in this capability

These are provider-catalog models, not a promise of availability on this site. Confirm callable IDs, endpoints, and pricing in the model marketplace.

Embeddings, retrieval, and reranking

Check live availability in the model marketplace

Use cases

  • Semantic search, clustering, and deduplication
  • RAG recall, reranking, and knowledge-base retrieval

FAQ

How should I choose a Cohere Embeddings, retrieval, and reranking model?

Start from the required input and output, then compare quality, latency, context or media limits, tool support, lifecycle status, and total request cost with representative production samples.

Does every model in the official catalog work on this site?

Not necessarily. The documentation records the provider catalog; callable IDs, channel support, endpoints, and current prices must be confirmed in the model marketplace.

What should be tested before production rollout?

Pin the exact model ID and endpoint, test success and error responses, measure quality and latency on real inputs, set timeouts and bounded retries, and monitor provider deprecations.

Official sources

  1. Cohere Chat API v2 Official
  2. Cohere Embed API v2 Official
  3. Cohere Rerank API v2 Official
  4. Cohere Rerank Model Guide Official
  5. Cohere API Keys and Rate Limits Official
  6. Cohere Model and Endpoint Deprecations Official
  7. Migrate Cohere API v1 to v2 Official