Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front. See what's new

Cohere: rerank-v4.0-fast

Cohere Rerank 4 Fast orders retrieved documents by relevance to a query, with a focus on latency and throughput. It supports multilingual text retrieval and is billed by reported search units, making it useful between retrieval and answer generation in a RAG pipeline.
Compare models
rerank-v4.0-fast
AvailableCohereRerankSync
Modalities
Chat

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    RerankingTokenLab endpoint
    POST/v1/rerank
    curl -X POST "https://api.tokenlab.sh/v1/rerank" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "operation": "rerank",
      "model": "rerank-v4.0-fast",
      "query": "What is the capital of France?",
      "documents": [
        "Paris is the capital of France.",
        "Berlin is the capital of Germany."
      ]
    }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Reranking

Per request
Official price
$0.002
TokenLab price
-
Discount
-

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours

No data yet

Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Cohere Rerank 4 Fast in Console with a prompt ready to edit or send.

Help me try rerank-v4.0-fast with rerank at /v1/rerank. Show the result and cost.

Use cases

01

Search quality

Rerank retrieval candidates and lift top-result precision for search and RAG.

02

Side-by-side test

Compare response quality, latency, and price side by side.

Prompt examples

Rerank these search results for the query: best image generation model for product photos.

Send the smallest possible request and show me the result and cost.

Compare this model with a cheaper one in the same category — when is each worth it?

FAQ

How much does Cohere Rerank 4 Fast cost?

On TokenLab, Cohere Rerank 4 Fast costs — Per request. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

How do I test Cohere Rerank 4 Fast?

Open Cohere Rerank 4 Fast in Create. A sample for /v1/rerank will be ready to try.

Which endpoint should Cohere Rerank 4 Fast use?

Use https://api.tokenlab.sh/v1/rerank for Cohere Rerank 4 Fast. The request example below shows the matching code shape.

Can I test Cohere Rerank 4 Fast before integrating it?

Yes. Open Console starts a ready draft for Cohere Rerank 4 Fast and keeps your prompt after sign-in, so you don’t lose context.

Which operations does Cohere Rerank 4 Fast support?

Cohere Rerank 4 Fast supports Reranking. Select an operation above to see its endpoint and request example.

More from Cohere Rerank

Related models