Text
Rerank Documents
Reranks documents by relevance to a query
Use GET /v1/models?recommended_for=rerank to get current rerank models.
Rerank documents using semantic similarity models. Useful for improving search results and RAG applications.
Request Body
This endpoint returns after ranking finishes. For large document sets, set your HTTP client timeout to at least 120s.
ID of the reranker model to use (e.g., qwen3-vl-rerank).
The query to rank documents against. Maximum length: 32,000 characters.
List of documents (strings) to rerank. Limits: up to 1,000 documents, each document up to 100,000 characters, and at most 2,000,000 total document characters.
Number of top results to return, from 1 to documents.length. If omitted, the selected model determines the number of results.
falseWhether to include original document text in response.
Response
Ranked list of documents with scores.
Each result contains:
index(integer): Original document indexrelevance_score(number): Relevance score. Its range depends on the selected model.document(string): Original text (ifreturn_documents=true)
The model used for reranking. When available.
Token usage statistics. When available.
Request
curl -X POST "https://api.tokenlab.sh/v1/rerank" \
-H "Authorization: Bearer sk-your-api-key" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-rerank",
"query": "What is machine learning?",
"documents": [
"Machine learning is a subset of AI",
"The weather is nice today",
"Deep learning uses neural networks"
],
"top_n": 2,
"return_documents": true
}'Response
{
"results": [
{
"index": 0,
"relevance_score": 0.95,
"document": "Machine learning is a subset of AI"
},
{
"index": 2,
"relevance_score": 0.82,
"document": "Deep learning uses neural networks"
}
],
"model": "qwen3-vl-rerank",
"usage": {
"prompt_tokens": 45,
"total_tokens": 45
}
}Authorization
BearerAuth API Key authentication. Create or manage API keys in Dashboard > API > API Keys.
In: header
Headers
Per-request Delivery policy. Overrides the API key and Workspace defaults. Auto tries TokenLab Verified first and may switch once to Official only before output, request acceptance, or persistent resource creation.
Value in
- "auto"
- "verified"
- "official"
Request Body
application/json
Response
application/json
application/json
application/json
application/json