文本

重排序文档

根据与查询的相关性对文档进行重排序

POST
/v1/rerank

根据查询内容重新排列文档,常用于搜索和 RAG。

请求体

输入较多时,请将 HTTP 客户端超时设为至少 120s;接口会在处理完成后返回。

modelstring必填

要使用的重排序模型 ID(例如 qwen3-vl-rerank)。

querystring必填

用于对文档进行排序的查询语句。最大长度:32,000 个字符。

documentsarray必填

需要重排序的文档列表(字符串数组)。限制:最多 1,000 篇文档;每篇最多 100,000 个字符;所有文档合计最多 2,000,000 个字符。

top_ninteger

返回的前若干个结果,范围为 1 到 documents.length。省略时由所选模型决定返回数量。

return_documentsboolean默认值: false

是否在响应中包含原始文档文本。

响应

resultsarray

已按相关性排列的结果。

每个结果包含:

  • index (integer): 原始文档索引
  • relevance_score (number): 相关性评分,取值范围由所选模型决定。
  • document (string): 原始文本(如果 return_documents=true)
modelstring

用于重排序的模型。 可用时返回。

usageobject

Token 使用统计。 可用时返回。

请求

curl -X POST "https://api.tokenlab.sh/v1/rerank" \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-vl-rerank",
    "query": "What is machine learning?",
    "documents": [
      "Machine learning is a subset of AI",
      "The weather is nice today",
      "Deep learning uses neural networks"
    ],
    "top_n": 2,
    "return_documents": true
  }'

响应

Response
{
  "results": [
    {
      "index": 0,
      "relevance_score": 0.95,
      "document": "Machine learning is a subset of AI"
    },
    {
      "index": 2,
      "relevance_score": 0.82,
      "document": "Deep learning uses neural networks"
    }
  ],
  "model": "qwen3-vl-rerank",
  "usage": {
    "prompt_tokens": 45,
    "total_tokens": 45
  }
}

授权

BearerAuth
AuthorizationBearer <token>

API Key 身份验证。在 Dashboard > API > API Keys 中创建或管理 API Key。

位置: header

请求头

X-TokenLab-Delivery-Policy?string

单次请求的交付策略。覆盖 API 密钥和 Workspace 的默认设置。自动优先尝试 TokenLab Verified,并在输出、请求接受或持久资源创建之前,可能会切换一次至仅限 Official。

可选值

  • "auto"
  • "verified"
  • "official"

请求体

application/json

响应

application/json

application/json

application/json

application/json