TokenLab

文字

重排文件

根據與查詢的相關性對文件進行重排

POST
/v1/rerank

使用語義相似度模型對文件進行重排。適用於優化搜尋結果和 RAG 應用程式。

請求主體

輸入較多時,請將 HTTP 用戶端逾時設為至少 120s;端點會在處理完成後回傳。

modelstring必填

要使用的重排模型 ID(例如:qwen3-vl-rerank)。

querystring必填

用於對文件進行排名的查詢語句。最大長度:32,000 個字元。

documentsarray必填

要進行重排的文件列表(字串)。限制:最多 1,000 份文件;每份文件最多 100,000 個字元;所有文件合計最多 2,000,000 個字元。

top_ninteger

回傳的前幾個結果數量,範圍為 1 到 documents.length。省略時由所選模型決定回傳數量。

return_documentsboolean預設值: false

是否在回應中包含原始文件文本。

回應

resultsarray

帶有評分的已排序文件列表。

每個結果包含:

  • index (integer):原始文件索引
  • relevance_score (number): 相關性評分,範圍取決於所選模型。
  • document (string):原始文本(若 return_documents=true)
modelstring

用於重排的模型。 可用時回傳。

usageobject

Token 使用量統計。 可用時回傳。

請求

curl -X POST "https://api.tokenlab.sh/v1/rerank" \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-vl-rerank",
    "query": "What is machine learning?",
    "documents": [
      "Machine learning is a subset of AI",
      "The weather is nice today",
      "Deep learning uses neural networks"
    ],
    "top_n": 2,
    "return_documents": true
  }'

回應

Response
{
  "results": [
    {
      "index": 0,
      "relevance_score": 0.95,
      "document": "Machine learning is a subset of AI"
    },
    {
      "index": 2,
      "relevance_score": 0.82,
      "document": "Deep learning uses neural networks"
    }
  ],
  "model": "qwen3-vl-rerank",
  "usage": {
    "prompt_tokens": 45,
    "total_tokens": 45
  }
}

授權

BearerAuth
AuthorizationBearer <token>

API Key 驗證。請在 Dashboard > API > API Keys 建立或管理 API 金鑰。

位置: header

請求標頭

X-TokenLab-Delivery-Policy?string

單次請求傳遞策略。會覆寫 API key 與 Workspace 的預設值。系統會自動優先嘗試 TokenLab Verified,並可能在輸出、請求接受或建立持久性資源前切換至 Official 模式一次。

可選值

  • "auto"
  • "verified"
  • "official"

請求主體

application/json

回應

application/json

application/json

application/json

application/json