协议端点决定载荷架构
TokenLab 不会在运行时使用动态格式提示标头(例如专有的 format-hint 标签)来指示响应架构。相反,载荷结构严格由所调用的端点决定。解析客户端响应需要将请求路由到目标原生协议端点,而不是通过检查响应标头来确定载荷类型:
- Chat Completions (
/v1/chat/completions):使用 OpenAI 兼容的架构,返回choices、message.content和usage块(prompt_tokens、completion_tokens、total_tokens)。 - Responses (
/v1/responses):遵循 OpenAI Responses API 格式,适用于后台任务、服务端工具和响应事件。 - Anthropic Messages (
/v1/messages):使用原生 Anthropic 架构(content块、thinking和output_tokens)与 Anthropic Claude 模型进行交互。配置 Anthropic SDK 时,请将基础 URL 设置为https://api.tokenlab.sh,且不带/v1前缀。 - Gemini (
/v1beta/models/:model:generateContent):接受原生 Gemini 架构(contents、parts),并返回标准 Gemini REST 候选对象。
在路由模型请求之前,请通过调用 Get a Model (GET /v1/models/{model}) 或查看 Models catalog 确认其接受的协议。检查响应中的 tokenlab.accepted_request_formats 列表。有关完整的端点映射规则,请参阅 API Formats guide。
记载的请求标头
对 TokenLab 端点的所有标准调用都需要特定的 HTTP 请求标头:
Authorization:以 Bearer 令牌形式传递凭据(Authorization: Bearer $TOKENLAB_API_KEY)。管理端点需要管理令牌(Authorization: Bearer mt-...)。Content-Type:对于包含 JSON 主体的 POST 请求,必须为application/json。
记载的响应标头
TokenLab 返回标准和自定义 HTTP 标头,用于速率限制、计费对账以及异步任务管理:
速率限制标头
当请求超出账户等级限制时,TokenLab 会返回 HTTP 429 rate_limit_exceeded 状态,并附带两个标头:
Retry-After:指定重试调用前需要等待的时间(以秒为单位)。X-RateLimit-Limit:报告您当前认证等级的每分钟请求数限制。
务必使用 Retry-After 标头值来处理重试,而不是硬编码退避限制。有关恢复处理的更多详细信息,请参阅 Rate Limits guide。
计费与可观测性标头
对于非流式和异步交互,TokenLab 提供了标识标头以追踪扣费和后台作业:
X-Billing-Transaction-ID:在 HTTP 响应发送前完成计费结算时返回。非流式 OpenAI 兼容端点会在 JSON 主体中包含billing_transaction_id,但 Gemini 和原生格式端点会通过此标头公开它。流式调用可能会在连接关闭后完成结算;如果不存在,请从工作区用量记录中检索该 ID。有关结算工作流,请查阅 Billing and Pricing guide。X-Task-ID:在为视频、音乐、3D 或基于任务的图像生成创建异步作业时,在响应标头中返回。它提供了与任务id对应的标头级关联 ID。有关日志记录规范,请参阅 Logs and Troubleshooting guide。
实现:捕获标头并在遇到 429 时重试
以下 Python 示例说明了如何向 Chat Completions 端点提交请求、检查交易标识符,以及在遭遇速率限制时处理 Retry-After 标头:
import os
import time
import requests
API_KEY = os.environ["TOKENLAB_API_KEY"]
ENDPOINT = "https://api.tokenlab.sh/v1/chat/completions"
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
payload = {
"model": "gpt-5.6-terra",
"messages": [{"role": "user", "content": "Summarize system status."}]
}
max_attempts = 3
for attempt in range(max_attempts):
response = requests.post(ENDPOINT, headers=headers, json=payload, timeout=30)
if response.status_code == 200:
# Check for billing transaction header on settled non-streaming calls
billing_id = response.headers.get("X-Billing-Transaction-ID")
data = response.json()
print(f"Settled Transaction ID: {billing_id}")
print(data["choices"][0]["message"]["content"])
break
elif response.status_code == 429:
retry_after = response.headers.get("Retry-After")
limit = response.headers.get("X-RateLimit-Limit")
wait_seconds = float(retry_after) if retry_after else 2 ** attempt
print(f"Rate limit reached ({limit} req/min). Retrying in {wait_seconds}s...")
time.sleep(wait_seconds)
else:
response.raise_for_status()
日志记录与可观测性实践
在配置请求监控时,请记录标头和载荷中返回的公共跟踪标识符以便对账,而无需保留用户提示词或凭据:
- 将
request_id、X-Billing-Transaction-ID和X-Task-ID与状态码和响应延迟一同保留。 - 始终从遥测管道中脱敏
Authorization标头、原始 API 密钥以及私有签名 URL。 - 如需进行服务端财务对账,请查询
GET /v1/management/api-keys/{keyId}/usage,而不是抓取仪表板页面或仅凭原始 Token 计数器估算总量。
来源
- https://docs.tokenlab.sh/api-reference/models/get-model资料更新于 2026-09-27
- https://docs.tokenlab.sh/guides/api-formats资料更新于 2026-09-27
- https://docs.tokenlab.sh/guides/rate-limits资料更新于 2026-09-27
- https://docs.tokenlab.sh/guides/billing资料更新于 2026-09-27
- https://docs.tokenlab.sh/guides/observability-troubleshooting资料更新于 2026-09-27



