文本

创建聊天补全

根据消息生成聊天回答

POST
/v1/chat/completions

请求体

可选参数的支持情况、取值和默认值取决于所选模型。设置采样、推理或工具选项前,请查看其模型详情。

modelstring必填

要使用的模型 ID。可用选项见 Models。

messagesarray必填

组成对话的消息列表。

每个消息对象包含:

  • role (字符串): system, developer, user, assistant, tool, function
  • content (string | array | null): 消息内容

包含 tool_calls 的 assistant 消息可以省略 content,或将其设为 null。

当 content 是数组时,TokenLab 为兼容模型支持结构化多模态块:

  • text: { "type": "text", "text": "..." }
  • 图片: { "type": "image_url", "image_url": { "url": "https://..." } }
  • 视频: { "type": "video_url", "video_url": { "url": "https://..." } }
  • 音频: { "type": "audio_url", "audio_url": { "url": "https://..." } }

多模态输入请使用可公开访问的 https URL。支持的媒体类型取决于所选模型。

temperaturenumber

采样温度。是否支持、允许的取值和默认值取决于所选模型;省略此字段即可使用模型默认值。

max_tokensinteger

要生成的最大 token 数。

streamboolean默认值: false

设为 true 时,通过 SSE 返回消息增量。

stream_optionsobject

流式响应设置。include_usage: true 会在流中返回 token 用量。

top_pnumber

Nucleus 采样参数。建议更改此参数或 temperature,而不是同时更改两者。

frequency_penaltynumber

值在 -2.0 到 2.0 之间。正值会惩罚重复出现的 token。

presence_penaltynumber

值在 -2.0 到 2.0 之间。正值会惩罚已出现在文本中的 token。

stopstring | array

停止序列或序列列表。是否支持及序列数量限制取决于所选模型。

toolsarray

模型可能调用的工具列表(函数调用)。

tool_choicestring | object

控制模型如何使用工具。选项:auto, none, required, 或特定工具对象。

parallel_tool_callsboolean

在所选模型支持时,允许助手在同一轮中发出多个工具调用。

max_completion_tokensinteger

补全的最大 token 数。作为 max_tokens 的替代,对于启用推理的新模型系列更有用。

reasoning_effortstring

支持推理的模型使用的推理强度,允许的取值取决于所选模型。

seedinteger

供支持此功能的模型使用的采样种子,不保证生成完全相同的输出。

ninteger

要生成的补全数量(1-128)。

logprobsboolean

是否返回对数概率(log probabilities)。

top_logprobsinteger

要返回的顶级对数概率数量(0-20)。需要 logprobs: true。

top_kinteger

供支持此功能的模型使用的 Top-K 采样参数。

response_formatobject

响应格式。JSON 模式使用 {"type": "json_object"},JSON Schema 使用 {"type": "json_schema", "json_schema": {...}}。支持情况取决于所选模型。

logit_biasobject

修改指定 token 出现的可能性。将 token ID(以字符串形式)映射到 -100 到 100 的偏差值。

userstring

终端用户标识,可用于安全与滥用防护。

响应

idstring

补全的唯一标识符。

objectstring

始终为 chat.completion。

createdinteger

补全创建时的 Unix 时间戳。

modelstring

用于补全的模型。

choicesarray

补全选项列表。

每个选项包含:

  • index (integer): 选项的索引
  • message (object): 生成的消息
  • finish_reason (string): 模型停止的原因,例如 stop、length 或 tool_calls
usageobject

token 使用统计。

  • prompt_tokens (integer): 提示中的 token 数
  • completion_tokens (integer): 补全中的 token 数
  • total_tokens (integer): 使用的总 token 数

请求

curl -X POST "https://api.tokenlab.sh/v1/chat/completions" \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-terra",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Hello!"}
    ],
    "max_tokens": 1000
  }'

多模态示例

{
  "model": "gemini-2.5-pro",
  "messages": [
    {
      "role": "user",
      "content": [
        { "type": "text", "text": "Describe this video briefly." },
        { "type": "video_url", "video_url": { "url": "https://example.com/demo.mp4" } }
      ]
    }
  ],
  "max_tokens": 64
}

响应

Response
{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "created": 1706000000,
  "model": "gpt-5.6-terra",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! How can I help you today?"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 20,
    "completion_tokens": 9,
    "total_tokens": 29
  }
}

授权

BearerAuth
AuthorizationBearer <token>

API Key 身份验证。在 Dashboard > API > API Keys 中创建或管理 API Key。

位置: header

请求头

X-TokenLab-Delivery-Policy?string

单次请求的交付策略。覆盖 API 密钥和 Workspace 的默认设置。自动优先尝试 TokenLab Verified,并在输出、请求接受或持久资源创建之前,可能会切换一次至仅限 Official。

可选值

  • "auto"
  • "verified"
  • "official"

请求体

application/json

响应

application/json

application/json

application/json

application/json

application/json

application/json

application/json