TokenLab

Text

Create Chat Completion

Creates a completion for the chat message

POST
/v1/chat/completions

Request Body

Supported optional parameters, accepted values, and defaults depend on the selected model. Check its model details before setting sampling, reasoning, or tool options.

modelstringrequired

ID of the model to use. See Models for available options.

messagesarrayrequired

A list of messages comprising the conversation.

Each message object contains:

  • role (string): system, developer, user, assistant, tool, function
  • content (string | array | null): The message content

For an assistant message containing tool_calls, content may be omitted or set to null.

When content is an array, TokenLab supports structured multimodal blocks for compatible models:

  • text: { "type": "text", "text": "..." }
  • image: { "type": "image_url", "image_url": { "url": "https://..." } }
  • video: { "type": "video_url", "video_url": { "url": "https://..." } }
  • audio: { "type": "audio_url", "audio_url": { "url": "https://..." } }

For multimodal input, use publicly reachable https URLs. Supported media types depend on the selected model.

temperaturenumber

Sampling temperature. Support, allowed values, and the default depend on the selected model; omit this field to use its default.

max_tokensinteger

Maximum number of tokens to generate.

streambooleandefault: false

If true, partial message deltas will be sent as SSE events.

stream_optionsobject

Options for streaming. Set include_usage: true to receive token usage in stream chunks.

top_pnumber

Nucleus sampling parameter. We recommend altering this or temperature, not both.

frequency_penaltynumber

Number between -2.0 and 2.0. Positive values penalize repeated tokens.

presence_penaltynumber

Number between -2.0 and 2.0. Positive values penalize tokens already in the text.

stopstring | array

Stop sequence or list of sequences. Support and sequence limits depend on the selected model.

toolsarray

A list of tools the model may call (function calling).

tool_choicestring | object

Controls how the model uses tools. Options: auto, none, required, or a specific tool object.

parallel_tool_callsboolean

Allow multiple tool calls in one assistant turn, when supported by the selected model.

max_completion_tokensinteger

Maximum tokens for the completion. Alternative to max_tokens, useful for newer reasoning-enabled model families.

reasoning_effortstring

Reasoning effort for models that support it. Accepted values depend on the selected model.

seedinteger

Sampling seed for models that support it. Identical output is not guaranteed.

ninteger

Number of completions to generate (1-128).

logprobsboolean

Whether to return log probabilities.

top_logprobsinteger

Number of top log probabilities to return (0-20). Requires logprobs: true.

top_kinteger

Top-K sampling for models that support it.

response_formatobject

Response format. Use {"type": "json_object"} for JSON mode or {"type": "json_schema", "json_schema": {...}} for a JSON schema. Support depends on the selected model.

logit_biasobject

Modify the likelihood of specified tokens appearing. Map token IDs (as strings) to bias values from -100 to 100.

userstring

A unique identifier representing your end-user for abuse monitoring.

Response

idstring

Unique identifier for the completion.

objectstring

Always chat.completion.

createdinteger

Unix timestamp of when the completion was created.

modelstring

The model used for completion.

choicesarray

List of completion choices.

Each choice contains:

  • index (integer): Index of the choice
  • message (object): The generated message
  • finish_reason (string): Why the model stopped, for example stop, length, or tool_calls
usageobject

Token usage statistics.

  • prompt_tokens (integer): Tokens in the prompt
  • completion_tokens (integer): Tokens in the completion
  • total_tokens (integer): Total tokens used

Request

curl -X POST "https://api.tokenlab.sh/v1/chat/completions" \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-terra",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Hello!"}
    ],
    "max_tokens": 1000
  }'

Multimodal Example

{
  "model": "gemini-2.5-pro",
  "messages": [
    {
      "role": "user",
      "content": [
        { "type": "text", "text": "Describe this video briefly." },
        { "type": "video_url", "video_url": { "url": "https://example.com/demo.mp4" } }
      ]
    }
  ],
  "max_tokens": 64
}

Response

Response
{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "created": 1706000000,
  "model": "gpt-5.6-terra",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! How can I help you today?"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 20,
    "completion_tokens": 9,
    "total_tokens": 29
  }
}

Authorization

BearerAuth
AuthorizationBearer <token>

API Key authentication. Create or manage API keys in Dashboard > API > API Keys.

In: header

Headers

X-TokenLab-Delivery-Policy?string

Per-request Delivery policy. Overrides the API key and Workspace defaults. Auto tries TokenLab Verified first and may switch once to Official only before output, request acceptance, or persistent resource creation.

Value in

  • "auto"
  • "verified"
  • "official"

Request Body

application/json

Response

application/json

application/json

application/json

application/json

application/json

application/json

application/json