Text
Create Chat Completion
Creates a completion for the chat message
Request Body
Supported optional parameters, accepted values, and defaults depend on the selected model. Check its model details before setting sampling, reasoning, or tool options.
ID of the model to use. See Models for available options.
A list of messages comprising the conversation.
Each message object contains:
role(string):system,developer,user,assistant,tool,functioncontent(string | array | null): The message content
For an assistant message containing tool_calls, content may be omitted or set to null.
When content is an array, TokenLab supports structured multimodal blocks for compatible models:
- text:
{ "type": "text", "text": "..." } - image:
{ "type": "image_url", "image_url": { "url": "https://..." } } - video:
{ "type": "video_url", "video_url": { "url": "https://..." } } - audio:
{ "type": "audio_url", "audio_url": { "url": "https://..." } }
For multimodal input, use publicly reachable https URLs. Supported media types depend on the selected model.
Sampling temperature. Support, allowed values, and the default depend on the selected model; omit this field to use its default.
Maximum number of tokens to generate.
falseIf true, partial message deltas will be sent as SSE events.
Options for streaming. Set include_usage: true to receive token usage in stream chunks.
Nucleus sampling parameter. We recommend altering this or temperature, not both.
Number between -2.0 and 2.0. Positive values penalize repeated tokens.
Number between -2.0 and 2.0. Positive values penalize tokens already in the text.
Stop sequence or list of sequences. Support and sequence limits depend on the selected model.
A list of tools the model may call (function calling).
Controls how the model uses tools. Options: auto, none, required, or a specific tool object.
Allow multiple tool calls in one assistant turn, when supported by the selected model.
Maximum tokens for the completion. Alternative to max_tokens, useful for newer reasoning-enabled model families.
Reasoning effort for models that support it. Accepted values depend on the selected model.
Sampling seed for models that support it. Identical output is not guaranteed.
Number of completions to generate (1-128).
Whether to return log probabilities.
Number of top log probabilities to return (0-20). Requires logprobs: true.
Top-K sampling for models that support it.
Response format. Use {"type": "json_object"} for JSON mode or {"type": "json_schema", "json_schema": {...}} for a JSON schema. Support depends on the selected model.
Modify the likelihood of specified tokens appearing. Map token IDs (as strings) to bias values from -100 to 100.
A unique identifier representing your end-user for abuse monitoring.
Response
Unique identifier for the completion.
Always chat.completion.
Unix timestamp of when the completion was created.
The model used for completion.
List of completion choices.
Each choice contains:
index(integer): Index of the choicemessage(object): The generated messagefinish_reason(string): Why the model stopped, for examplestop,length, ortool_calls
Token usage statistics.
prompt_tokens(integer): Tokens in the promptcompletion_tokens(integer): Tokens in the completiontotal_tokens(integer): Total tokens used
Request
curl -X POST "https://api.tokenlab.sh/v1/chat/completions" \
-H "Authorization: Bearer sk-your-api-key" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-terra",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"max_tokens": 1000
}'Multimodal Example
{
"model": "gemini-2.5-pro",
"messages": [
{
"role": "user",
"content": [
{ "type": "text", "text": "Describe this video briefly." },
{ "type": "video_url", "video_url": { "url": "https://example.com/demo.mp4" } }
]
}
],
"max_tokens": 64
}Response
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1706000000,
"model": "gpt-5.6-terra",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I help you today?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 20,
"completion_tokens": 9,
"total_tokens": 29
}
}Authorization
BearerAuth API Key authentication. Create or manage API keys in Dashboard > API > API Keys.
In: header
Headers
Per-request Delivery policy. Overrides the API key and Workspace defaults. Auto tries TokenLab Verified first and may switch once to Official only before output, request acceptance, or persistent resource creation.
Value in
- "auto"
- "verified"
- "official"
Request Body
application/json
Response
application/json
application/json
application/json
application/json
application/json
application/json
application/json