Core Guides
API formats
Choose Chat Completions, Responses, Messages, or Gemini
One TokenLab API key works with four common API formats. Keep the format your application already uses unless you need a model-specific feature.
| Format | Endpoint | Choose it when |
|---|---|---|
| Chat Completions | /v1/chat/completions | You use an OpenAI-compatible chat client or want broad model compatibility |
| Responses | /v1/responses | Your application uses Responses events, background responses, or Responses tools |
| Anthropic Messages | /v1/messages | You need Claude Messages fields such as thinking or Claude tool use |
| Gemini | /v1beta/models/:model:generateContent | You already use Gemini contents, parts, files, cache, or Gemini tools |
Check tokenlab.accepted_request_formats on the model page or in GET /v1/models/{model}. A model does not necessarily accept all four formats.
Chat Completions
Chat Completions is the easiest choice for existing OpenAI-compatible clients and portable chat features.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["TOKENLAB_API_KEY"],
base_url="https://api.tokenlab.sh/v1",
)
response = client.chat.completions.create(
model="gpt-5.6-luna",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)Use this format for:
- existing OpenAI-compatible chat integrations
- standard message history and streaming
- function calling supported by the selected model
Fields unique to another API format are not guaranteed to work here.
Responses
Use Responses only when the selected model lists openai_responses in accepted_request_formats.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["TOKENLAB_API_KEY"],
base_url="https://api.tokenlab.sh/v1",
)
response = client.responses.create(
model="gpt-5.6-terra",
input="Explain why the sky is blue in two sentences.",
)
print(response.output_text)The Responses API supports create, retrieve, compact, delete, streaming, WebSocket creation and continuation, and background responses where the selected model supports them. Delete removes a stored response; it does not cancel an active response.
Anthropic Messages
Use the Anthropic SDK with the TokenLab host, without adding /v1 to the base URL.
import os
from anthropic import Anthropic
client = Anthropic(
api_key=os.environ["TOKENLAB_API_KEY"],
base_url="https://api.tokenlab.sh",
)
message = client.messages.create(
model="claude-sonnet-5",
max_tokens=512,
messages=[{"role": "user", "content": "Hello!"}],
)
print(message.content[0].text)Choose Messages for Claude-specific request and response fields. Keep Claude tool calls, thinking blocks, and prompt-cache fields in this format.
Gemini
Gemini requests use the TokenLab host and the Gemini REST path.
curl "https://api.tokenlab.sh/v1beta/models/gemini-3.5-flash:generateContent" \
-H "Authorization: Bearer $TOKENLAB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"contents": [{"parts": [{"text": "Hello!"}]}]
}'Choose Gemini when your application already uses Gemini content parts, function declarations, files, cached content, or other Gemini fields. Both lowerCamelCase ProtoJSON names and original snake_case proto names are accepted; avoid sending both spellings of the same field in one request.
Do not mix formats in one conversation
Messages, Responses, Chat Completions, and Gemini represent conversation state and tool calls differently. Keep one format for a conversation. If you must migrate stored history, convert and validate it in your application before sending the next request.
Unknown fields
TokenLab may pass through fields it does not recognize, but that does not mean every model accepts them. A feature is safe to depend on only when the selected model documents it and your application handles an unsupported-field error.