Text
Create Message
Creates a message using the Anthropic Messages API format
Overview
This endpoint provides native Anthropic Messages API compatibility. Use this for Claude models with features like extended thinking.
This endpoint keeps the native Anthropic contract. messages must be an array of user / assistant messages, system belongs in the top-level system field, and max_tokens is required. If your payload uses OpenAI roles such as system, developer, or tool inside messages, send it to /v1/chat/completions instead.
Base URL for Anthropic SDK: https://api.tokenlab.sh (no /v1 suffix)
Request Headers
Your TokenLab API key. Send either x-api-key or Authorization: Bearer <API_KEY>.
Anthropic API version. Use 2023-06-01.
Request Body
Supported optional parameters, accepted values, and defaults depend on the selected model. Check its model details before setting sampling, reasoning, or tool options.
ID of a model that supports the anthropic_messages format. Check its model details for current capabilities.
Array of message objects with role and content.
For Claude models with vision support, content can be either a plain string or an array of content blocks. To send images, use structured content blocks rather than placing image URLs or Base64 strings directly into plain text.
Example content blocks:
- text block:
{ "type": "text", "text": "Describe this image" } - image block via URL:
{ "type": "image", "source": { "type": "url", "url": "https://example.com/image.jpg" } } - image block via Base64:
{ "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0KGgoAAA..." } }
Maximum tokens to generate.
System prompt as a string or an array of content blocks, separate from the messages array.
Sampling temperature. Support, allowed values, and the default depend on the selected model; omit this field to use its default.
falseEnable streaming responses.
Native thinking configuration. Supported modes, budgets, and combinations with other parameters depend on the selected model.
Available tools for the model.
Native tool-choice object, for example {"type":"auto"}. Supported types depend on the selected model.
Nucleus sampling parameter. Use either temperature or top_p, not both.
Only sample from the top K options for each token.
Custom stop sequences that will cause the model to stop generating.
Metadata to attach to the request for tracking purposes.
Response
Unique message identifier.
Always message.
Always assistant.
Array of content blocks, such as text, thinking, or tool_use.
Model used.
Reason generation stopped, such as end_turn, max_tokens, stop_sequence, or tool_use.
Token usage with input_tokens and output_tokens.
Request
curl -X POST "https://api.tokenlab.sh/v1/messages" \
-H "x-api-key: sk-your-api-key" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"system": "You are a helpful assistant.",
"messages": [
{"role": "user", "content": "Hello, Claude!"}
]
}'Response
{
"id": "msg_abc123",
"type": "message",
"role": "assistant",
"content": [
{
"type": "text",
"text": "Hello! How can I help you today?"
}
],
"model": "claude-sonnet-5",
"stop_reason": "end_turn",
"usage": {
"input_tokens": 15,
"output_tokens": 10
}
}Vision Input Example
For Claude models with vision support, place images inside messages[].content as structured image blocks.
{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Please describe this image."
},
{
"type": "image",
"source": {
"type": "url",
"url": "https://example.com/demo.jpg"
}
}
]
}
]
}{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Please describe this image."
},
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/jpeg",
"data": "/9j/4AAQSkZJRgABAQ..."
}
}
]
}
]
}Extended Thinking Example
Set TOKENLAB_THINKING_MODEL to a model that supports manual thinking with type: "enabled" and budget_tokens. Adjust both token budgets to that model's limits; other models may require a different thinking configuration.
import os
from anthropic import Anthropic
client = Anthropic(
api_key="sk-your-api-key",
base_url="https://api.tokenlab.sh"
)
message = client.messages.create(
model=os.environ["TOKENLAB_THINKING_MODEL"],
max_tokens=16000,
thinking={
"type": "enabled",
"budget_tokens": 10000
},
messages=[{"role": "user", "content": "Solve this math problem..."}]
)
for block in message.content:
if block.type == "thinking":
print(f"Thinking: {block.thinking}")
elif block.type == "text":
print(f"Response: {block.text}")Anthropic Message Batches
TokenLab now exposes the native Anthropic Message Batches flow alongside /v1/messages.
Available routes:
POST /v1/messages/batchesGET /v1/messages/batchesGET /v1/messages/batches/:message_batch_idGET /v1/messages/batches/:message_batch_id/resultsPOST /v1/messages/batches/:message_batch_id/cancelDELETE /v1/messages/batches/:message_batch_id
Operational notes:
- Use the same TokenLab API key plus Anthropic-native headers.
- If batch items reference
file_id, also includeanthropic-beta: files-api-2025-04-14. - Batch jobs keep Anthropic-native request/response shapes while TokenLab tracks billing status for reconciliation.
Authorization
BearerAuth API Key authentication. Create or manage API keys in Dashboard > API > API Keys.
In: header
Headers
Per-request Delivery policy. Overrides the API key and Workspace defaults. Auto tries TokenLab Verified first and may switch once to Official only before output, request acceptance, or persistent resource creation.
Value in
- "auto"
- "verified"
- "official"
Request Body
application/json
Response
application/json
application/json
application/json
application/json
application/json
application/json