TokenLab

Text

Create Message

Creates a message using the Anthropic Messages API format

POST
/v1/messages

Overview

This endpoint provides native Anthropic Messages API compatibility. Use this for Claude models with features like extended thinking.

This endpoint keeps the native Anthropic contract. messages must be an array of user / assistant messages, system belongs in the top-level system field, and max_tokens is required. If your payload uses OpenAI roles such as system, developer, or tool inside messages, send it to /v1/chat/completions instead.

Base URL for Anthropic SDK: https://api.tokenlab.sh (no /v1 suffix)

Request Headers

x-api-keystringheader

Your TokenLab API key. Send either x-api-key or Authorization: Bearer <API_KEY>.

anthropic-versionstringheaderrequired

Anthropic API version. Use 2023-06-01.

Request Body

Supported optional parameters, accepted values, and defaults depend on the selected model. Check its model details before setting sampling, reasoning, or tool options.

modelstringrequired

ID of a model that supports the anthropic_messages format. Check its model details for current capabilities.

messagesarrayrequired

Array of message objects with role and content.

For Claude models with vision support, content can be either a plain string or an array of content blocks. To send images, use structured content blocks rather than placing image URLs or Base64 strings directly into plain text.

Example content blocks:

  • text block: { "type": "text", "text": "Describe this image" }
  • image block via URL: { "type": "image", "source": { "type": "url", "url": "https://example.com/image.jpg" } }
  • image block via Base64: { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0KGgoAAA..." } }
max_tokensintegerrequired

Maximum tokens to generate.

systemstring | array

System prompt as a string or an array of content blocks, separate from the messages array.

temperaturenumber

Sampling temperature. Support, allowed values, and the default depend on the selected model; omit this field to use its default.

streambooleandefault: false

Enable streaming responses.

thinkingobject

Native thinking configuration. Supported modes, budgets, and combinations with other parameters depend on the selected model.

toolsarray

Available tools for the model.

tool_choiceobject

Native tool-choice object, for example {"type":"auto"}. Supported types depend on the selected model.

top_pnumber

Nucleus sampling parameter. Use either temperature or top_p, not both.

top_kinteger

Only sample from the top K options for each token.

stop_sequencesarray

Custom stop sequences that will cause the model to stop generating.

metadataobject

Metadata to attach to the request for tracking purposes.

Response

idstring

Unique message identifier.

typestring

Always message.

rolestring

Always assistant.

contentarray

Array of content blocks, such as text, thinking, or tool_use.

modelstring

Model used.

stop_reasonstring

Reason generation stopped, such as end_turn, max_tokens, stop_sequence, or tool_use.

usageobject

Token usage with input_tokens and output_tokens.

Request

curl -X POST "https://api.tokenlab.sh/v1/messages" \
  -H "x-api-key: sk-your-api-key" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5",
    "max_tokens": 1024,
    "system": "You are a helpful assistant.",
    "messages": [
      {"role": "user", "content": "Hello, Claude!"}
    ]
  }'

Response

Response
{
  "id": "msg_abc123",
  "type": "message",
  "role": "assistant",
  "content": [
    {
      "type": "text",
      "text": "Hello! How can I help you today?"
    }
  ],
  "model": "claude-sonnet-5",
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 15,
    "output_tokens": 10
  }
}

Vision Input Example

For Claude models with vision support, place images inside messages[].content as structured image blocks.

{
  "model": "claude-sonnet-5",
  "max_tokens": 1024,
  "messages": [
    {
      "role": "user",
      "content": [
        {
          "type": "text",
          "text": "Please describe this image."
        },
        {
          "type": "image",
          "source": {
            "type": "url",
            "url": "https://example.com/demo.jpg"
          }
        }
      ]
    }
  ]
}
{
  "model": "claude-sonnet-5",
  "max_tokens": 1024,
  "messages": [
    {
      "role": "user",
      "content": [
        {
          "type": "text",
          "text": "Please describe this image."
        },
        {
          "type": "image",
          "source": {
            "type": "base64",
            "media_type": "image/jpeg",
            "data": "/9j/4AAQSkZJRgABAQ..."
          }
        }
      ]
    }
  ]
}

Extended Thinking Example

Set TOKENLAB_THINKING_MODEL to a model that supports manual thinking with type: "enabled" and budget_tokens. Adjust both token budgets to that model's limits; other models may require a different thinking configuration.

import os
from anthropic import Anthropic

client = Anthropic(
    api_key="sk-your-api-key",
    base_url="https://api.tokenlab.sh"
)

message = client.messages.create(
    model=os.environ["TOKENLAB_THINKING_MODEL"],
    max_tokens=16000,
    thinking={
        "type": "enabled",
        "budget_tokens": 10000
    },
    messages=[{"role": "user", "content": "Solve this math problem..."}]
)

for block in message.content:
    if block.type == "thinking":
        print(f"Thinking: {block.thinking}")
    elif block.type == "text":
        print(f"Response: {block.text}")

Anthropic Message Batches

TokenLab now exposes the native Anthropic Message Batches flow alongside /v1/messages.

Available routes:

  • POST /v1/messages/batches
  • GET /v1/messages/batches
  • GET /v1/messages/batches/:message_batch_id
  • GET /v1/messages/batches/:message_batch_id/results
  • POST /v1/messages/batches/:message_batch_id/cancel
  • DELETE /v1/messages/batches/:message_batch_id

Operational notes:

  • Use the same TokenLab API key plus Anthropic-native headers.
  • If batch items reference file_id, also include anthropic-beta: files-api-2025-04-14.
  • Batch jobs keep Anthropic-native request/response shapes while TokenLab tracks billing status for reconciliation.

Authorization

BearerAuth
AuthorizationBearer <token>

API Key authentication. Create or manage API keys in Dashboard > API > API Keys.

In: header

Headers

X-TokenLab-Delivery-Policy?string

Per-request Delivery policy. Overrides the API key and Workspace defaults. Auto tries TokenLab Verified first and may switch once to Official only before output, request acceptance, or persistent resource creation.

Value in

  • "auto"
  • "verified"
  • "official"

Request Body

application/json

Response

application/json

application/json

application/json

application/json

application/json

application/json