Protocol Endpoints Determine Payload Schemas
TokenLab does not use dynamic format hint headers (such as proprietary format-hint tags) to indicate response schemas at runtime. Instead, payload structures are governed strictly by the endpoint called. Parsing client responses requires routing requests to the target native protocol endpoint rather than inspecting response headers for payload types:
- Chat Completions (
/v1/chat/completions): Uses OpenAI-compatible schemas returningchoices,message.content, and ausageblock (prompt_tokens,completion_tokens,total_tokens). - Responses (
/v1/responses): Adheres to the OpenAI Responses API format for background tasks, server tools, and response events. - Anthropic Messages (
/v1/messages): Interacts with Anthropic Claude models using the native Anthropic schema (contentblocks,thinking, andoutput_tokens). When configuring the Anthropic SDK, set the base URL tohttps://api.tokenlab.shwithout the/v1prefix. - Gemini (
/v1beta/models/:model:generateContent): Accepts native Gemini schemas (contents, parts) and returns standard Gemini REST candidate objects.
Before routing a model request, verify which protocols it accepts by calling Get a Model (GET /v1/models/{model}) or reviewing the Models catalog. Inspect the tokenlab.accepted_request_formats list in the response. Consult the API Formats guide for comprehensive endpoint mapping rules.
Documented Request Headers
All standard calls to TokenLab endpoints require specific HTTP request headers:
Authorization: Passes credentials as a bearer token (Authorization: Bearer $TOKENLAB_API_KEY). Management endpoints require a management token (Authorization: Bearer mt-...).Content-Type: Must beapplication/jsonfor POST requests containing JSON bodies.
Documented Response Headers
TokenLab returns standard and custom HTTP headers for rate limits, billing reconciliation, and asynchronous task management:
Rate Limiting Headers
When a request exceeds account tier limits, TokenLab returns an HTTP 429 rate_limit_exceeded status accompanied by two headers:
Retry-After: Specifies the required waiting period in seconds before retrying the call.X-RateLimit-Limit: Reports your active requests-per-minute limit for the authenticated tier.
Always use the Retry-After header value to handle retries rather than hardcoding backoff limits. More details on recovery handling appear in the Rate Limits guide.
Billing and Observability Headers
For non-streaming and asynchronous interactions, TokenLab provides identification headers to trace charges and background work:
X-Billing-Transaction-ID: Returned when billing settles before the HTTP response is dispatched. Non-streaming OpenAI-compatible endpoints includebilling_transaction_idin the JSON body, but Gemini and native format endpoints expose it through this header. Streaming calls may settle after the connection closes; when absent, retrieve the ID from workspace usage records. Review settlement workflows in the Billing and Pricing guide.X-Task-ID: Returned on response headers when creating asynchronous jobs for video, music, 3D, or task-based image generation. It provides a header-level correlation ID corresponding to the taskid. Consult the Logs and Troubleshooting guide for logging standards.
Implementation: Capturing Headers and Retrying on 429
The following Python example illustrates how to submit a request to the Chat Completions endpoint, inspect transaction identifiers, and handle Retry-After headers during rate limits:
import os
import time
import requests
API_KEY = os.environ["TOKENLAB_API_KEY"]
ENDPOINT = "https://api.tokenlab.sh/v1/chat/completions"
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
payload = {
"model": "gpt-5.6-terra",
"messages": [{"role": "user", "content": "Summarize system status."}]
}
max_attempts = 3
for attempt in range(max_attempts):
response = requests.post(ENDPOINT, headers=headers, json=payload, timeout=30)
if response.status_code == 200:
# Check for billing transaction header on settled non-streaming calls
billing_id = response.headers.get("X-Billing-Transaction-ID")
data = response.json()
print(f"Settled Transaction ID: {billing_id}")
print(data["choices"][0]["message"]["content"])
break
elif response.status_code == 429:
retry_after = response.headers.get("Retry-After")
limit = response.headers.get("X-RateLimit-Limit")
wait_seconds = float(retry_after) if retry_after else 2 ** attempt
print(f"Rate limit reached ({limit} req/min). Retrying in {wait_seconds}s...")
time.sleep(wait_seconds)
else:
response.raise_for_status()
Logging and Observability Practices
When instrumenting request monitoring, log public tracking identifiers returned in headers and payloads to reconcile records without persisting user prompts or credentials:
- Retain
request_id,X-Billing-Transaction-ID, andX-Task-IDalongside status codes and response latencies. - Always redact
Authorizationheaders, raw API keys, and private signed URLs from telemetry pipelines. - For server-side financial reconciliation, query
GET /v1/management/api-keys/{keyId}/usageinstead of scraping dashboard pages or estimating totals from raw token counters alone.
Sources
- https://docs.tokenlab.sh/api-reference/models/get-modelSources checked 2026-09-27
- https://docs.tokenlab.sh/guides/api-formatsSources checked 2026-09-27
- https://docs.tokenlab.sh/guides/rate-limitsSources checked 2026-09-27
- https://docs.tokenlab.sh/guides/billingSources checked 2026-09-27
- https://docs.tokenlab.sh/guides/observability-troubleshootingSources checked 2026-09-27



