Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Understanding TokenLab HTTP Headers and Native Protocol Endpoints

·September 19, 2026·4 min read·Updated September 26, 2026·1320 views
#feature#api-formats#developer-experience#agents
Understanding TokenLab HTTP Headers and Native Protocol Endpoints

Protocol Endpoints Determine Payload Schemas

TokenLab does not use dynamic format hint headers (such as proprietary format-hint tags) to indicate response schemas at runtime. Instead, payload structures are governed strictly by the endpoint called. Parsing client responses requires routing requests to the target native protocol endpoint rather than inspecting response headers for payload types:

  • Chat Completions (/v1/chat/completions): Uses OpenAI-compatible schemas returning choices, message.content, and a usage block (prompt_tokens, completion_tokens, total_tokens).
  • Responses (/v1/responses): Adheres to the OpenAI Responses API format for background tasks, server tools, and response events.
  • Anthropic Messages (/v1/messages): Interacts with Anthropic Claude models using the native Anthropic schema (content blocks, thinking, and output_tokens). When configuring the Anthropic SDK, set the base URL to https://api.tokenlab.sh without the /v1 prefix.
  • Gemini (/v1beta/models/:model:generateContent): Accepts native Gemini schemas (contents, parts) and returns standard Gemini REST candidate objects.

Before routing a model request, verify which protocols it accepts by calling Get a Model (GET /v1/models/{model}) or reviewing the Models catalog. Inspect the tokenlab.accepted_request_formats list in the response. Consult the API Formats guide for comprehensive endpoint mapping rules.

Documented Request Headers

All standard calls to TokenLab endpoints require specific HTTP request headers:

  • Authorization: Passes credentials as a bearer token (Authorization: Bearer $TOKENLAB_API_KEY). Management endpoints require a management token (Authorization: Bearer mt-...).
  • Content-Type: Must be application/json for POST requests containing JSON bodies.

Documented Response Headers

TokenLab returns standard and custom HTTP headers for rate limits, billing reconciliation, and asynchronous task management:

Rate Limiting Headers

When a request exceeds account tier limits, TokenLab returns an HTTP 429 rate_limit_exceeded status accompanied by two headers:

  • Retry-After: Specifies the required waiting period in seconds before retrying the call.
  • X-RateLimit-Limit: Reports your active requests-per-minute limit for the authenticated tier.

Always use the Retry-After header value to handle retries rather than hardcoding backoff limits. More details on recovery handling appear in the Rate Limits guide.

Billing and Observability Headers

For non-streaming and asynchronous interactions, TokenLab provides identification headers to trace charges and background work:

  • X-Billing-Transaction-ID: Returned when billing settles before the HTTP response is dispatched. Non-streaming OpenAI-compatible endpoints include billing_transaction_id in the JSON body, but Gemini and native format endpoints expose it through this header. Streaming calls may settle after the connection closes; when absent, retrieve the ID from workspace usage records. Review settlement workflows in the Billing and Pricing guide.
  • X-Task-ID: Returned on response headers when creating asynchronous jobs for video, music, 3D, or task-based image generation. It provides a header-level correlation ID corresponding to the task id. Consult the Logs and Troubleshooting guide for logging standards.

Implementation: Capturing Headers and Retrying on 429

The following Python example illustrates how to submit a request to the Chat Completions endpoint, inspect transaction identifiers, and handle Retry-After headers during rate limits:

import os
import time
import requests

API_KEY = os.environ["TOKENLAB_API_KEY"]
ENDPOINT = "https://api.tokenlab.sh/v1/chat/completions"

headers = {
    "Authorization": f"Bearer {API_KEY}",
    "Content-Type": "application/json",
}

payload = {
    "model": "gpt-5.6-terra",
    "messages": [{"role": "user", "content": "Summarize system status."}]
}

max_attempts = 3
for attempt in range(max_attempts):
    response = requests.post(ENDPOINT, headers=headers, json=payload, timeout=30)

    if response.status_code == 200:
        # Check for billing transaction header on settled non-streaming calls
        billing_id = response.headers.get("X-Billing-Transaction-ID")
        data = response.json()
        print(f"Settled Transaction ID: {billing_id}")
        print(data["choices"][0]["message"]["content"])
        break

    elif response.status_code == 429:
        retry_after = response.headers.get("Retry-After")
        limit = response.headers.get("X-RateLimit-Limit")
        wait_seconds = float(retry_after) if retry_after else 2 ** attempt
        print(f"Rate limit reached ({limit} req/min). Retrying in {wait_seconds}s...")
        time.sleep(wait_seconds)
    else:
        response.raise_for_status()

Logging and Observability Practices

When instrumenting request monitoring, log public tracking identifiers returned in headers and payloads to reconcile records without persisting user prompts or credentials:

  • Retain request_id, X-Billing-Transaction-ID, and X-Task-ID alongside status codes and response latencies.
  • Always redact Authorization headers, raw API keys, and private signed URLs from telemetry pipelines.
  • For server-side financial reconciliation, query GET /v1/management/api-keys/{keyId}/usage instead of scraping dashboard pages or estimating totals from raw token counters alone.

Sources

Related models

Recent model releases

Try the models from this article

Chat, create images, or make video with the same TokenLab balance.