TokenLab

Core Guides

Errors agents can act on

Use error codes, retry timing, and model suggestions without parsing prose

This page describes machine-readable public API errors for applications and coding agents. It does not grant access to workspace request investigations or support. For a request in your workspace, start with request troubleshooting.

OpenAI-compatible TokenLab errors can include structured hints for an agent or application. Use these fields when present; do not parse the human-readable message to decide what to do.

Anthropic Messages and Gemini APIs keep their native error formats, so the extensions on this page apply only to OpenAI-compatible Chat Completions and Responses errors.

Optional error fields

All fields below appear inside the error object and may be absent.

FieldTypeUse
did_you_meanstringClosest available model ID
suggestionsarrayModels that may fit the request
hintstringA short explanation or suggested action
retryablebooleanWhether the same request may succeed later
retry_afternumberSeconds to wait before trying again
balance_usdnumberCurrent balance in USD
estimated_cost_usdnumberEstimated cost of the rejected request

Your client should still handle every error by its HTTP status and code. Treat these extra fields as useful context, not required fields.

Unknown model

A misspelled or unavailable model returns 400 model_not_found. If did_you_mean is present, show it to the user or retry only when your product already has permission to change the selected model.

{
  "error": {
    "message": "Model not found: please check the model name",
    "type": "invalid_request_error",
    "param": "model",
    "code": "model_not_found",
    "did_you_mean": "gpt-5.6-terra",
    "suggestions": [
      {"id": "gpt-5.6-terra"},
      {"id": "gpt-5.6-luna"}
    ],
    "hint": "Did you mean 'gpt-5.6-terra'? Use GET https://api.tokenlab.sh/v1/models to list all available models."
  }
}

Insufficient balance

402 insufficient_balance can include the current balance and the estimated amount required. Your application can offer a top-up link, a less expensive model, or a smaller request.

{
  "error": {
    "message": "Insufficient balance: need ~$0.3500 for claude-sonnet-4-6, but balance is $0.1200.",
    "type": "insufficient_balance",
    "code": "insufficient_balance",
    "balance_usd": 0.12,
    "estimated_cost_usd": 0.35,
    "suggestions": [
      {"id": "gpt-5.6-luna"},
      {"id": "deepseek-v3-2"}
    ],
    "hint": "Try a cheaper model, or top up at https://tokenlab.sh/dashboard/billing."
  }
}

Model unavailable

A 503 all_channels_failed or 503 delivery_tier_unavailable does not always mean a temporary outage. If the requested operation has no supply in the selected Delivery tier, retryable is false and retry_after is omitted. Do not repeat the same request. Check the operation and Delivery availability with GET /v1/models before selecting another model. Similar names do not prove availability; unverified alternatives are omitted.

{
  "error": {
    "message": "This model is unavailable for the requested operation and Delivery tier.",
    "type": "all_channels_failed",
    "code": "all_channels_failed",
    "retryable": false,
    "hint": "Check the model's operation and Delivery availability with GET /v1/models. Repeating the same request will not resolve this."
  }
}

Rate limit

For 429 rate_limit_exceeded, wait for retry_after seconds or use the standard Retry-After response header.

{
  "error": {
    "message": "Rate limit: 1000 rpm exceeded",
    "type": "rate_limit_exceeded",
    "code": "rate_limit_exceeded",
    "retryable": true,
    "retry_after": 8,
    "hint": "Retry after 8s."
  }
}

Context too long

400 context_length_exceeded is not fixed by sending the same request again. Shorten the input or let the user choose a model with a larger context window.

{
  "error": {
    "message": "This model's maximum context length is 128000 tokens...",
    "type": "invalid_request_error",
    "code": "context_length_exceeded",
    "retryable": false,
    "suggestions": [
      {"id": "gemini-2.5-pro"},
      {"id": "claude-sonnet-5"}
    ],
    "hint": "Reduce your input or switch to a model with a larger context window."
  }
}

Find the right API format

Read tokenlab.accepted_request_formats from GET /v1/models/{model} before using a model-specific API.

ValueEndpoint
openai_chat_completions/v1/chat/completions
openai_responses/v1/responses
anthropic_messages/v1/messages
gemini_generate_content/v1beta/models/{model}:generateContent

An accepted format confirms the endpoint. Individual tools and fields can still vary by model; check the model page before depending on them.

Find a model by task

The Models API can return a current shortlist for non-chat tasks:

curl "https://api.tokenlab.sh/v1/models?recommended_for=image"

Valid recommended_for values are image, video, music, 3d, tts, stt, embedding, rerank, and translation. Send the chosen model ID explicitly in the create request. TokenLab does not silently replace it with a different model.

Machine-readable overview

Agents can read a compact API overview at:

GET https://api.tokenlab.sh/llms.txt

It includes a first request, common endpoints, model filters, and error-handling guidance.

Handle an error without replaying the request

This example makes one request, preserves the selected model, and reports structured recovery information. SDK retries are disabled so your application can decide whether a retry is safe. Present model suggestions for an explicit choice; do not replay an accepted or timed-out generation automatically.

import os
from openai import OpenAI, APIStatusError

with OpenAI(
    api_key=os.environ["TOKENLAB_API_KEY"],
    base_url="https://api.tokenlab.sh/v1",
    timeout=30.0,
    max_retries=0,
) as client:
    try:
        response = client.chat.completions.create(
            model="gpt-5.6-terra",
            messages=[{"role": "user", "content": "Reply only with OK."}],
        )
        print(response.choices[0].message.content)
    except APIStatusError as exc:
        body = exc.body if isinstance(exc.body, dict) else {}
        error = body.get("error", body)
        if not isinstance(error, dict):
            error = {}
        print({
            "status": exc.status_code,
            "request_id": exc.request_id,
            "code": error.get("code"),
            "hint": error.get("hint"),
            "suggested_model": error.get("did_you_mean"),
            "retry_after": exc.response.headers.get("Retry-After") or error.get("retry_after"),
        })
        raise

On this page