Core Guides
Errors agents can act on
Use error codes, retry timing, and model suggestions without parsing prose
This page describes machine-readable public API errors for applications and coding agents. It does not grant access to workspace request investigations or support. For a request in your workspace, start with request troubleshooting.
OpenAI-compatible TokenLab errors can include structured hints for an agent or application. Use these fields when present; do not parse the human-readable message to decide what to do.
Anthropic Messages and Gemini APIs keep their native error formats, so the extensions on this page apply only to OpenAI-compatible Chat Completions and Responses errors.
Optional error fields
All fields below appear inside the error object and may be absent.
| Field | Type | Use |
|---|---|---|
did_you_mean | string | Closest available model ID |
suggestions | array | Models that may fit the request |
hint | string | A short explanation or suggested action |
retryable | boolean | Whether the same request may succeed later |
retry_after | number | Seconds to wait before trying again |
balance_usd | number | Current balance in USD |
estimated_cost_usd | number | Estimated cost of the rejected request |
Your client should still handle every error by its HTTP status and code. Treat these extra fields as useful context, not required fields.
Unknown model
A misspelled or unavailable model returns 400 model_not_found. If did_you_mean is present, show it to the user or retry only when your product already has permission to change the selected model.
{
"error": {
"message": "Model not found: please check the model name",
"type": "invalid_request_error",
"param": "model",
"code": "model_not_found",
"did_you_mean": "gpt-5.6-terra",
"suggestions": [
{"id": "gpt-5.6-terra"},
{"id": "gpt-5.6-luna"}
],
"hint": "Did you mean 'gpt-5.6-terra'? Use GET https://api.tokenlab.sh/v1/models to list all available models."
}
}Insufficient balance
402 insufficient_balance can include the current balance and the estimated amount required. Your application can offer a top-up link, a less expensive model, or a smaller request.
{
"error": {
"message": "Insufficient balance: need ~$0.3500 for claude-sonnet-4-6, but balance is $0.1200.",
"type": "insufficient_balance",
"code": "insufficient_balance",
"balance_usd": 0.12,
"estimated_cost_usd": 0.35,
"suggestions": [
{"id": "gpt-5.6-luna"},
{"id": "deepseek-v3-2"}
],
"hint": "Try a cheaper model, or top up at https://tokenlab.sh/dashboard/billing."
}
}Model unavailable
A 503 all_channels_failed or 503 delivery_tier_unavailable does not always mean a temporary outage. If the requested operation has no supply in the selected Delivery tier, retryable is false and retry_after is omitted. Do not repeat the same request. Check the operation and Delivery availability with GET /v1/models before selecting another model. Similar names do not prove availability; unverified alternatives are omitted.
{
"error": {
"message": "This model is unavailable for the requested operation and Delivery tier.",
"type": "all_channels_failed",
"code": "all_channels_failed",
"retryable": false,
"hint": "Check the model's operation and Delivery availability with GET /v1/models. Repeating the same request will not resolve this."
}
}Rate limit
For 429 rate_limit_exceeded, wait for retry_after seconds or use the standard Retry-After response header.
{
"error": {
"message": "Rate limit: 1000 rpm exceeded",
"type": "rate_limit_exceeded",
"code": "rate_limit_exceeded",
"retryable": true,
"retry_after": 8,
"hint": "Retry after 8s."
}
}Context too long
400 context_length_exceeded is not fixed by sending the same request again. Shorten the input or let the user choose a model with a larger context window.
{
"error": {
"message": "This model's maximum context length is 128000 tokens...",
"type": "invalid_request_error",
"code": "context_length_exceeded",
"retryable": false,
"suggestions": [
{"id": "gemini-2.5-pro"},
{"id": "claude-sonnet-5"}
],
"hint": "Reduce your input or switch to a model with a larger context window."
}
}Find the right API format
Read tokenlab.accepted_request_formats from GET /v1/models/{model} before using a model-specific API.
| Value | Endpoint |
|---|---|
openai_chat_completions | /v1/chat/completions |
openai_responses | /v1/responses |
anthropic_messages | /v1/messages |
gemini_generate_content | /v1beta/models/{model}:generateContent |
An accepted format confirms the endpoint. Individual tools and fields can still vary by model; check the model page before depending on them.
Find a model by task
The Models API can return a current shortlist for non-chat tasks:
curl "https://api.tokenlab.sh/v1/models?recommended_for=image"Valid recommended_for values are image, video, music, 3d, tts, stt, embedding, rerank, and translation. Send the chosen model ID explicitly in the create request. TokenLab does not silently replace it with a different model.
Machine-readable overview
Agents can read a compact API overview at:
GET https://api.tokenlab.sh/llms.txtIt includes a first request, common endpoints, model filters, and error-handling guidance.
Handle an error without replaying the request
This example makes one request, preserves the selected model, and reports structured recovery information. SDK retries are disabled so your application can decide whether a retry is safe. Present model suggestions for an explicit choice; do not replay an accepted or timed-out generation automatically.
import os
from openai import OpenAI, APIStatusError
with OpenAI(
api_key=os.environ["TOKENLAB_API_KEY"],
base_url="https://api.tokenlab.sh/v1",
timeout=30.0,
max_retries=0,
) as client:
try:
response = client.chat.completions.create(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "Reply only with OK."}],
)
print(response.choices[0].message.content)
except APIStatusError as exc:
body = exc.body if isinstance(exc.body, dict) else {}
error = body.get("error", body)
if not isinstance(error, dict):
error = {}
print({
"status": exc.status_code,
"request_id": exc.request_id,
"code": error.get("code"),
"hint": error.get("hint"),
"suggested_model": error.get("did_you_mean"),
"retry_after": exc.response.headers.get("Retry-After") or error.get("retry_after"),
})
raise