What you’ll learn
- Do I need an API key to use the TokenLab MCP server?
- Which MCP profile should I start with?
- Which model fields should an agent trust when choosing a model?
- Why did my agent get a 503 all_channels_failed error?
- Does the MCP server support Gemini Files or cachedContents?
A coding agent that picks a model ID from memory will eventually pick one that no longer exists, and it only finds out after a 404 or a surprising bill. TokenLab MCP gives the agent a live catalog to check first, so it can verify the ID, the accepted request format, and the price before it writes any integration code. We originally described the server as strictly read-only. The docs observed on 2026-10-03 say otherwise, so this version corrects that and adds a worked flow.
Key Takeaways
- The TokenLab MCP server has three profiles:
catalog(no API key),core, andfull. Onlycatalogis key-free. - With a key, it can also send model requests, create media, handle files, and check async tasks. It is not read-only.
- Route on
tokenlab.accepted_request_formats,tokenlab.pricing,tokenlab.lifecycle, andtokenlab.deliveryAvailability. Don't hard-code recommendation order. - Gemini Files, resumable uploads, and
cachedContentsare not in any MCP profile. - Never paste an API key into a prompt or tool argument.
What the TokenLab MCP server gives a coding agent
Per the MCP server docs (observed 2026-10-03), the TokenLab MCP Server lets a client browse current models and prices, send model requests, create media, work with files, and check async tasks. The docs name these capabilities:
- list models and read one model's capabilities (
list_models,get_model) - read current prices or compare several models
- send Chat Completions, Responses, Anthropic Messages, or Gemini requests
- evaluate typed decisions with
evaluate_decisions - create or edit images; create video, music, 3D, speech, transcription, or translation
- upload and retrieve files through the OpenAI-compatible
/v1/filesAPI - create embeddings or rerank documents
- check and cancel supported async tasks (
get_task_statusfor polling)
The docs don't list tool names for pricing or the API overview in this version. Our older draft named get_model_pricing and get_api_overview. Check your connected client's tool list before you rely on those two names.
Tool availability depends on the profile:
| Profile | API key | Includes |
|---|---|---|
catalog |
Not required | Model list, model details, prices, comparisons, API overview |
core |
Required for paid calls | Common chat, decision, media, audio, file, task, embedding, rerank, and translation tools |
full |
Required for paid calls | core plus additional developer APIs |
Start with catalog if you only want better model selection. Use core when the client should create content or call a model.
Install the TokenLab MCP server in your client
The package needs Node.js 18.17 or newer and npx. It runs locally over stdio, so you need no global install. Back up your active configuration first, and add only the TokenLab entry. These commands come from the docs, observed 2026-10-03.
Claude Code:
claude mcp add \
--env TOKENLAB_MCP_TOOL_PROFILE=catalog \
--scope user \
tokenlab -- \
npx -y @tokenlabai/mcp-server@0.6.26
Codex:
codex mcp add \
--env TOKENLAB_MCP_TOOL_PROFILE=catalog \
tokenlab -- \
npx -y @tokenlabai/mcp-server@0.6.26
Cursor (~/.cursor/mcp.json or .cursor/mcp.json):
{
"mcpServers": {
"tokenlab": {
"command": "npx",
"args": ["-y", "@tokenlabai/mcp-server@0.6.26"],
"env": {
"TOKENLAB_MCP_TOOL_PROFILE": "catalog"
}
}
}
}
VS Code uses .vscode/mcp.json with a servers key and "type": "stdio". Claude Desktop uses the same shape as Cursor in claude_desktop_config.json. Copy both from the docs page.
To enable paid tools, create a key in Console → API keys and set both variables in the server environment:
{
"env": {
"TOKENLAB_API_KEY": "<TOKENLAB_API_KEY>",
"TOKENLAB_MCP_TOOL_PROFILE": "core"
}
}
Then restart the client and run claude mcp list or codex mcp list. Ask the agent to call list_models. A non-empty list confirms the package started and reached TokenLab. If a real key lands in a shared file, log, or shell history, revoke it and create a new one.
One agent flow: discover, check, call
Here is the flow we use. Imagine an agent asked to add image generation to a Node.js app. Every value below comes from docs and live model pages observed 2026-10-03.
1. Discover. Ask for the current shortlist, with the MCP tool or plain HTTP:
{ "tool": "list_models", "arguments": { "recommended_for": "image" } }
curl "https://api.tokenlab.sh/v1/models?recommended_for=image"
The valid recommended_for values are image, video, music, 3d, tts, stt, embedding, rerank, and translation. Say the agent picks nano-banana-pro.
2. Check formats and price. Call get_model, or GET /v1/models/nano-banana-pro (live model API, observed 2026-10-03). It reports:
- accepted request formats:
gemini_generate_content, which maps to/v1beta/models/{model}:generateContent - capabilities:
image-edit,image-to-image,text-to-image - price:
per_requestof 0.067 USD, with a price range of 0.067 to 0.12 (pricing updated 2026-10-02T16:53:30.068Z)
An agent that assumed Chat Completions would have written the wrong code. Compare gpt-image-2 (live model API). It lists no accepted request formats and is token-priced at 3.5 USD input and 21 USD output per 1M tokens. The price shape differs per model, so the agent must read it per model.
3. Make the call. For a chat model, the format check decides the endpoint. gpt-5.6-terra accepts openai_chat_completions and openai_responses (live model API, observed 2026-10-03), so the standard SDK works:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["TOKENLAB_API_KEY"],
base_url="https://api.tokenlab.sh/v1",
)
response = client.chat.completions.create(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "Reply only with OK."}],
)
print(response.choices[0].message.content)
Send the chosen model ID explicitly. The docs state that TokenLab does not silently replace it. The client should ask for approval before a paid call if price or model choice isn't already confirmed.
For a cost estimate, gpt-5.6-terra charges 0.6 USD per 1M input tokens up to 272K input tokens. A 10,000-token prompt then costs about 10,000 / 1,000,000 × 0.6 = 0.006 USD for input (estimate, before output). Above 272K input tokens, the whole request moves to the higher tier at 1.2 USD input and 5.4 USD output.
Which model API fields an agent should trust for routing
Read these from GET /v1/models/{model} (Get a Model, observed 2026-10-03):
| Field | What it means for routing |
|---|---|
tokenlab.accepted_request_formats |
Which endpoint family to use: openai_chat_completions is /v1/chat/completions, openai_responses is /v1/responses, anthropic_messages is /v1/messages |
tokenlab.pricing / pricing_unit |
Current public price and its billing unit, such as per_token or per_image |
tokenlab.max_input_tokens, max_output_tokens |
Context and output limits. For gpt-5.6-terra: 1,050,000 and 128,000 |
tokenlab.supported_operations |
Operations such as text-to-image or image-to-video |
tokenlab.lifecycle |
Availability, release date, deprecation date, replacement model |
tokenlab.deliveryAvailability |
Configured verified and official support. A missing field means unknown |
Two cautions apply. First, an accepted format confirms the endpoint, but individual tools and fields can still vary by model. Second, deliveryAvailability is configured support, not a real-time guarantee. Treat recommended_for results as a shortlist, because the docs say not to pin their order.
For prices alone, GET /v1/models/{model}/pricing is the pricing-only endpoint. Complex entries can carry tiers. seedance-2.0, for example, has resolution- and video-input-dependent output prices from 2.04 to 6.545 USD per 1M tokens (live model API, observed 2026-10-03).
Error fields that guide recovery
On OpenAI-compatible Chat Completions and Responses errors, the error guide (observed 2026-10-03) lists optional did_you_mean, suggestions, hint, retryable, and retry_after. Handle the HTTP status and code first. A 400 model_not_found may carry did_you_mean. Show it to the user rather than swapping models silently. A 503 all_channels_failed can have retryable: false, and repeating it won't help. Anthropic Messages and Gemini keep their native error formats.
What the MCP server does not do
The docs state these limits:
- It does not change your client's main model provider. Use that client's own setup guide.
- It does not cover Gemini Files, resumable uploads, or
cachedContents. Those need HTTP calls, per Gemini Files and cache. - It is not the Skill. The TokenLab Skill installs instructions with
npx skills addand starts no MCP server. - It does not poll for you on timeout. If a status check times out, don't create a second task.
- It does not make decisions trustworthy by itself. A Noul answer from
evaluate_decisionsis a probability, not a Boolean. Validate against your own labeled cases.
The catalog profile can't make paid calls at all. Image tools return either a result or a task, depending on the model. Video, music, and 3D always return tasks.
Where you still need the plain HTTP surfaces
In our pipeline we kept the HTTP discovery endpoints alongside MCP for non-MCP agents. https://api.tokenlab.sh/llms.txt is a compact overview with a first request, common endpoints, and error guidance. The docs observed 2026-10-03 don't cover the llms-full.txt file or the model-data snapshot files from our earlier draft. Verify those URLs yourself before you depend on them. For live status and costs, see the public Models catalog.
FAQ
Do I need an API key to use the TokenLab MCP server?
No, not for browsing. The catalog profile lists models, details, prices, and comparisons without a key. Paid model or media requests need TOKENLAB_API_KEY in the server environment, with the core or full profile.
Which MCP profile should I start with?
Start with catalog if you only want better model selection. Move to core when the client should call models or create media. Use full only if the agent genuinely needs the extra developer APIs.
Which model fields should an agent trust when choosing a model?
Trust accepted_request_formats for the endpoint, pricing with its unit for cost, the token limits, supported_operations, and lifecycle. Treat deliveryAvailability as configured support, not live availability.
Why did my agent get a 503 all_channels_failed error?
The operation may have no supply in the selected Delivery tier. When retryable is false, don't repeat the request. Check availability with GET /v1/models and choose another model with the user's approval.
Does the MCP server support Gemini Files or cachedContents?
No. The docs say Gemini Files, resumable uploads, and cachedContents currently require HTTP calls. The MCP file tools use the OpenAI-compatible /v1/files API.
Create a key in Console → API keys, then add the catalog profile to your client with the commands above.
Sources
Prices checked 2026-10-03
- TokenLab Docs: TokenLab MCP ServerSources checked 2026-10-03
- TokenLab Docs: Errors agents can act onSources checked 2026-10-03
- TokenLab Docs: List ModelsSources checked 2026-10-03
- TokenLab Docs: Get a ModelSources checked 2026-10-03
- TokenLab Docs: Get PricingSources checked 2026-10-03
- TokenLab Docs: TokenLab API skill for coding agentsSources checked 2026-10-03
- TokenLab live model API: gpt-5.6-terraSources checked 2026-10-03
- TokenLab live model API: gpt-image-2Sources checked 2026-10-03



