Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

TokenLab for Agents: Machine-Readable Models, Pricing, SDKs, and MCP

·September 19, 2026·8 min read·Updated October 3, 2026·1502 views
#feature#agents#mcp#llms-txt#sdk
TokenLab for Agents: Machine-Readable Models, Pricing, SDKs, and MCP

What you’ll learn

  • Do I need an API key to use the TokenLab MCP server?
  • Which MCP profile should I start with?
  • Which model fields should an agent trust when choosing a model?
  • Why did my agent get a 503 all_channels_failed error?
  • Does the MCP server support Gemini Files or cachedContents?

A coding agent that picks a model ID from memory will eventually pick one that no longer exists, and it only finds out after a 404 or a surprising bill. TokenLab MCP gives the agent a live catalog to check first, so it can verify the ID, the accepted request format, and the price before it writes any integration code. We originally described the server as strictly read-only. The docs observed on 2026-10-03 say otherwise, so this version corrects that and adds a worked flow.

Key Takeaways

  • The TokenLab MCP server has three profiles: catalog (no API key), core, and full. Only catalog is key-free.
  • With a key, it can also send model requests, create media, handle files, and check async tasks. It is not read-only.
  • Route on tokenlab.accepted_request_formats, tokenlab.pricing, tokenlab.lifecycle, and tokenlab.deliveryAvailability. Don't hard-code recommendation order.
  • Gemini Files, resumable uploads, and cachedContents are not in any MCP profile.
  • Never paste an API key into a prompt or tool argument.

What the TokenLab MCP server gives a coding agent

Per the MCP server docs (observed 2026-10-03), the TokenLab MCP Server lets a client browse current models and prices, send model requests, create media, work with files, and check async tasks. The docs name these capabilities:

  • list models and read one model's capabilities (list_models, get_model)
  • read current prices or compare several models
  • send Chat Completions, Responses, Anthropic Messages, or Gemini requests
  • evaluate typed decisions with evaluate_decisions
  • create or edit images; create video, music, 3D, speech, transcription, or translation
  • upload and retrieve files through the OpenAI-compatible /v1/files API
  • create embeddings or rerank documents
  • check and cancel supported async tasks (get_task_status for polling)

The docs don't list tool names for pricing or the API overview in this version. Our older draft named get_model_pricing and get_api_overview. Check your connected client's tool list before you rely on those two names.

Tool availability depends on the profile:

Profile API key Includes
catalog Not required Model list, model details, prices, comparisons, API overview
core Required for paid calls Common chat, decision, media, audio, file, task, embedding, rerank, and translation tools
full Required for paid calls core plus additional developer APIs

Start with catalog if you only want better model selection. Use core when the client should create content or call a model.

Install the TokenLab MCP server in your client

The package needs Node.js 18.17 or newer and npx. It runs locally over stdio, so you need no global install. Back up your active configuration first, and add only the TokenLab entry. These commands come from the docs, observed 2026-10-03.

Claude Code:

claude mcp add \
  --env TOKENLAB_MCP_TOOL_PROFILE=catalog \
  --scope user \
  tokenlab -- \
  npx -y @tokenlabai/mcp-server@0.6.26

Codex:

codex mcp add \
  --env TOKENLAB_MCP_TOOL_PROFILE=catalog \
  tokenlab -- \
  npx -y @tokenlabai/mcp-server@0.6.26

Cursor (~/.cursor/mcp.json or .cursor/mcp.json):

{
  "mcpServers": {
    "tokenlab": {
      "command": "npx",
      "args": ["-y", "@tokenlabai/mcp-server@0.6.26"],
      "env": {
        "TOKENLAB_MCP_TOOL_PROFILE": "catalog"
      }
    }
  }
}

VS Code uses .vscode/mcp.json with a servers key and "type": "stdio". Claude Desktop uses the same shape as Cursor in claude_desktop_config.json. Copy both from the docs page.

To enable paid tools, create a key in Console → API keys and set both variables in the server environment:

{
  "env": {
    "TOKENLAB_API_KEY": "<TOKENLAB_API_KEY>",
    "TOKENLAB_MCP_TOOL_PROFILE": "core"
  }
}

Then restart the client and run claude mcp list or codex mcp list. Ask the agent to call list_models. A non-empty list confirms the package started and reached TokenLab. If a real key lands in a shared file, log, or shell history, revoke it and create a new one.

One agent flow: discover, check, call

Here is the flow we use. Imagine an agent asked to add image generation to a Node.js app. Every value below comes from docs and live model pages observed 2026-10-03.

1. Discover. Ask for the current shortlist, with the MCP tool or plain HTTP:

{ "tool": "list_models", "arguments": { "recommended_for": "image" } }
curl "https://api.tokenlab.sh/v1/models?recommended_for=image"

The valid recommended_for values are image, video, music, 3d, tts, stt, embedding, rerank, and translation. Say the agent picks nano-banana-pro.

2. Check formats and price. Call get_model, or GET /v1/models/nano-banana-pro (live model API, observed 2026-10-03). It reports:

  • accepted request formats: gemini_generate_content, which maps to /v1beta/models/{model}:generateContent
  • capabilities: image-edit, image-to-image, text-to-image
  • price: per_request of 0.067 USD, with a price range of 0.067 to 0.12 (pricing updated 2026-10-02T16:53:30.068Z)

An agent that assumed Chat Completions would have written the wrong code. Compare gpt-image-2 (live model API). It lists no accepted request formats and is token-priced at 3.5 USD input and 21 USD output per 1M tokens. The price shape differs per model, so the agent must read it per model.

3. Make the call. For a chat model, the format check decides the endpoint. gpt-5.6-terra accepts openai_chat_completions and openai_responses (live model API, observed 2026-10-03), so the standard SDK works:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["TOKENLAB_API_KEY"],
    base_url="https://api.tokenlab.sh/v1",
)

response = client.chat.completions.create(
    model="gpt-5.6-terra",
    messages=[{"role": "user", "content": "Reply only with OK."}],
)
print(response.choices[0].message.content)

Send the chosen model ID explicitly. The docs state that TokenLab does not silently replace it. The client should ask for approval before a paid call if price or model choice isn't already confirmed.

For a cost estimate, gpt-5.6-terra charges 0.6 USD per 1M input tokens up to 272K input tokens. A 10,000-token prompt then costs about 10,000 / 1,000,000 × 0.6 = 0.006 USD for input (estimate, before output). Above 272K input tokens, the whole request moves to the higher tier at 1.2 USD input and 5.4 USD output.

Which model API fields an agent should trust for routing

Read these from GET /v1/models/{model} (Get a Model, observed 2026-10-03):

Field What it means for routing
tokenlab.accepted_request_formats Which endpoint family to use: openai_chat_completions is /v1/chat/completions, openai_responses is /v1/responses, anthropic_messages is /v1/messages
tokenlab.pricing / pricing_unit Current public price and its billing unit, such as per_token or per_image
tokenlab.max_input_tokens, max_output_tokens Context and output limits. For gpt-5.6-terra: 1,050,000 and 128,000
tokenlab.supported_operations Operations such as text-to-image or image-to-video
tokenlab.lifecycle Availability, release date, deprecation date, replacement model
tokenlab.deliveryAvailability Configured verified and official support. A missing field means unknown

Two cautions apply. First, an accepted format confirms the endpoint, but individual tools and fields can still vary by model. Second, deliveryAvailability is configured support, not a real-time guarantee. Treat recommended_for results as a shortlist, because the docs say not to pin their order.

For prices alone, GET /v1/models/{model}/pricing is the pricing-only endpoint. Complex entries can carry tiers. seedance-2.0, for example, has resolution- and video-input-dependent output prices from 2.04 to 6.545 USD per 1M tokens (live model API, observed 2026-10-03).

Error fields that guide recovery

On OpenAI-compatible Chat Completions and Responses errors, the error guide (observed 2026-10-03) lists optional did_you_mean, suggestions, hint, retryable, and retry_after. Handle the HTTP status and code first. A 400 model_not_found may carry did_you_mean. Show it to the user rather than swapping models silently. A 503 all_channels_failed can have retryable: false, and repeating it won't help. Anthropic Messages and Gemini keep their native error formats.

What the MCP server does not do

The docs state these limits:

  • It does not change your client's main model provider. Use that client's own setup guide.
  • It does not cover Gemini Files, resumable uploads, or cachedContents. Those need HTTP calls, per Gemini Files and cache.
  • It is not the Skill. The TokenLab Skill installs instructions with npx skills add and starts no MCP server.
  • It does not poll for you on timeout. If a status check times out, don't create a second task.
  • It does not make decisions trustworthy by itself. A Noul answer from evaluate_decisions is a probability, not a Boolean. Validate against your own labeled cases.

The catalog profile can't make paid calls at all. Image tools return either a result or a task, depending on the model. Video, music, and 3D always return tasks.

Where you still need the plain HTTP surfaces

In our pipeline we kept the HTTP discovery endpoints alongside MCP for non-MCP agents. https://api.tokenlab.sh/llms.txt is a compact overview with a first request, common endpoints, and error guidance. The docs observed 2026-10-03 don't cover the llms-full.txt file or the model-data snapshot files from our earlier draft. Verify those URLs yourself before you depend on them. For live status and costs, see the public Models catalog.

FAQ

Do I need an API key to use the TokenLab MCP server?

No, not for browsing. The catalog profile lists models, details, prices, and comparisons without a key. Paid model or media requests need TOKENLAB_API_KEY in the server environment, with the core or full profile.

Which MCP profile should I start with?

Start with catalog if you only want better model selection. Move to core when the client should call models or create media. Use full only if the agent genuinely needs the extra developer APIs.

Which model fields should an agent trust when choosing a model?

Trust accepted_request_formats for the endpoint, pricing with its unit for cost, the token limits, supported_operations, and lifecycle. Treat deliveryAvailability as configured support, not live availability.

Why did my agent get a 503 all_channels_failed error?

The operation may have no supply in the selected Delivery tier. When retryable is false, don't repeat the request. Check availability with GET /v1/models and choose another model with the user's approval.

Does the MCP server support Gemini Files or cachedContents?

No. The docs say Gemini Files, resumable uploads, and cachedContents currently require HTTP calls. The MCP file tools use the OpenAI-compatible /v1/files API.

Create a key in Console → API keys, then add the catalog profile to your client with the commands above.

Sources

Prices checked 2026-10-03

Related models

Recent model releases

Try the models from this article

Chat, create images, or make video with the same TokenLab balance.