
Prompt Caching Cost Guide: Cache Hits, Prefixes, and Real API Spend
A 20,000-token system prompt costs an estimated $18.18 per 1,000 requests uncached on claude-sonnet-5 and about $2.00 with cache hits, using TokenLab prices observed on 2026-10-03.
Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new
TokenLab Blog

A 20,000-token system prompt costs an estimated $18.18 per 1,000 requests uncached on claude-sonnet-5 and about $2.00 with cache hits, using TokenLab prices observed on 2026-10-03.

A technical comparison of OpenRouter and TokenLab, examining single-schema normalization versus multi-format native gateway routing across OpenAI, Anthropic, and Gemini endpoints.

Migrating from OpenAI to TokenLab starts with changing base_url and api_key, then checking model IDs, streaming, tool calls, error shapes, timeouts and billing against dated documentation.

A coding agent should not break because a model name changed. Here is how an MCP model catalog makes model choice a runtime lookup, with request examples, routing rules, and verification habits.

A dated four-model streaming sample from TokenLab shows why first visible token, total time and tokens per second must be measured separately, with a harness, headers to log and rate-limit math.

Hitting a 404 because a model string moved is a bad way to start an agent loop. Here is how to pin `gemini-3.5-flash`, read Google Cloud pricing, and build a loop that survives deprecations.