What you’ll learn
- Why do I see two prices for DeepSeek V4 Flash?
- Can I use the external reference prices for budgeting?
- Which current SSOT model is cheapest for a production agent?
- Why are PixVerse, fal, and BFL cited if this article is about LLM pricing?
- How do tool-calling retries change the cheapest model choice?
In our agent pipeline, the cheapest model per token is rarely the cheapest model per completed task. When we tested the best cheap AI models for agents, output tokens from tool calls, JSON repair, and chain-of-thought scratch dominated spend. DeepSeek V4 Flash, Gemini 3.5 Flash, Claude Sonnet 5, GPT-5.5, GLM-5.2, and Laguna XS 2.1 are the current SSOT examples we can discuss. TokenLab catalog prices below are observed at 2026-09-19; external prices remain undated and need provider confirmation. We do not publish a latency benchmark because no controlled TTFT, tokens/sec, or retry-rate dataset was available for these exact routes at refresh time. Treat this as a cost-and-failure-mode shortlist, then benchmark latency in your own loop.
Key Takeaways
- Output cost drives agent spend more than input cost. Agents generate far more tokens through tool calls, retries, and JSON than they consume per turn.
- Among current SSOT examples with TokenLab catalog prices, DeepSeek V4 Flash is the lowest output-cost row we can verify here at $0.147 input / $0.294 output per MTok.
- Model name collisions are real. DeepSeek V4 Flash appears at two price points: $0.147/$0.294 in the TokenLab catalog and $0.09/$0.18 in an external listing.
- Video and image generation pricing (PixVerse, fal, Black Forest Labs) is not evidence for LLM token pricing. It should never appear as a citation for text-agent costs, and we separated it below.
- External comparison prices for Claude Sonnet 5, GLM-5.2, Kimi K2.7 Code, Qwen3.7 Plus, and others lack a dated public source URL. Confirm them against each provider's own pricing page before you budget.
- Cheapest-per-token is not automatically cheapest-per-task. A model that fails tool calls or truncates JSON forces retries, which erases savings.
Source Snapshot and Failure Modes
| Source | Content Type | Observed Date | Relevance to This Article |
|---|---|---|---|
| PixVerse Platform Docs (docs.platform.pixverse.ai) | Video generation pricing | 2026-07-08 | Not used for LLM claims - video-only, included for transparency |
| fal PixVerse V6 model page (fal.ai/pixverse-v6) | Video generation pricing | 2026-07-08 | Not used for LLM claims - video-only |
| Black Forest Labs pricing docs (docs.bfl.ml) | Image generation pricing | 2026-07-08 | Not used for LLM claims - image-only |
| TokenLab internal model catalog | First-party LLM & media pricing | 2026-09-19 | Primary source for all TokenLab catalog prices below |
| Individual provider pricing pages (OpenAI, Anthropic, Google, DeepSeek, Moonshot, Qwen, MiniMax, Z.ai) | LLM pricing | Not dated in this dataset | Must be independently verified before public budgeting claims |
The PixVerse, fal, and BFL sources exist in our research set because we pulled them during a broader media-pricing sweep. They describe per-second video billing and per-image credit costs, which have no bearing on per-token LLM pricing. We list them here so readers can see why we exclude them. We do not use them for LLM claims.
In our pipeline, we log tool-calling retry rates, truncated JSON, missed tool calls, and hallucinated function arguments alongside token spend. We do not have a controlled retry-rate dataset for these exact routes at refresh time, so we cannot publish rates. For example, a $0.05/MTok model that fails 15% of tool calls can cost more in retries than a $1/MTok model that fails 2%. That failure-mode math, not sticker price, should drive your cheap-tier choice.
Model And Cost Comparison for Best Cheap AI Models for Agents
All prices below are TokenLab's own first-party catalog listings (per-token, expressed as $/MTok), observed at 2026-09-19. They are the correct basis for LLM agent cost claims in this article.
| Model | Provider | Input $/MTok | Output $/MTok | Agent Fit |
|---|---|---|---|---|
| DeepSeek V4 Flash | DeepSeek | $0.147 | $0.294 | Cheapest output-heavy current SSOT option in TokenLab's catalog; confirm exact model ID before comparing to external quotes |
| Gemini 3.5 Flash | $1.50 | $9.00 | Multimodal agent steps (image + text tool use) | |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | Reserve for final verification/QA pass on agent output |
| GPT-5.5 | OpenAI | $5.00 | $30.00 | Fallback for complex planning or ambiguous tool selection |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | Flagship comparison; rarely justified inside an agent loop |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | Premium ceiling |
| GLM-5.2 | Z.ai | $0.93 | $3.00 | Open-weight example for low-cost routing |
| Kimi K2.7 Code | Moonshot AI | $0.74 | $3.50 | Coding-agent example |
| Qwen3.7 Plus | Alibaba/Qwen | $0.32 | $1.28 | Low-cost routing example |
| MiniMax M3 | MiniMax | $0.30 | $1.20 | Low-cost routing example |
| DeepSeek V4 Pro | DeepSeek | $0.441 | $0.882 | Heavier reasoning subtasks within the same family |
We keep the external figures below for audit continuity and mark each row as unverified. Do not use them for budgeting. The exact detail must be confirmed against the current docs.
| Model | Provider | Input $/MTok | Output $/MTok | Verification Status |
|---|---|---|---|---|
| DeepSeek V4 Flash (external listing) | DeepSeek | $0.09 | $0.18 | Differs from TokenLab catalog above - confirm which product/tier this refers to; no dated source URL in our source set |
| MiniMax M3 | MiniMax | $0.30 | $1.20 | No dated source URL in our source set; confirm against provider docs |
| Qwen3.7 Plus | Alibaba/Qwen | $0.32 | $1.28 | No dated source URL in our source set; confirm against provider docs |
| Kimi K2.7 Code | Moonshot AI | $0.74 | $3.50 | No dated source URL in our source set; confirm against provider docs |
| GLM-5.2 | Z.ai | $0.93 | $3.00 | No dated source URL in our source set; confirm against provider docs |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | No dated source URL in our source set; confirm against provider docs |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | No dated source URL in our source set; confirm against provider docs |
| GPT-5.5 Batch/Flex | OpenAI | $2.50 | $15.00 | No dated source URL in our source set; not the standard GPT-5.5 row |
| GPT-5.5 | OpenAI | $5.00 | $30.00 | No dated source URL in our source set; confirm against provider docs |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | No dated source URL in our source set; confirm against provider docs |
Retired catalog price points no longer match a current SSOT model name. We preserve the numbers for audit continuity, but we do not route new agent work to these unnamed tiers without provider confirmation.
| Retired or non-SSOT tier | Input $/MTok | Output $/MTok | Status |
|---|---|---|---|
| Retired ultra-low-cost tier | $0.05 | $0.40 | No longer a current SSOT example; confirm successor before use |
| Retired low-cost flash-lite tier | $0.25 | $1.50 | No longer a current SSOT example; map to Gemini 3.5 Flash or DeepSeek V4 Flash and confirm |
| Retired low-cost mini tier | $0.25 | $2.00 | No longer a current SSOT example; map to Claude Sonnet 5 or DeepSeek V4 Flash and confirm |
| Retired small validation tier | $1.00 | $5.00 | No longer a current SSOT example; map to Claude Sonnet 5 and confirm |
| Retired mid-tier fallback | $1.25 | $10.00 | No longer a current SSOT example; map to GPT-5.5 and confirm |
| Retired premium ceiling | $15.00 | $120.00 | No longer a current SSOT example; map to GPT-5.5 or Claude Fable 5 and confirm |
Implementation Notes for Best Cheap AI Models for Agents
- Tier your agent by task, not by a single cheapest model. Route routing, classification, and tool selection to DeepSeek V4 Flash or Gemini 3.5 Flash. Escalate multi-step planning to DeepSeek V4 Pro or GPT-5.5. Reserve Claude Sonnet 5 for final-answer verification.
- Cap output tokens aggressively at the cheap tier. Since output cost dominates agent spend, set tight max-token limits on DeepSeek V4 Flash and Gemini 3.5 Flash calls. Only allow longer generations at the escalation tier.
- Log failure modes per model, not just cost per call. Track truncated JSON, missed tool calls, and hallucinated function arguments separately from token spend. For example, a $0.05/MTok model that fails 15% of tool calls can cost more in retries than a $1/MTok model that fails 2%.
- Pin exact model IDs in code, never aliases. Given the naming collision shown above (two different DeepSeek V4 Flash price points), hardcode the full vendor and model string. Re-verify pricing on each provider update.
- Re-check external pricing quarterly. Anything without a dated source URL in this article should be re-pulled from the provider's live pricing page before it appears in a customer-facing cost estimate.
Here is a routing sketch that uses the catalog model names in this article. The exact TokenLab request shape and tool-call error fields must be confirmed in the TokenLab API docs before production use.
Routing sketch, not a TokenLab API contract. Confirm the request shape in the TokenLab API docs.
for model in ['DeepSeek V4 Flash', 'Gemini 3.5 Flash', 'Claude Sonnet 5']:
retry once if the tool call is missing or the JSON is truncated
if the response contains a valid tool call and valid JSON:
return the response
log the failure mode for this model
raise an error after the final tier fails
This sketch shows the retry boundary: catch a tool-call error, log the failure mode, then move to the next model. Confirm the request shape and error fields in the current TokenLab API docs before production.
Where TokenLab Fits Agent Routing
- A single catalog reduces per-vendor account juggling. TokenLab exposes OpenAI, Google, Anthropic, and DeepSeek pricing tiers under one API and billing surface. Switching a workflow step from DeepSeek V4 Flash to Gemini 3.5 Flash becomes a config change, not a new integration.
- First-party, dated pricing you can audit. Unlike scattered third-party quotes, TokenLab catalog numbers in this article are pulled from the platform's own listings observed at 2026-09-19. That is why they are the only prices here presented without a must-verify caveat.
- Built-in tiering for agent pipelines. Because TokenLab hosts both cheap-tier and escalation-tier models side by side, you can A/B cost and failure rate across DeepSeek V4 Flash, Gemini 3.5 Flash, and Claude Sonnet 5. You do not need to re-architect your agent's routing logic.
Related Reading
FAQ
Why do I see two prices for DeepSeek V4 Flash?
DeepSeek V4 Flash appears at two price points in our source set: $0.147/$0.294 in TokenLab's catalog and $0.09/$0.18 in an external listing. Model names get reused and re-priced across providers and resellers. Pin the exact vendor and model ID in code, then re-verify against the provider's live docs rather than trusting a name alone.
Can I use the external reference prices for budgeting?
Not yet. Those rows lack a dated source URL in our source set, so we mark them as needing provider confirmation. The TokenLab catalog figures are the planning basis, and the external table is a starting point for your own verification, not a final number. The exact detail must be confirmed against the current docs.
Which current SSOT model is cheapest for a production agent?
By output cost among current SSOT examples with TokenLab catalog prices, DeepSeek V4 Flash is the lowest we can verify here at $0.294/MTok output. The external listing shows $0.18/MTok output, but that number needs provider confirmation. Measure cheapest after accounting for retry rate on your specific tool-calling schema.
Why are PixVerse, fal, and BFL cited if this article is about LLM pricing?
They are not cited as evidence for any LLM price claim. They appear only in the Source Snapshot table for transparency about our broader research sweep and are explicitly marked not used for LLM claims. Every dollar figure in the Model And Cost Comparison table comes from TokenLab's own catalog or is flagged as unverified.
How do tool-calling retries change the cheapest model choice?
A cheap model that fails tool calls or truncates JSON forces retries, and those retries add output tokens. For example, a $0.05/MTok model that fails 15% of tool calls can cost more in retries than a $1/MTok model that fails 2%. In our pipeline, we log retry rates per model and compare cost per completed task, not cost per token.
Explore the TokenLab model catalog and start comparing costs today: TokenLab model catalog.
Sources
Prices checked 2026-07-07
- PixVerse Platform DocsSources checked 2026-07-08
- fal PixVerse V6 model pageSources checked 2026-07-08
- Black Forest Labs pricing docsSources checked 2026-07-08
- fal FLUX.2 model pageSources checked 2026-07-08
- Google AI Gemini API pricingSources checked 2026-07-08
- Claude Platform pricingSources checked 2026-07-08
- OpenAI API pricingSources checked 2026-07-08
- DeepSeek API pricingSources checked 2026-07-08



