What you’ll learn
- Can I call deepseek-v4-pro through the Anthropic Messages format?
- Should I retry a 503 from deepseek-v4-pro or deepseek-v4-flash?
- Does deepseek-v4.1-flash replace deepseek-v4-flash?
- Do cached tokens make deepseek-v4-flash cheaper in an agent loop?
- Will routing to deepseek-v4-flash raise my rate limit?
Sending every step of a coding session to one model is the easiest routing policy, and usually the most expensive one. This tutorial shows how to use the DeepSeek V4 API for coding on TokenLab by splitting work between deepseek-v4-pro and deepseek-v4-flash. We read both model records from the live API on 2026-10-03, and everything below comes from those records and the TokenLab docs. You get a comparison table, a worked cost estimate, a tool-calling request, retry and fallback code, and a preflight check.
Key Takeaways
- Both models list a 1,000,000-token input limit, a 384,000-token output limit, and the same three request formats. Price is the main difference between them.
- At list prices,
deepseek-v4-procosts 4.4 times as much asdeepseek-v4-flashper input token and 3.3 times as much per output token. - In our 20-call example, routing 4 calls to pro and 16 to flash costs about $0.18 off-peak. Sending all 20 to pro costs about $0.46.
- Retry
429afterRetry-After. Retry500–504only whenretryableistrue. Never retry400,401,402,403,404, or413unchanged. - The catalog lists
deepseek-v4.1-flashas active. Neitherdeepseek-v4-pronordeepseek-v4-flashnames a replacement model. - Read limits, formats, and prices from
GET /v1/models/:modelbefore you route. Do not hard-code a copied table.
The DeepSeek V4 API for Coding: What the Catalog Says
We fetched both records on 2026-10-03. The table below compares them side by side. Prices are USD per 1M tokens, and the catalog pricing was last updated 2026-10-02T16:53:30.068Z.
| Item | deepseek-v4-pro |
deepseek-v4-flash |
Source, observed 2026-10-03 |
|---|---|---|---|
| Context limit (max input tokens) | 1,000,000 | 1,000,000 | pro, flash |
| Output limit (max output tokens) | 384,000 | 384,000 | pro, flash |
| Accepted request formats | anthropic_messages, openai_chat_completions, openai_responses |
anthropic_messages, openai_chat_completions, openai_responses |
pro, flash |
| Capabilities | json-mode, prompt-cache, tool-use |
json-mode, prompt-cache, tool-use |
pro, flash |
| Off-peak input | $0.66 | $0.15 | pro, flash |
| Off-peak output | $1.98 | $0.60 | pro, flash |
| Off-peak cache read | $0.022 | $0.003 | pro, flash |
| Off-peak cache write | $0.66 | not listed | pro, flash |
| Peak input | $1.32 | $0.30 | pro, flash |
| Peak output | $3.96 | $1.20 | pro, flash |
| Peak cache read | $0.044 | $0.006 | pro, flash |
| Lifecycle stage | active, released 2026-04-24 | active, released 2026-04-24 | pro, flash |
The default price block in each record matches the off-peak entry. Peak timing differs by record. For deepseek-v4-pro, peak prices apply at 09:00-12:00 and 14:00-18:00 Beijing time. For deepseek-v4-flash, the record says peak windows apply on weekdays, excluding China public holidays. It says off-peak includes weekends and those holidays, but it gives no hours. Check the pricing endpoint before you budget around the flash window.
Lifecycle and newer DeepSeek models
Both records show lifecycle stage active, with replacement model, deprecated_at, and retired_at all empty. So the catalog does not schedule either model for removal, and it names no successor for either.
The catalog also lists deepseek-v4.1-flash. Its record, observed 2026-10-03, is active with no release date and no replacement. It has the same limits, formats, and list prices as deepseek-v4-flash. It adds reasoning and vision to the capability list and shows a cache write price of $0.15 off-peak.
That is a separate model ID, so this article keeps its subject. We would test deepseek-v4.1-flash on your own tasks before swapping it in. The catalog also lists deepseek-v4-flash-vision-exp, but we did not read its record. Verify it on the Models page if you need it.
Routing deepseek-v4-pro and deepseek-v4-flash by Task
Imagine an agent session that plans a change across five modules, writes the edits, then generates a dozen test stubs. The first step needs the most context and care. The last one is repetitive and cheap to redo. The catalog cannot tell you where the quality line sits. The TokenLab guide on coding-agent models, observed 2026-10-03, says leaderboard results do not predict how a model follows your own instructions and tools.
Our starting heuristic assumes the pricier model earns its cost on cross-file work. Treat it as a hypothesis to test, not a finding:
+-------------------------------------------------------------+
| Incoming Task |
+-------------------------------------------------------------+
|
[Does task involve multi-file context,
backward compatibility, or security review?]
|
+---------------+---------------+
| |
[Yes] [No]
| |
v v
deepseek-v4-pro deepseek-v4-flash
Criteria that push a step to deepseek-v4-pro:
- Modifying logic across multiple imported files.
- Security or vulnerability assessments.
- Strict backward compatibility on public interfaces.
- Multi-turn work where accuracy matters more than turnaround.
Standalone test scaffolding, schema formatting, docstrings, and syntax completion go to deepseek-v4-flash.
To test the heuristic, follow the same guide. Give each model the same repository state, instructions, tools, and time limit. Then compare correctness, tests passed, unnecessary changes, total tokens, final cost, and how often a person had to step in. Keep results by task type, because one model may review well and implement poorly.
Estimating a Coding-Agent Loop Cost
Agents resend instructions, history, code, and tool results on every call. The cost guide, observed 2026-10-03, notes that long sessions can cost much more than a single chat request. We worked the arithmetic below from list prices. The result is an estimate, not a measured bill.
Assumptions (ours, not measured): a loop of 20 model calls, each with 30,000 input tokens and 1,500 output tokens. That gives 600,000 input tokens and 30,000 output tokens in total.
The formula is input_tokens / 1M × input price + output_tokens / 1M × output price. Prices come from the table above.
All 20 calls to deepseek-v4-pro:
- Off-peak: 0.6 × $0.66 = $0.396 input, plus 0.03 × $1.98 = $0.0594 output, for $0.4554.
- Peak: 0.6 × $1.32 = $0.792, plus 0.03 × $3.96 = $0.1188, for $0.9108.
All 20 calls to deepseek-v4-flash:
- Off-peak: 0.6 × $0.15 = $0.09, plus 0.03 × $0.60 = $0.018, for $0.108.
- Peak: 0.6 × $0.30 = $0.18, plus 0.03 × $1.20 = $0.036, for $0.216.
Mixed: 4 calls to pro, 16 to flash. Pro carries 120,000 input and 6,000 output tokens. Flash carries 480,000 input and 24,000 output tokens.
- Off-peak: pro is 0.12 × $0.66 + 0.006 × $1.98 = $0.0792 + $0.01188 = $0.09108. Flash is 0.48 × $0.15 + 0.024 × $0.60 = $0.072 + $0.0144 = $0.0864. The total is $0.17748.
- Peak: pro is 0.12 × $1.32 + 0.006 × $3.96 = $0.1584 + $0.02376 = $0.18216. Flash is 0.48 × $0.30 + 0.024 × $1.20 = $0.144 + $0.0288 = $0.1728. The total is $0.35496.
| Scenario | Off-peak estimate | Peak estimate |
|---|---|---|
20 calls on deepseek-v4-pro |
$0.4554 | $0.9108 |
20 calls on deepseek-v4-flash |
$0.1080 | $0.2160 |
| 4 pro + 16 flash | $0.1775 | $0.3550 |
Estimates from list prices observed 2026-10-03 (pro, flash).
Cache variant (off-peak, assumption: 80% of input tokens are cache reads). That means 480,000 cache-read tokens and 120,000 uncached tokens per loop.
- Pro: 0.48 × $0.022 = $0.01056, plus 0.12 × $0.66 = $0.0792, plus $0.0594 output, for $0.14916.
- Flash: 0.48 × $0.003 = $0.00144, plus 0.12 × $0.15 = $0.018, plus $0.018 output, for $0.03744.
This variant bills uncached tokens at the plain input price and ignores cache-write charges on flash, which the record does not list. Confirm cached-token counts in the response or in Usage before you count on this discount. The billing guide also warns that the lowest price per token is not always the lowest cost per completed task, because retries add up.
A Tool-Calling Request for a Coding Agent
Both records list tool-use, and both accept openai_chat_completions. The request below uses only fields from the tool-calling guide (observed 2026-10-03): model, messages, and tools with type: "function". We added max_tokens, which the billing guide lists as a way to cap response length. We left out tool_choice, because that guide documents it only for the Responses format.
curl https://api.tokenlab.sh/v1/chat/completions \
-H "Authorization: Bearer $TOKENLAB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-pro",
"max_tokens": 2000,
"messages": [
{"role": "system", "content": "You are a software engineering assistant."},
{"role": "user", "content": "The pagination test in tests/test_api.py fails. Find the cause."}
],
"tools": [
{
"type": "function",
"function": {
"name": "read_file",
"description": "Read a file from the repository",
"parameters": {
"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"]
}
}
},
{
"type": "function",
"function": {
"name": "run_tests",
"description": "Run the test suite for one path",
"parameters": {
"type": "object",
"properties": {"path": {"type": "string"}},
"required": ["path"]
}
}
}
]
}'
The model returns a function name and arguments in tool_calls. Your backend runs the tool. The loop then runs in five steps:
- Send messages plus tool definitions.
- Read the response for
tool_calls. - Execute the tool in your own backend.
- Append the tool result in the same API format.
- Continue until the model returns a final answer.
The guide does not show the tool-result message shape inline. Take it from the Create Chat Completion reference (/api-reference/chat/create-completion) rather than guessing.
Before executing any call, validate the arguments and apply your own permission checks. Make execution idempotent, because a client retry can repeat the same tool call. Keep one API format for the whole exchange, since the formats represent tool state differently.
Retries, Backoff, and Fallback Between the Two Models
The error guide and rate-limit guide, both observed 2026-10-03, set the policy. Branch on HTTP status and code, never on message.
| Status | Repeat the same request? | Action |
|---|---|---|
400, 401, 402, 403, 404, 413 |
No | Fix the request, key, balance, permissions, or input |
429 |
Yes | Wait for Retry-After; if absent, use exponential backoff with jitter |
500–504 |
Only if retryable is true |
Respect retry_after and cap attempts |
| Connection closed before a response | Sometimes | Retry with care if a tool call could repeat a side effect |
| Stream interrupted after output arrived | No | Treat it as incomplete; a repeat may produce different output or a second charge |
Two cases need extra care. A 503 all_channels_failed or 503 delivery_tier_unavailable is not always temporary. When retryable is false and retry_after is missing, do not repeat the request. Check GET /v1/models before choosing another model. Also, context_length_exceeded will not be fixed by switching between these two models, since both list the same 1,000,000-token input limit.
The code below applies that policy. It sets max_retries=0 so the SDK does not retry behind your back. Each model gets four attempts, and the fallback runs only after the first model exhausts retryable errors.
import os
import random
import time
from openai import OpenAI, APIStatusError, APIConnectionError
client = OpenAI(
api_key=os.environ["TOKENLAB_API_KEY"],
base_url="https://api.tokenlab.sh/v1",
timeout=30.0,
max_retries=0,
)
FALLBACK = {
"deepseek-v4-pro": "deepseek-v4-flash",
"deepseek-v4-flash": "deepseek-v4-pro",
}
def error_fields(exc):
body = getattr(exc, "body", None)
if isinstance(body, dict):
return body.get("error", body)
return {}
def backoff(attempt):
return min(30, 2 ** attempt + random.random())
def retry_delay(exc, attempt):
"""Seconds to wait, or None when the request must not be repeated."""
if isinstance(exc, APIConnectionError):
return backoff(attempt)
fields = error_fields(exc)
header = exc.response.headers.get("Retry-After")
if exc.status_code == 429:
return float(header) if header else backoff(attempt)
if exc.status_code >= 500 and fields.get("retryable") is True:
wait = fields.get("retry_after") or header
return float(wait) if wait else backoff(attempt)
return None
def chat_with_fallback(model, messages, tools=None, attempts=4):
last_exc = None
for candidate in (model, FALLBACK[model]):
kwargs = {"model": candidate, "messages": messages}
if tools:
kwargs["tools"] = tools
for attempt in range(attempts):
try:
return candidate, client.chat.completions.create(**kwargs)
except (APIStatusError, APIConnectionError) as exc:
delay = retry_delay(exc, attempt)
if delay is None:
raise # 4xx or non-retryable 5xx: do not repeat or fall back
last_exc = exc
if attempt < attempts - 1:
time.sleep(delay)
print(f"{candidate} exhausted retries, trying {FALLBACK[candidate]}")
raise last_exc
def pick_model(is_complex):
return "deepseek-v4-pro" if is_complex else "deepseek-v4-flash"
used, response = chat_with_fallback(
pick_model(is_complex=False),
[{"role": "user", "content": "Write a pytest case: an empty list returns 0 for sum_items()."}],
)
print(used, response.choices[0].message.content)
Always log which model answered. The coding-agent guide warns that a fallback may change price, context limit, tool format, or output style, so tell the user when the model changes. Falling back from flash to pro roughly quadruples input cost at list prices, so alert on it. Save the Request ID from the response headers with each call so support can trace a failure.
Read Limits, Formats, and Price Before Routing
The Get a Model reference, observed 2026-10-03, describes GET /v1/models/:model. The response carries a tokenlab object with capabilities, pricing, max_input_tokens, max_output_tokens, accepted_request_formats, and lifecycle. An unknown model returns 404 model_not_found. The billing guide also points to GET /v1/models/:model/pricing for the current price.
import json
import urllib.request
def read_model(model_id):
url = f"https://api.tokenlab.sh/v1/models/{model_id}"
with urllib.request.urlopen(url, timeout=10) as resp:
meta = json.load(resp)["tokenlab"]
return {
"max_input_tokens": meta.get("max_input_tokens"),
"max_output_tokens": meta.get("max_output_tokens"),
"formats": meta.get("accepted_request_formats"),
"capabilities": meta.get("capabilities"),
"lifecycle": meta.get("lifecycle"),
"pricing": meta.get("pricing"),
}
def preflight(model_id, input_tokens):
info = read_model(model_id)
problems = []
if "openai_chat_completions" not in (info["formats"] or []):
problems.append("chat completions not accepted")
if "tool-use" not in (info["capabilities"] or []):
problems.append("no tool-use capability")
if info["max_input_tokens"] and input_tokens > info["max_input_tokens"]:
problems.append("input exceeds max_input_tokens")
return info, problems
for model_id in ("deepseek-v4-pro", "deepseek-v4-flash"):
info, problems = preflight(model_id, input_tokens=30_000)
print(model_id, json.dumps(info, indent=2), problems)
We print lifecycle and pricing raw because this article's evidence does not show their exact JSON layout inside that response. Inspect the output once, then parse the fields you need. The docs advise against hard-coding copied price tables, so run the check at startup or on a schedule. Public discovery endpoints such as GET /v1/models have their own rate limits, so cache the result instead of calling it per request.
For rate limits, the standard User tier allows 1,000 requests per minute per API key, as observed 2026-10-03. The guide says the active configuration may differ. On a 429, trust the returned X-RateLimit-Limit and Retry-After values over any copied number.
FAQ
Can I call deepseek-v4-pro through the Anthropic Messages format?
Yes. Both records list anthropic_messages among the accepted formats, observed 2026-10-03. The coding-agent guide gives the Anthropic Messages base URL as https://api.tokenlab.sh, without the /v1 suffix that Chat Completions uses. Tool schemas differ by format, so keep one format for the whole conversation.
Should I retry a 503 from deepseek-v4-pro or deepseek-v4-flash?
Only if the error body says retryable is true, and then wait for retry_after. A 503 all_channels_failed with retryable: false means the request has no supply in the selected Delivery tier. Repeating it will not help. Check GET /v1/models before you pick another model, as the error guide describes.
Does deepseek-v4.1-flash replace deepseek-v4-flash?
The catalog does not say so. On 2026-10-03 the deepseek-v4-flash record showed no replacement model, and deepseek-v4.1-flash showed active status. The two share limits and list prices, and the newer one adds reasoning and vision capabilities. Test it on your tasks and switch deliberately by model ID.
Do cached tokens make deepseek-v4-flash cheaper in an agent loop?
They can. The record lists an off-peak cache read price of $0.003 per 1M tokens, against $0.15 for plain input. The cost guide says to confirm cached-token usage in the response or in Usage before counting on a discount. Cache behavior and prices differ by model.
Will routing to deepseek-v4-flash raise my rate limit?
No. The rate-limit guide says a faster model does not raise your account's request limit. Model speed, token limits, and account rate limits are separate constraints, and limits apply per API key.
Check the current deepseek-v4-pro and deepseek-v4-flash entries on the TokenLab models page before you wire up your router.
Sources
Prices checked 2026-10-03
- TokenLab Docs: QuickstartSources checked 2026-10-03
- TokenLab Docs: Choose a model for coding agentsSources checked 2026-10-03
- TokenLab Docs: Control coding agent costsSources checked 2026-10-03
- TokenLab Docs: Structured Outputs & Tool CallingSources checked 2026-10-03
- TokenLab Docs: Handle API errorsSources checked 2026-10-03
- TokenLab Docs: Rate limitsSources checked 2026-10-03
- TokenLab Docs: Get a ModelSources checked 2026-10-03
- TokenLab Docs: Billing and pricingSources checked 2026-10-03



