DeepSeek: DeepSeek V4 Pro
deepseek-v4-pro- Input / Output
- $0.66 / $1.98
- Context
- 1M
- Released
- Apr 24, 2026
- Max output
- 384K
- Modalities
- Chat
- Capabilities
- Tool usePrompt CacheReasoning
About DeepSeek V4 Pro
DeepSeek V4 Pro is the larger model of DeepSeek's V4 generation, an open-weight mixture-of-experts model with 1.6T total and 49B active parameters. DeepSeek aims it at agentic coding, math, STEM reasoning and broad world knowledge, where it wants to rival closed frontier models. It offers thinking and non-thinking modes with three effort levels, tool calls and JSON output, and reads text only.
Where it works well
- DeepSeek calls it open-source state of the art on agentic coding benchmarks and says its world knowledge trails only Gemini 3.1 Pro.
- Thinking mode has low, high and max effort, so hard math, STEM and debugging problems can be given far more reasoning than routine ones.
- Tool calls, JSON output and prompt caching are supported, so it fits agent loops that repeat a long system prompt.
- Weights are open, so teams that also self-host can evaluate the same model family before committing.
When to choose another model
- It reads text only; screenshot and diagram work needs deepseek-v4.1-flash or the Flash vision-exp variant.
- It is the heavier V4 model, so V4 Flash answers faster when a task does not need the extra reasoning depth.
- Max effort produces long reasoning that adds delay and tokens, and tool conversations must pass the reasoning text back each turn.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textResponses APIPOST/v1/responsesAPI format:curl https://api.tokenlab.sh/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxx" \ -d '{ "model": "deepseek-v4-pro", "input": "Hello!" }'
Pricing
TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.
Rate window · Off-peak
per 1M tokens- Official price
- Input $0.66 / Output $1.98 / Cache read $0.022 / Cache write $0.66
- Official
- Input $0.66 / Output $1.98 / Cache read $0.022 / Cache write $0.66
- Discount
- —
Rate window · Peak
per 1M tokens- Official price
- Input $1.32 / Output $3.96 / Cache read $0.044 / Cache write $1.32
- Official
- Input $1.32 / Output $3.96 / Cache read $0.044 / Cache write $1.32
- Discount
- —
Cache read
- Official price
- $0.022
- Official
- $0.022
- Discount
- —
Cache write
- Official price
- $0.66
- Official
- $0.66
- Discount
- —
| Official priceper 1M tokens | Officialper 1M tokens | Discount | |
|---|---|---|---|
| Rate window · Off-peak | Input $0.66 / Output $1.98 / Cache read $0.022 / Cache write $0.66 | Input $0.66 / Output $1.98 / Cache read $0.022 / Cache write $0.66 | — |
| Rate window · Peak | Input $1.32 / Output $3.96 / Cache read $0.044 / Cache write $1.32 | Input $1.32 / Output $3.96 / Cache read $0.044 / Cache write $1.32 | — |
| Prompt cache pricing | |||
| Cache read | $0.022 | $0.022 | — |
| Cache write | $0.66 | $0.66 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
- 30-day success rate
- 99.9%
- 30-day requests
- 100+
- Last active
- 53 seconds ago
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open DeepSeek V4 Pro in Console with a prompt ready to edit or send.
Help me try deepseek-v4-pro with a short message at /v1/responses. Show the reply, latency, and cost.
Use cases
Best for- Reasoning
Autonomous coding agents
Run it as the planner in a terminal or IDE agent that edits many files, runs tests and reacts to failures over a long session.
Hard technical reasoning
Give it proofs, algorithm design, scientific derivations or tricky concurrency bugs and set reasoning effort to max for the final pass.
Knowledge-heavy question answering
Ask questions that need broad factual recall across fields, such as comparing standards or explaining how two technologies interact.
Escalation tier behind Flash
Send routine requests to V4 Flash and pass only the ones that fail validation or need deeper analysis to Pro.
Prompt examples
This backend deadlocks under load. The thread dump and the two lock-acquiring modules are attached. Find the ordering problem and write the patch.
Prove that the algorithm below terminates and state its worst-case complexity, then point out any input that breaks the argument.
Plan a migration from our REST handlers to the new gateway, list the affected modules in order, and run the tests after each step using the shell tool.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
When should I use DeepSeek V4 Pro instead of V4 Flash?
Use Pro for agentic coding, difficult reasoning and questions that need wide factual knowledge. DeepSeek says Flash comes close on reasoning and simple agent work, so Pro matters most when Flash gets a task wrong or when a mistake is expensive.
How do the thinking effort levels work on DeepSeek V4 Pro?
You enable thinking and choose low, high or max. High is the default and suits everyday agent tasks, low suits simple ones, and max is for the most complex problems. Higher effort means more reasoning text and a longer wait before the answer.
Is DeepSeek V4 Pro open source?
DeepSeek released the V4 generation with open weights, and the Pro model has 1.6T total parameters with 49B active per token. Calling it through the API does not require hosting anything; the open release matters if you want to evaluate or self-host the same family.
Does DeepSeek V4 Pro accept images?
No. It reads and writes text only. DeepSeek ships picture understanding in separate Flash-based models, so a workflow that needs both screenshots and hard reasoning has to split the steps across two models.
Does DeepSeek V4 Pro work with the OpenAI SDK or the Anthropic SDK?
Both. DeepSeek documents compatibility with the OpenAI and Anthropic API formats, so you can keep your existing client and change the model name and base URL.
How much does DeepSeek V4 Pro cost?
On TokenLab, DeepSeek V4 Pro costs Input $0.66 / Output $1.98 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of DeepSeek V4 Pro?
DeepSeek V4 Pro accepts up to 1,000,000 tokens of context and returns up to 384,000 tokens in one response.
Which endpoint should DeepSeek V4 Pro use?
Use https://api.tokenlab.sh/v1/responses for DeepSeek V4 Pro. The request example below shows the matching code shape.
Which operations does DeepSeek V4 Pro support?
DeepSeek V4 Pro supports Text to text. Select an operation above to see its endpoint and request example.
Compare DeepSeek V4 Pro
Guides that use DeepSeek V4 Pro
Sources
Reviewed Oct 2, 2026