Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

DeepSeek: DeepSeek V4 Pro

DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes
Compare models
deepseek-v4-pro
AvailableDeepSeekChat
Input / Output
$0.66 / $1.98
Context
1M
Released
Apr 24, 2026
Max output
384K
Modalities
Chat
Capabilities
Tool usePrompt CacheReasoning

About DeepSeek V4 Pro

DeepSeek V4 Pro is the larger model of DeepSeek's V4 generation, an open-weight mixture-of-experts model with 1.6T total and 49B active parameters. DeepSeek aims it at agentic coding, math, STEM reasoning and broad world knowledge, where it wants to rival closed frontier models. It offers thinking and non-thinking modes with three effort levels, tool calls and JSON output, and reads text only.

Where it works well

  • DeepSeek calls it open-source state of the art on agentic coding benchmarks and says its world knowledge trails only Gemini 3.1 Pro.
  • Thinking mode has low, high and max effort, so hard math, STEM and debugging problems can be given far more reasoning than routine ones.
  • Tool calls, JSON output and prompt caching are supported, so it fits agent loops that repeat a long system prompt.
  • Weights are open, so teams that also self-host can evaluate the same model family before committing.

When to choose another model

  • It reads text only; screenshot and diagram work needs deepseek-v4.1-flash or the Flash vision-exp variant.
  • It is the heavier V4 model, so V4 Flash answers faster when a task does not need the extra reasoning depth.
  • Max effort produces long reasoning that adds delay and tokens, and tool conversations must pass the reasoning text back each turn.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textResponses API
    POST/v1/responses
    API format:
    curl https://api.tokenlab.sh/v1/responses \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer sk-xxx" \
      -d '{
        "model": "deepseek-v4-pro",
        "input": "Hello!"
      }'

Pricing

TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.

Rate window · Off-peak

per 1M tokens
Official price
Input $0.66 / Output $1.98 / Cache read $0.022 / Cache write $0.66
Official
Input $0.66 / Output $1.98 / Cache read $0.022 / Cache write $0.66
Discount
—

Rate window · Peak

per 1M tokens
Official price
Input $1.32 / Output $3.96 / Cache read $0.044 / Cache write $1.32
Official
Input $1.32 / Output $3.96 / Cache read $0.044 / Cache write $1.32
Discount
—
Prompt cache pricing

Cache read

Official price
$0.022
Official
$0.022
Discount
—

Cache write

Official price
$0.66
Official
$0.66
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
30-day success rate
99.9%
30-day requests
100+
Last active
53 seconds ago

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open DeepSeek V4 Pro in Console with a prompt ready to edit or send.

Help me try deepseek-v4-pro with a short message at /v1/responses. Show the reply, latency, and cost.

Use cases

Best for
  • Reasoning
  • Autonomous coding agents

    Run it as the planner in a terminal or IDE agent that edits many files, runs tests and reacts to failures over a long session.

  • Hard technical reasoning

    Give it proofs, algorithm design, scientific derivations or tricky concurrency bugs and set reasoning effort to max for the final pass.

  • Knowledge-heavy question answering

    Ask questions that need broad factual recall across fields, such as comparing standards or explaining how two technologies interact.

  • Escalation tier behind Flash

    Send routine requests to V4 Flash and pass only the ones that fail validation or need deeper analysis to Pro.

Prompt examples

This backend deadlocks under load. The thread dump and the two lock-acquiring modules are attached. Find the ordering problem and write the patch.

Prove that the algorithm below terminates and state its worst-case complexity, then point out any input that breaks the argument.

Plan a migration from our REST handlers to the new gateway, list the affected modules in order, and run the tests after each step using the shell tool.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

When should I use DeepSeek V4 Pro instead of V4 Flash?

Use Pro for agentic coding, difficult reasoning and questions that need wide factual knowledge. DeepSeek says Flash comes close on reasoning and simple agent work, so Pro matters most when Flash gets a task wrong or when a mistake is expensive.

How do the thinking effort levels work on DeepSeek V4 Pro?

You enable thinking and choose low, high or max. High is the default and suits everyday agent tasks, low suits simple ones, and max is for the most complex problems. Higher effort means more reasoning text and a longer wait before the answer.

Is DeepSeek V4 Pro open source?

DeepSeek released the V4 generation with open weights, and the Pro model has 1.6T total parameters with 49B active per token. Calling it through the API does not require hosting anything; the open release matters if you want to evaluate or self-host the same family.

Does DeepSeek V4 Pro accept images?

No. It reads and writes text only. DeepSeek ships picture understanding in separate Flash-based models, so a workflow that needs both screenshots and hard reasoning has to split the steps across two models.

Does DeepSeek V4 Pro work with the OpenAI SDK or the Anthropic SDK?

Both. DeepSeek documents compatibility with the OpenAI and Anthropic API formats, so you can keep your existing client and change the model name and base URL.

How much does DeepSeek V4 Pro cost?

On TokenLab, DeepSeek V4 Pro costs Input $0.66 / Output $1.98 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of DeepSeek V4 Pro?

DeepSeek V4 Pro accepts up to 1,000,000 tokens of context and returns up to 384,000 tokens in one response.

Which endpoint should DeepSeek V4 Pro use?

Use https://api.tokenlab.sh/v1/responses for DeepSeek V4 Pro. The request example below shows the matching code shape.

Which operations does DeepSeek V4 Pro support?

DeepSeek V4 Pro supports Text to text. Select an operation above to see its endpoint and request example.

Compare DeepSeek V4 Pro

Guides that use DeepSeek V4 Pro

Sources

Reviewed Oct 2, 2026

More from DeepSeek V4

Related models