DeepSeek: DeepSeek V4 Flash
deepseek-v4-flash- Input / Output
- $0.15 / $0.60
- Context
- 1M
- Released
- Apr 24, 2026
- Max output
- 384K
- Modalities
- VisionChat
- Capabilities
- Tool usePrompt CacheReasoning
About DeepSeek V4 Flash
DeepSeek V4 Flash is the smaller model of DeepSeek's V4 generation, an open-weight mixture-of-experts language model with 284B total and 13B active parameters. DeepSeek positions it for speed and low serving cost while keeping reasoning close to V4 Pro. It runs in thinking and non-thinking modes, calls tools, caches prompts, and answers in JSON on request. It is a text model.
Where it works well
- DeepSeek reports that its reasoning closely matches V4 Pro and that it holds up on basic agent tasks, at a much smaller active-parameter count.
- Thinking can be switched on per request with low, high or max effort, so one model covers quick replies and slower problem solving.
- A very long context window lets an assistant keep a large codebase or a substantial document collection in view across a conversation.
- It supports tool calls, structured JSON output and prompt caching, which helps long agent loops that resend the same instructions.
When to choose another model
- It is a text model, so work that involves screenshots or diagrams belongs on DeepSeek V4.1 Flash or the vision-exp variant.
- DeepSeek says V4 Pro leads on agentic coding benchmarks and world knowledge, so the hardest coding and research tasks fit Pro better.
- In thinking mode the reasoning text comes back separately and has to be passed back on later tool turns, which agent code must handle.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textResponses APIPOST/v1/responsesAPI format:curl https://api.tokenlab.sh/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxx" \ -d '{ "model": "deepseek-v4-flash", "input": "Hello!" }'
Pricing
TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.
Rate window · Off-peak
per 1M tokens- Official price
- Input $0.15 / Output $0.60 / Cache read $0.003
- Official
- Input $0.15 / Output $0.60 / Cache read $0.003
- Discount
- —
Rate window · Peak
per 1M tokens- Official price
- Input $0.30 / Output $1.20 / Cache read $0.006
- Official
- Input $0.30 / Output $1.20 / Cache read $0.006
- Discount
- —
Cache read
- Official price
- $0.003
- Official
- $0.003
- Discount
- —
| Official priceper 1M tokens | Officialper 1M tokens | Discount | |
|---|---|---|---|
| Rate window · Off-peak | Input $0.15 / Output $0.60 / Cache read $0.003 | Input $0.15 / Output $0.60 / Cache read $0.003 | — |
| Rate window · Peak | Input $0.30 / Output $1.20 / Cache read $0.006 | Input $0.30 / Output $1.20 / Cache read $0.006 | — |
| Prompt cache pricing | |||
| Cache read | $0.003 | $0.003 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
- 30-day success rate
- 99.5%
- 30-day requests
- 100+
- Last active
- 12 hours ago
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open DeepSeek V4 Flash in Console with a prompt ready to edit or send.
Help me try deepseek-v4-flash with a short message at /v1/responses. Show the reply, latency, and cost.
Use cases
Best for- Reasoning
- Vision
High-volume agent steps
Use it for the many small decisions inside an agent loop, such as choosing a tool, reading its result and planning the next call, where response time and cost add up.
Whole-repository questions
Paste a large codebase or a pile of documents into the long context and ask where a behavior is implemented or which files a change would touch.
Structured extraction
Turn contracts, tickets or logs into JSON that follows a schema, with thinking left off so each record returns quickly.
Drop-in replacement for older DeepSeek names
Point existing clients that used the retired chat and reasoner names at deepseek-v4-flash, then choose thinking or non-thinking mode per request.
Prompt examples
The full source of our billing service follows. Which modules read the invoice status field, and what breaks if I add a new status?
Extract every renewal date, notice period and cancellation fee from these five contracts and return one JSON array that matches this schema.
You can call search_docs and open_ticket. A user reports that webhooks stopped arriving yesterday. Investigate with the tools, then summarize the likely cause.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What is the difference between DeepSeek V4 Flash and V4 Pro?
Flash is the smaller model, with 13B active parameters against 49B for Pro. DeepSeek says Flash reasons almost as well as Pro and matches it on simple agent tasks, while Pro is stronger on agentic coding and world knowledge. Start on Flash and move a task to Pro when it fails.
Does DeepSeek V4 Flash support thinking mode?
Yes. Enable thinking in the request and pick low, high or max reasoning effort; high is the default. The reasoning arrives in a separate field, and tool-using conversations must send it back on each following turn. Leave thinking off for short, latency-sensitive answers.
Can DeepSeek V4 Flash read images?
No. It is a text model: text goes in and text comes out. DeepSeek V4.1 Flash and the V4 Flash vision-exp variant are separate models that understand images, so use one of those when the prompt includes screenshots or diagrams.
What is the long context useful for?
It lets one request carry a whole repository or a pile of documents, so the model can answer where a behavior is implemented or which files a change touches without retrieval in between. Pair it with prompt caching when the same material is sent repeatedly.
What happened to the old deepseek-chat and deepseek-reasoner names?
DeepSeek discontinued them on 24 July 2026. During the transition they mapped to V4 Flash in non-thinking and thinking modes. New code should call deepseek-v4-flash and set the thinking option, rather than relying on separate names for each mode.
How much does DeepSeek V4 Flash cost?
On TokenLab, DeepSeek V4 Flash costs Input $0.15 / Output $0.60 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of DeepSeek V4 Flash?
DeepSeek V4 Flash accepts up to 1,000,000 tokens of context and returns up to 384,000 tokens in one response.
Which endpoint should DeepSeek V4 Flash use?
Use https://api.tokenlab.sh/v1/responses for DeepSeek V4 Flash. The request example below shows the matching code shape.
Which operations does DeepSeek V4 Flash support?
DeepSeek V4 Flash supports Text to text. Select an operation above to see its endpoint and request example.
Compare DeepSeek V4 Flash
Guides that use DeepSeek V4 Flash
Sources
Reviewed Oct 2, 2026