DeepSeek: deepseek-v4.1-flash
deepseek-v4.1-flash- Input / Output
- $0.15 / $0.60
- Context
- 1M
- Max output
- 384K
- Modalities
- VisionChat
- Capabilities
- Tool usePrompt CacheReasoning
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
messagesNative (Messages API)POST/v1/messagesAPI format:curl https://api.tokenlab.sh/v1/messages \ -H "Content-Type: application/json" \ -H "x-api-key: sk-xxx" \ -H "anthropic-version: 2023-06-01" \ -d '{ "model": "deepseek-v4.1-flash", "max_tokens": 1024, "messages": [ {"role": "user", "content": "Hello!"} ] }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Rate window · Off-peak
Per 1M tokens- Official price
- Input $0.15 / Output $0.60 / Cache read $0.003 / Cache write $0.15
- TokenLab price
- Input $0.15 / Output $0.60 / Cache read $0.003 / Cache write $0.15
- Discount
- -
Rate window · Peak
Per 1M tokens- Official price
- Input $0.30 / Output $1.20 / Cache read $0.006 / Cache write $0.30
- TokenLab price
- Input $0.30 / Output $1.20 / Cache read $0.006 / Cache write $0.30
- Discount
- -
Cache read
- Official price
- $0.003
- TokenLab price
- $0.003
- Discount
- -
Cache write
- Official price
- $0.15
- TokenLab price
- $0.15
- Discount
- -
| Official pricePer 1M tokens | TokenLab pricePer 1M tokens | Discount | |
|---|---|---|---|
| Rate window · Off-peak | Input $0.15 / Output $0.60 / Cache read $0.003 / Cache write $0.15 | Input $0.15 / Output $0.60 / Cache read $0.003 / Cache write $0.15 | - |
| Rate window · Peak | Input $0.30 / Output $1.20 / Cache read $0.006 / Cache write $0.30 | Input $0.30 / Output $1.20 / Cache read $0.006 / Cache write $0.30 | - |
| Cache read | $0.003 | $0.003 | - |
| Cache write | $0.15 | $0.15 | - |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hoursNo data yet
- 30-day success rate
- 95.9%
- 7-day median latency
- 4.7 sn=1,318
- 7-day P95 latency
- 22.1 sn=1,318
- 30-day requests
- 1K+
- Last active
- 16 hours ago
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open DeepSeek V4.1 Flash in Console with a prompt ready to edit or send.
Help me try deepseek-v4.1-flash with a short message at /v1/messages. Show the reply, latency, and cost.
Use cases
Best forVision
Reading images, parsing documents, and answering visual questions
Reasoning
Multi-step reasoning, analysis, and research workflows
Agents and tools
Handle reasoning, support triage, tool calls, and multi-step tasks.
Coding
Generate, review, or debug code in the tools you already use.
Knowledge assistants
Build chat, search, and retrieval with a clear price and capability profile.
Side-by-side test
Compare response quality, latency, and price side by side.
Prompt examples
Write a concise support reply and list the assumptions behind it.
Review this API design and call out the top three integration risks.
Turn a long changelog into release notes a non-engineer would read.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
How much does DeepSeek V4.1 Flash cost?
On TokenLab, DeepSeek V4.1 Flash costs Input $0.15 / Output $0.60 Per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What is DeepSeek V4.1 Flash best for?
DeepSeek V4.1 Flash supports JSON mode, Prompt Cache, Tool use. You can open it directly in Create.
How do I test DeepSeek V4.1 Flash?
Open DeepSeek V4.1 Flash in Create. A sample for /v1/messages will be ready to try.
Which endpoint should DeepSeek V4.1 Flash use?
Use https://api.tokenlab.sh/v1/messages for DeepSeek V4.1 Flash. The request example below shows the matching code shape.
Can I test DeepSeek V4.1 Flash before integrating it?
Yes. Open Console starts a ready draft for DeepSeek V4.1 Flash and keeps your prompt after sign-in, so you don’t lose context.
Which operations does DeepSeek V4.1 Flash support?
DeepSeek V4.1 Flash supports messages. Select an operation above to see its endpoint and request example.