Alibaba: Qwen Flash
qwen-flash- Input / Output
- $0.05 / $0.40
- Context
- 998K
- Max output
- 32K
- Modalities
- Chat
- Capabilities
- Prompt CacheReasoning
About Qwen Flash
Qwen Flash is Alibaba's fast, low-cost general text model in the Qwen line, aimed at high-volume work such as summarization, rewriting and simple assistants. It sits below Qwen Plus and Qwen Max in capability and is the one to try first when volume matters more than depth. Alibaba lists it among its legacy Qwen names, with newer Qwen3 generations recommended for new projects.
Where it works well
- Summarizing and reformatting large amounts of text at low latency per request.
- Keeping a lengthy source document available in the conversation while it summarizes or answers.
- An optional thinking mode lets you spend reasoning only on the requests that need it.
- Prompt caching reduces repeat work when a long system prompt or document is reused.
When to choose another model
- Text input only, with no tool calling.
- Reasoning depth and writing quality trail Qwen Plus and Qwen Max on hard analysis tasks.
- Alibaba lists it as a legacy name, so a current Qwen3 model may suit a new build better.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textResponses APIPOST/v1/responsesAPI format:curl https://api.tokenlab.sh/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxx" \ -d '{ "model": "qwen-flash", "input": "Hello!" }'
Pricing
TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.
Input tokens <= 125K
per 1M tokens- Official price
- Input $0.05 / Output $0.40 / Cache read $0.01 / Cache write $0.0625 / Text Input $0.05 / Text Output $0.40
- Official
- Input $0.05 / Output $0.40 / Cache read $0.01 / Cache write $0.0625 / Text Input $0.05 / Text Output $0.40
- Discount
- —
Input tokens <= 250K
per 1M tokens- Official price
- Input $0.05 / Output $0.40 / Cache read $0.01 / Cache write $0.0625 / Text Input $0.05 / Text Output $0.40
- Official
- Input $0.05 / Output $0.40 / Cache read $0.01 / Cache write $0.0625 / Text Input $0.05 / Text Output $0.40
- Discount
- —
Input tokens <= 1M
per 1M tokens- Official price
- Input $0.25 / Output $2.00 / Cache read $0.05 / Cache write $0.3125 / Text Input $0.25 / Text Output $2.00
- Official
- Input $0.25 / Output $2.00 / Cache read $0.05 / Cache write $0.3125 / Text Input $0.25 / Text Output $2.00
- Discount
- —
Cache read
- Official price
- $0.01
- Official
- $0.01
- Discount
- —
Cache write
- Official price
- $0.0625
- Official
- $0.0625
- Discount
- —
| Official priceper 1M tokens | Officialper 1M tokens | Discount | |
|---|---|---|---|
| Input tokens <= 125K | Input $0.05 / Output $0.40 / Cache read $0.01 / Cache write $0.0625 / Text Input $0.05 / Text Output $0.40 | Input $0.05 / Output $0.40 / Cache read $0.01 / Cache write $0.0625 / Text Input $0.05 / Text Output $0.40 | — |
| Input tokens <= 250K | Input $0.05 / Output $0.40 / Cache read $0.01 / Cache write $0.0625 / Text Input $0.05 / Text Output $0.40 | Input $0.05 / Output $0.40 / Cache read $0.01 / Cache write $0.0625 / Text Input $0.05 / Text Output $0.40 | — |
| Input tokens <= 1M | Input $0.25 / Output $2.00 / Cache read $0.05 / Cache write $0.3125 / Text Input $0.25 / Text Output $2.00 | Input $0.25 / Output $2.00 / Cache read $0.05 / Cache write $0.3125 / Text Input $0.25 / Text Output $2.00 | — |
| Prompt cache pricing | |||
| Cache read | $0.01 | $0.01 | — |
| Cache write | $0.0625 | $0.0625 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Qwen Flash in Console with a prompt ready to edit or send.
Help me try qwen-flash with a short message at /v1/responses. Show the reply, latency, and cost.
Use cases
Best for- Reasoning
Bulk summarization
Condense thousands of articles, call transcripts or reviews into short summaries, with the same prompt applied to every item.
Content transformation
Convert text between tones, languages or formats, such as turning notes into a polished email or a table into prose.
Lightweight chat assistants
Back an FAQ bot or in-app helper where answers are short and response time matters more than nuance.
Long-document Q&A
Paste a long manual or policy set into the prompt and answer staff questions directly from it.
Prompt examples
Rewrite the following product description in a friendly tone, 60 words maximum, and keep the model number unchanged.
Summarize this meeting transcript as decisions, action items with owners, and unresolved questions.
Using only the policy text above, answer: can a contractor claim travel expenses? Quote the clause you used.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What is Qwen Flash best for?
High-volume, everyday text tasks: summaries, rewrites, translation-style transformations and simple assistants. It is the quick, inexpensive tier of the Qwen line, so use it when throughput matters and the task does not need deep reasoning.
How is Qwen Flash different from Qwen Plus?
Flash is the speed-oriented tier and Plus is the balanced one. Plus gives stronger analysis and writing; Flash gets simple jobs done faster and more cheaply. Start on Flash and move specific tasks to Plus if the answers fall short.
Does Qwen Flash support thinking mode?
Yes. Thinking can be switched on for requests that need step-by-step reasoning and left off for simple ones. Alibaba's documentation lists thinking mode as supported for this model.
Can Qwen Flash take images or call tools?
No. It is a text-in, text-out model without image input or tool calling. Pick a Qwen vision model or another family if the request needs either.
Is Qwen Flash a good choice for a new project?
Existing pipelines can stay on it, and it remains usable for stable high-volume jobs. For a fresh build, Alibaba points to newer Qwen3 generations, so compare it with Qwen3.8 Flash first.
How much does Qwen Flash cost?
On TokenLab, Qwen Flash costs Input $0.05 / Output $0.40 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of Qwen Flash?
Qwen Flash accepts up to 997,952 tokens of context and returns up to 32,768 tokens in one response.
Which endpoint should Qwen Flash use?
Use https://api.tokenlab.sh/v1/responses for Qwen Flash. The request example below shows the matching code shape.
Which operations does Qwen Flash support?
Qwen Flash supports Text to text. Select an operation above to see its endpoint and request example.
Compare Qwen Flash
Sources
Reviewed Oct 2, 2026