Alibaba: Qwen3.5-Flash
qwen3.5-flash- Input / Output
- $0.10 / $0.40
- Context
- 992K
- Max output
- 64K
- Modalities
- Chat
- Capabilities
- Prompt Cache
About Qwen3.5-Flash
Qwen3.5-Flash is the lighter hosted tier of Alibaba's Qwen3.5 family, built for faster and cheaper text processing than Qwen3.5-Plus. It has a very large context and prompt caching, which fits summarization and repeated passes over the same large source. It shares the family's broad language coverage.
Where it works well
- Prompt caching helps repeated consultation of one large source.
- Tuned for speed over the Plus tier.
- Shares the Qwen3.5 family's 201-language coverage.
- Good fit for pipeline steps such as summarizing or tagging.
When to choose another model
- Qwen3.5-Plus is better on harder analysis.
- It takes text only, so it cannot replace a vision model.
- Older than Qwen3.8-Flash, which reports agentic coding results.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textResponses APIPOST/v1/responsesAPI format:curl https://api.tokenlab.sh/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxx" \ -d '{ "model": "qwen3.5-flash", "input": "Hello!" }'
Pricing
TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.
Input tokens <= 125K
per 1M tokens- Official price
- Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
- Official
- Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
- Discount
- —
Input tokens <= 250K
per 1M tokens- Official price
- Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
- Official
- Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
- Discount
- —
Input tokens <= 1M
per 1M tokens- Official price
- Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
- Official
- Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
- Discount
- —
Cache read
- Official price
- $0.01
- Official
- $0.01
- Discount
- —
Cache write
- Official price
- $0.125
- Official
- $0.125
- Discount
- —
| Official priceper 1M tokens | Officialper 1M tokens | Discount | |
|---|---|---|---|
| Input tokens <= 125K | Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40 | Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40 | — |
| Input tokens <= 250K | Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40 | Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40 | — |
| Input tokens <= 1M | Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40 | Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40 | — |
| Prompt cache pricing | |||
| Cache read | $0.01 | $0.01 | — |
| Cache write | $0.125 | $0.125 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Qwen3.5-Flash in Console with a prompt ready to edit or send.
Help me try qwen3.5-flash with a short message at /v1/responses. Show the reply, latency, and cost.
Use cases
Batch summarization
Condense long reports in a nightly pipeline, with the same large source cached across the sections you ask about.
Document tagging and routing
Label long text by topic, urgency, or owner and return the result in a fixed format that downstream code can parse.
Repeated Q&A over one source
Cache a manual or policy once and ask many questions against it without paying to reprocess the whole text every turn.
Draft cleanup
Fix grammar, tone, and structure in long drafts while keeping the author's claims and figures untouched.
Prompt examples
Summarize each section of this handbook in two lines.
Tag this ticket with product area and urgency and return JSON.
Answer questions about the pasted policy; reply 'not stated' if absent.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What is Qwen3.5-Flash exactly?
It is the lighter hosted tier of Alibaba's Qwen3.5 family, aimed at fast processing of long text. It sits below Qwen3.5-Plus in strength and is meant for repeated, routine document steps.
When should I pick Flash over Qwen3.5-Plus?
Pick Flash when speed and cost matter more than depth, such as bulk summarization, tagging, or Q&A over a fixed source. Move to Plus when answers need more careful analysis across many passages.
Does Qwen3.5-Flash support prompt caching?
Yes. Prompt caching helps when the same long source is sent again and again, as in repeated document-processing steps where only the question changes between calls.
Can Qwen3.5-Flash read images?
No. Qwen3.5-Flash takes text only. Use Qwen3.8-Flash, Qwen3.7-Plus, or Qwen3-VL Plus when screenshots, scans, or charts are part of the task, and send extracted text to this model instead.
Is there a newer Flash model from Alibaba?
Yes. Qwen3.8-Flash is a newer multimodal mixture-of-experts model aimed at agentic coding and long-horizon agent work. Keep Qwen3.5-Flash for plain text pipelines that already perform well.
How much does Qwen3.5-Flash cost?
On TokenLab, Qwen3.5-Flash costs Input $0.10 / Output $0.40 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of Qwen3.5-Flash?
Qwen3.5-Flash accepts up to 991,808 tokens of context and returns up to 65,536 tokens in one response.
Which endpoint should Qwen3.5-Flash use?
Use https://api.tokenlab.sh/v1/responses for Qwen3.5-Flash. The request example below shows the matching code shape.
Which operations does Qwen3.5-Flash support?
Qwen3.5-Flash supports Text to text. Select an operation above to see its endpoint and request example.
Compare Qwen3.5-Flash
Sources
Reviewed Oct 2, 2026