Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Alibaba: Qwen3.5-Flash

Qwen3.5 Flash is the efficiency-oriented option for long-context text work in the 3.5 line. It can support summarization and repeated document-processing steps where the same large source needs to be consulted throughout a task.
Compare models
qwen3.5-flash
AvailableAlibabaChat
Input / Output
$0.10 / $0.40
Context
992K
Max output
64K
Modalities
Chat
Capabilities
Prompt Cache

About Qwen3.5-Flash

Qwen3.5-Flash is the lighter hosted tier of Alibaba's Qwen3.5 family, built for faster and cheaper text processing than Qwen3.5-Plus. It has a very large context and prompt caching, which fits summarization and repeated passes over the same large source. It shares the family's broad language coverage.

Where it works well

  • Prompt caching helps repeated consultation of one large source.
  • Tuned for speed over the Plus tier.
  • Shares the Qwen3.5 family's 201-language coverage.
  • Good fit for pipeline steps such as summarizing or tagging.

When to choose another model

  • Qwen3.5-Plus is better on harder analysis.
  • It takes text only, so it cannot replace a vision model.
  • Older than Qwen3.8-Flash, which reports agentic coding results.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textResponses API
    POST/v1/responses
    API format:
    curl https://api.tokenlab.sh/v1/responses \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer sk-xxx" \
      -d '{
        "model": "qwen3.5-flash",
        "input": "Hello!"
      }'

Pricing

TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.

Input tokens <= 125K

per 1M tokens
Official price
Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
Official
Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
Discount
—

Input tokens <= 250K

per 1M tokens
Official price
Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
Official
Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
Discount
—

Input tokens <= 1M

per 1M tokens
Official price
Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
Official
Input $0.10 / Output $0.40 / Cache read $0.01 / Cache write $0.125 / Text Input $0.10 / Text Output $0.40
Discount
—
Prompt cache pricing

Cache read

Official price
$0.01
Official
$0.01
Discount
—

Cache write

Official price
$0.125
Official
$0.125
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Qwen3.5-Flash in Console with a prompt ready to edit or send.

Help me try qwen3.5-flash with a short message at /v1/responses. Show the reply, latency, and cost.

Use cases

  • Batch summarization

    Condense long reports in a nightly pipeline, with the same large source cached across the sections you ask about.

  • Document tagging and routing

    Label long text by topic, urgency, or owner and return the result in a fixed format that downstream code can parse.

  • Repeated Q&A over one source

    Cache a manual or policy once and ask many questions against it without paying to reprocess the whole text every turn.

  • Draft cleanup

    Fix grammar, tone, and structure in long drafts while keeping the author's claims and figures untouched.

Prompt examples

Summarize each section of this handbook in two lines.

Tag this ticket with product area and urgency and return JSON.

Answer questions about the pasted policy; reply 'not stated' if absent.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

What is Qwen3.5-Flash exactly?

It is the lighter hosted tier of Alibaba's Qwen3.5 family, aimed at fast processing of long text. It sits below Qwen3.5-Plus in strength and is meant for repeated, routine document steps.

When should I pick Flash over Qwen3.5-Plus?

Pick Flash when speed and cost matter more than depth, such as bulk summarization, tagging, or Q&A over a fixed source. Move to Plus when answers need more careful analysis across many passages.

Does Qwen3.5-Flash support prompt caching?

Yes. Prompt caching helps when the same long source is sent again and again, as in repeated document-processing steps where only the question changes between calls.

Can Qwen3.5-Flash read images?

No. Qwen3.5-Flash takes text only. Use Qwen3.8-Flash, Qwen3.7-Plus, or Qwen3-VL Plus when screenshots, scans, or charts are part of the task, and send extracted text to this model instead.

Is there a newer Flash model from Alibaba?

Yes. Qwen3.8-Flash is a newer multimodal mixture-of-experts model aimed at agentic coding and long-horizon agent work. Keep Qwen3.5-Flash for plain text pipelines that already perform well.

How much does Qwen3.5-Flash cost?

On TokenLab, Qwen3.5-Flash costs Input $0.10 / Output $0.40 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of Qwen3.5-Flash?

Qwen3.5-Flash accepts up to 991,808 tokens of context and returns up to 65,536 tokens in one response.

Which endpoint should Qwen3.5-Flash use?

Use https://api.tokenlab.sh/v1/responses for Qwen3.5-Flash. The request example below shows the matching code shape.

Which operations does Qwen3.5-Flash support?

Qwen3.5-Flash supports Text to text. Select an operation above to see its endpoint and request example.

Compare Qwen3.5-Flash

Sources

Reviewed Oct 2, 2026

More from Qwen 3 / QwQ

Related models