Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Alibaba: Fun-ASR Flash 2026-06-15

Fun-ASR Flash 2026-06-15 is a dated speech-recognition release for audio-file transcription. Pinning this version is useful when a transcription pipeline needs a consistent model choice across repeated processing runs.
Compare models
fun-asr-flash-2026-06-15
AvailableAlibabaSpeech to TextSync
Price
From $0.000035
Modalities
Audio

About Fun-ASR Flash 2026-06-15

Fun-ASR Flash 2026-06-15 is a dated release of Alibaba's Flash speech recognition model, surfaced by Alibaba as Qwen-Audio-3.0-ASR-Flash. It transcribes short audio files over a synchronous HTTP call and takes prior conversation turns as context to bias recognition. Pinning the dated ID keeps a pipeline on one version, unlike an undated alias that may be updated.

Where it works well

  • Accepts earlier conversation turns alongside the audio, so vocabulary from the ongoing discussion steers recognition.
  • Alibaba lists word-level timestamps and sentence boundaries in the same synchronous response.
  • The date in the ID pins one release, which keeps repeated runs comparable.

When to choose another model

  • Clips are capped at about five minutes, so long recordings need the asynchronous Fun-ASR or a filetrans model.
  • It handles finished clips only; for live captions or dictation, Fun-ASR Realtime is the streaming model.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Speech to textTokenLab endpoint
    POST/v1/audio/transcriptions
    curl -X POST "https://api.tokenlab.sh/v1/audio/transcriptions" \
      -H "Authorization: Bearer sk-xxx" \
      -F file="@audio.mp3" \
      -F operation="stt" \
      -F response_format="json" \
      -F model="fun-asr-flash-2026-06-15" \
      -F language="en"

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Speech to text

per second
Official price
$0.000035
Official
$0.000035
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Fun-ASR Flash 2026-06-15 in Console with a prompt ready to edit or send.

Help me try fun-asr-flash-2026-06-15 with a short audio request at /v1/audio/transcriptions. Show the result and cost.

Use cases

  • Voice notes

    Transcribe short recorded notes and messages in a single request, without queueing an asynchronous job.

  • Medical dictation

    Pass the preceding dialogue as context so drug names and clinical terms are more likely to be recognised.

  • Pinned test pipelines

    Hold the version constant while comparing downstream summarisation changes against the same transcripts.

Prompt examples

Transcribe this 3-minute clinic recording; context: the previous turn discussed metformin dosage.

Transcribe this field-engineer voice note about a PLC fault and return word timestamps.

Transcribe this short customer-call clip, using the earlier chat turns as context.

FAQ

What does the date in fun-asr-flash-2026-06-15 mean?

It identifies a specific release, 15 June 2026. Calling the dated ID keeps your pipeline on that release. Alibaba presents it under the name Qwen-Audio-3.0-ASR-Flash.

How long can an audio clip be?

Alibaba documents a limit of about five minutes per call, with files up to 2 GB. For longer recordings use Fun-ASR or Qwen3 ASR Flash Filetrans, both of which take asynchronous file jobs.

Can I give it context about the conversation?

Yes. It uses a chat-style message format, so you can pass earlier turns before the audio. The model uses them to favour vocabulary that fits the topic, which helps with specialist terms.

Does it return timestamps?

Yes. Alibaba describes word-level timestamps and sentence boundaries in the response, so you can align the text to the audio for subtitles, review, or search without a separate alignment step.

How much does Fun-ASR Flash 2026-06-15 cost?

On TokenLab, Fun-ASR Flash 2026-06-15 costs $0.000035 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Fun-ASR Flash 2026-06-15 use?

Use https://api.tokenlab.sh/v1/audio/transcriptions for Fun-ASR Flash 2026-06-15. The request example below shows the matching code shape.

Which operations does Fun-ASR Flash 2026-06-15 support?

Fun-ASR Flash 2026-06-15 supports Speech to text. Select an operation above to see its endpoint and request example.

Compare Fun-ASR Flash 2026-06-15

Sources

Reviewed Oct 2, 2026

More from Fun-ASR

Related models