Alibaba: Fun-ASR Realtime
fun-asr-realtime- Price
- From $0.00009
- Modalities
- Audio
About Fun-ASR Realtime
Fun-ASR Realtime is Alibaba's streaming speech recognition model: it receives audio over a live session and returns text as people speak. It is the Fun-ASR variant for captions, voice input, and in-call assistance, where the recorded-audio Fun-ASR would force users to wait until the recording ends. Alibaba lists hotwords, timestamps, and Mandarin with dialects such as Cantonese and Sichuanese.
Where it works well
- Streams partial text while the speaker is still talking, so captions appear without waiting for a finished recording.
- Recognises Mandarin with high accuracy and several dialects, including Cantonese and Sichuanese, per Alibaba's documentation.
- Custom hotwords steer live recognition toward names and terms specific to your product.
- Word- and sentence-level timestamps arrive with the stream, which helps align captions.
When to choose another model
- It needs a WebSocket session, which is more integration work than uploading a recording to Fun-ASR.
- Live decoding cannot revisit earlier audio, so for archival accuracy the recorded-audio Fun-ASR is the safer pick.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Speech to textWebSocketWS/v1/realtime// npm install ws import WebSocket from 'ws'; const socket = new WebSocket("wss://api.tokenlab.sh/v1/realtime?model=fun-asr-realtime", { headers: { Authorization: "Bearer sk-xxx" } }); socket.on('open', () => { // Send the input events supported by this model with socket.send(...). // https://tokenlab.sh/docs/en/api-reference/realtime/connect }); socket.on('message', (data) => console.log(data.toString())); socket.on('error', (error) => console.error(error)); socket.on('close', (code, reason) => console.log(code, reason.toString())); // Close the session when finished. process.on('SIGINT', () => socket.close());
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Realtime / Speech to text
per second- Official price
- $0.00009
- Official
- $0.00009
- Discount
- —
| Official priceper second | Officialper second | Discount | |
|---|---|---|---|
| Realtime / Speech to text | $0.00009 | $0.00009 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Fun-ASR Realtime in Console with a prompt ready to edit or send.
Help me try fun-asr-realtime with a short audio request at /v1/realtime. Show the result and cost.
Use cases
Live captions
Show subtitles for a webinar or event as the speaker talks, updating each sentence as it settles.
Voice input
Let users dictate into a form or chat box and see their words appear while they speak.
Dialect-aware assistants
Handle spoken requests in Mandarin and regional dialects in an app that cannot wait for a full recording.
Prompt examples
Stream microphone audio for a Mandarin product launch and print captions as each sentence ends.
Transcribe a live Cantonese customer call with the hotwords 'refund' and 'order number'.
Stream a training session and keep word timestamps for later clipping.
FAQ
When should I pick Fun-ASR Realtime over Fun-ASR?
Pick Realtime when text must appear while the audio is still being spoken, such as captions or dictation. Pick the recorded-audio Fun-ASR when you have a finished recording and want diarization or archival accuracy.
How do I send audio to it?
Over a streaming WebSocket session, either through Alibaba's DashScope SDK or the raw protocol. You push audio chunks and receive recognised text events back as the session continues.
Does it support dialects?
Yes. Alibaba's real-time documentation lists Mandarin plus Cantonese, Sichuanese, and other dialects, alongside custom hotwords that help with names and product terms. Dialect speech is recognised in the same live session as standard Mandarin.
Can it handle silence between sentences?
Yes, it includes voice activity detection that filters non-speech audio and decides where sentences end. The silence thresholds are configurable in the session settings, so you can tune how quickly captions settle.
How much does Fun-ASR Realtime cost?
On TokenLab, Fun-ASR Realtime costs $0.00009 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should Fun-ASR Realtime use?
Use https://api.tokenlab.sh/v1/realtime for Fun-ASR Realtime. The request example below shows the matching code shape.
Which operations does Fun-ASR Realtime support?
Fun-ASR Realtime supports Speech to text. Select an operation above to see its endpoint and request example.
Compare Fun-ASR Realtime
Sources
- Model Studio: real-time speech recognition
- Model Studio: speech-to-text models for real-time and file transcription
Reviewed Oct 2, 2026