Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Alibaba: Fun-ASR Realtime

Fun-ASR Realtime consumes incoming audio as a session progresses. It is designed for live transcription experiences such as captions and voice input, where waiting for a complete recording would interrupt the interaction.
Compare models
fun-asr-realtime
AvailableAlibabaSpeech to TextSync
Price
From $0.00009
Modalities
Audio

About Fun-ASR Realtime

Fun-ASR Realtime is Alibaba's streaming speech recognition model: it receives audio over a live session and returns text as people speak. It is the Fun-ASR variant for captions, voice input, and in-call assistance, where the recorded-audio Fun-ASR would force users to wait until the recording ends. Alibaba lists hotwords, timestamps, and Mandarin with dialects such as Cantonese and Sichuanese.

Where it works well

  • Streams partial text while the speaker is still talking, so captions appear without waiting for a finished recording.
  • Recognises Mandarin with high accuracy and several dialects, including Cantonese and Sichuanese, per Alibaba's documentation.
  • Custom hotwords steer live recognition toward names and terms specific to your product.
  • Word- and sentence-level timestamps arrive with the stream, which helps align captions.

When to choose another model

  • It needs a WebSocket session, which is more integration work than uploading a recording to Fun-ASR.
  • Live decoding cannot revisit earlier audio, so for archival accuracy the recorded-audio Fun-ASR is the safer pick.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Speech to textWebSocket
    WS/v1/realtime
    // npm install ws
    import WebSocket from 'ws';
    
    const socket = new WebSocket("wss://api.tokenlab.sh/v1/realtime?model=fun-asr-realtime", {
      headers: { Authorization: "Bearer sk-xxx" }
    });
    
    socket.on('open', () => {
      // Send the input events supported by this model with socket.send(...).
      // https://tokenlab.sh/docs/en/api-reference/realtime/connect
    });
    socket.on('message', (data) => console.log(data.toString()));
    socket.on('error', (error) => console.error(error));
    socket.on('close', (code, reason) => console.log(code, reason.toString()));
    
    // Close the session when finished.
    process.on('SIGINT', () => socket.close());

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Realtime / Speech to text

per second
Official price
$0.00009
Official
$0.00009
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Fun-ASR Realtime in Console with a prompt ready to edit or send.

Help me try fun-asr-realtime with a short audio request at /v1/realtime. Show the result and cost.

Use cases

  • Live captions

    Show subtitles for a webinar or event as the speaker talks, updating each sentence as it settles.

  • Voice input

    Let users dictate into a form or chat box and see their words appear while they speak.

  • Dialect-aware assistants

    Handle spoken requests in Mandarin and regional dialects in an app that cannot wait for a full recording.

Prompt examples

Stream microphone audio for a Mandarin product launch and print captions as each sentence ends.

Transcribe a live Cantonese customer call with the hotwords 'refund' and 'order number'.

Stream a training session and keep word timestamps for later clipping.

FAQ

When should I pick Fun-ASR Realtime over Fun-ASR?

Pick Realtime when text must appear while the audio is still being spoken, such as captions or dictation. Pick the recorded-audio Fun-ASR when you have a finished recording and want diarization or archival accuracy.

How do I send audio to it?

Over a streaming WebSocket session, either through Alibaba's DashScope SDK or the raw protocol. You push audio chunks and receive recognised text events back as the session continues.

Does it support dialects?

Yes. Alibaba's real-time documentation lists Mandarin plus Cantonese, Sichuanese, and other dialects, alongside custom hotwords that help with names and product terms. Dialect speech is recognised in the same live session as standard Mandarin.

Can it handle silence between sentences?

Yes, it includes voice activity detection that filters non-speech audio and decides where sentences end. The silence thresholds are configurable in the session settings, so you can tune how quickly captions settle.

How much does Fun-ASR Realtime cost?

On TokenLab, Fun-ASR Realtime costs $0.00009 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Fun-ASR Realtime use?

Use https://api.tokenlab.sh/v1/realtime for Fun-ASR Realtime. The request example below shows the matching code shape.

Which operations does Fun-ASR Realtime support?

Fun-ASR Realtime supports Speech to text. Select an operation above to see its endpoint and request example.

Compare Fun-ASR Realtime

Sources

Reviewed Oct 2, 2026

More from Fun-ASR

Related models