Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Alibaba: Paraformer Realtime v2

Paraformer Realtime v2 provides streaming speech recognition for live audio. It suits captions, spoken input, and meeting interfaces where transcription needs to update as people continue speaking.
Compare models
paraformer-realtime-v2
AvailableAlibabaSpeech to TextSync
Price
From $0.000035
Modalities
Audio

About Paraformer Realtime v2

Paraformer Realtime v2 is Alibaba's general-purpose streaming speech recognizer, returning text while audio is still arriving. It targets captions, spoken input, and meeting interfaces where the transcript should update as people keep talking. Unlike the 8K variant it is not tied to telephone bandwidth, and unlike Fun-ASR Realtime it belongs to the older Paraformer line.

Where it works well

  • Updates the transcript continuously, which suits live captions and meeting screens.
  • Works on general wideband audio such as headsets and room microphones, not only phone lines.
  • Supports hotwords for names and terms that appear often in your sessions.
  • Provides timestamps with the streamed results for aligning captions.

When to choose another model

  • Telephone-quality 8 kHz audio is better matched by Paraformer Realtime 8K v2.
  • Streaming means no second pass over the audio; use a file model when final accuracy matters most.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Speech to textWebSocket
    WS/v1/realtime
    // npm install ws
    import WebSocket from 'ws';
    
    const socket = new WebSocket("wss://api.tokenlab.sh/v1/realtime?model=paraformer-realtime-v2", {
      headers: { Authorization: "Bearer sk-xxx" }
    });
    
    socket.on('open', () => {
      // Send the input events supported by this model with socket.send(...).
      // https://tokenlab.sh/docs/en/api-reference/realtime/connect
    });
    socket.on('message', (data) => console.log(data.toString()));
    socket.on('error', (error) => console.error(error));
    socket.on('close', (code, reason) => console.log(code, reason.toString()));
    
    // Close the session when finished.
    process.on('SIGINT', () => socket.close());

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Realtime / Speech to text

per second
Official price
$0.00003529
Official
$0.00003529
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Paraformer Realtime v2 in Console with a prompt ready to edit or send.

Help me try paraformer-realtime-v2 with a short audio request at /v1/realtime. Show the result and cost.

Use cases

  • Meeting captions

    Display a live transcript during a video meeting so participants can follow along or catch missed words.

  • Voice typing

    Convert dictation into text in an editor or form as the user speaks.

  • Live event subtitles

    Add subtitles to a talk or broadcast, updating each sentence as the speaker finishes.

Prompt examples

Stream this live meeting audio and show captions as people speak.

Dictate a support ticket by voice and fill the text box while I talk.

Stream a conference keynote with hotwords for the speaker and product names.

FAQ

Is Paraformer Realtime v2 for phone calls?

Not primarily. For 8 kHz telephone audio Alibaba offers Paraformer Realtime 8K v2. This model targets general live audio from microphones, headsets, and meeting software.

How does it compare with Fun-ASR Realtime?

Both stream. Fun-ASR Realtime is the newer Fun-ASR line with listed dialect support, while Paraformer Realtime v2 is the established Paraformer option. Test both on your audio and keep the one that fits.

What is the integration method?

A WebSocket session through the DashScope SDK or the raw protocol. Your client sends audio in chunks, and recognition events come back continuously while the session stays open.

Can it transcribe a finished file?

It is designed for streams. For finished recordings use Paraformer v2, the file-based counterpart, which processes the whole file and returns timestamps with every result.

How much does Paraformer Realtime v2 cost?

On TokenLab, Paraformer Realtime v2 costs $0.00003529 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Paraformer Realtime v2 use?

Use https://api.tokenlab.sh/v1/realtime for Paraformer Realtime v2. The request example below shows the matching code shape.

Which operations does Paraformer Realtime v2 support?

Paraformer Realtime v2 supports Speech to text. Select an operation above to see its endpoint and request example.

Compare Paraformer Realtime v2

Sources

Reviewed Oct 2, 2026

More from Paraformer

Related models