Alibaba: Paraformer Realtime v2
paraformer-realtime-v2- Price
- From $0.000035
- Modalities
- Audio
About Paraformer Realtime v2
Paraformer Realtime v2 is Alibaba's general-purpose streaming speech recognizer, returning text while audio is still arriving. It targets captions, spoken input, and meeting interfaces where the transcript should update as people keep talking. Unlike the 8K variant it is not tied to telephone bandwidth, and unlike Fun-ASR Realtime it belongs to the older Paraformer line.
Where it works well
- Updates the transcript continuously, which suits live captions and meeting screens.
- Works on general wideband audio such as headsets and room microphones, not only phone lines.
- Supports hotwords for names and terms that appear often in your sessions.
- Provides timestamps with the streamed results for aligning captions.
When to choose another model
- Telephone-quality 8 kHz audio is better matched by Paraformer Realtime 8K v2.
- Streaming means no second pass over the audio; use a file model when final accuracy matters most.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Speech to textWebSocketWS/v1/realtime// npm install ws import WebSocket from 'ws'; const socket = new WebSocket("wss://api.tokenlab.sh/v1/realtime?model=paraformer-realtime-v2", { headers: { Authorization: "Bearer sk-xxx" } }); socket.on('open', () => { // Send the input events supported by this model with socket.send(...). // https://tokenlab.sh/docs/en/api-reference/realtime/connect }); socket.on('message', (data) => console.log(data.toString())); socket.on('error', (error) => console.error(error)); socket.on('close', (code, reason) => console.log(code, reason.toString())); // Close the session when finished. process.on('SIGINT', () => socket.close());
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Realtime / Speech to text
per second- Official price
- $0.00003529
- Official
- $0.00003529
- Discount
- —
| Official priceper second | Officialper second | Discount | |
|---|---|---|---|
| Realtime / Speech to text | $0.00003529 | $0.00003529 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Paraformer Realtime v2 in Console with a prompt ready to edit or send.
Help me try paraformer-realtime-v2 with a short audio request at /v1/realtime. Show the result and cost.
Use cases
Meeting captions
Display a live transcript during a video meeting so participants can follow along or catch missed words.
Voice typing
Convert dictation into text in an editor or form as the user speaks.
Live event subtitles
Add subtitles to a talk or broadcast, updating each sentence as the speaker finishes.
Prompt examples
Stream this live meeting audio and show captions as people speak.
Dictate a support ticket by voice and fill the text box while I talk.
Stream a conference keynote with hotwords for the speaker and product names.
FAQ
Is Paraformer Realtime v2 for phone calls?
Not primarily. For 8 kHz telephone audio Alibaba offers Paraformer Realtime 8K v2. This model targets general live audio from microphones, headsets, and meeting software.
How does it compare with Fun-ASR Realtime?
Both stream. Fun-ASR Realtime is the newer Fun-ASR line with listed dialect support, while Paraformer Realtime v2 is the established Paraformer option. Test both on your audio and keep the one that fits.
What is the integration method?
A WebSocket session through the DashScope SDK or the raw protocol. Your client sends audio in chunks, and recognition events come back continuously while the session stays open.
Can it transcribe a finished file?
It is designed for streams. For finished recordings use Paraformer v2, the file-based counterpart, which processes the whole file and returns timestamps with every result.
How much does Paraformer Realtime v2 cost?
On TokenLab, Paraformer Realtime v2 costs $0.00003529 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should Paraformer Realtime v2 use?
Use https://api.tokenlab.sh/v1/realtime for Paraformer Realtime v2. The request example below shows the matching code shape.
Which operations does Paraformer Realtime v2 support?
Paraformer Realtime v2 supports Speech to text. Select an operation above to see its endpoint and request example.
Compare Paraformer Realtime v2
Sources
- Model Studio: real-time speech recognition
- Model Studio: speech-to-text models for real-time and file transcription
Reviewed Oct 2, 2026