Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Alibaba: Paraformer Realtime 8K v2

Paraformer Realtime 8K v2 recognizes Chinese speech from a live telephone audio stream. It is useful for call-center captions and in-call assistance that must work with narrowband telephony input.
Compare models
paraformer-realtime-8k-v2
AvailableAlibabaSpeech to TextSync
Price
From $0.000035
Modalities
Audio

About Paraformer Realtime 8K v2

Paraformer Realtime 8K v2 is Alibaba's streaming recognizer for live Chinese telephone audio at 8 kHz. It combines the real-time behaviour of Paraformer Realtime v2 with the narrowband tuning of Paraformer 8K v2. Choose it when a call is in progress and an agent screen, captions, or an assistant needs text before the call ends.

Where it works well

  • Returns text while a phone call is still going, so agents see the conversation as it happens.
  • Tuned for 8 kHz narrowband input, matching the audio telephone systems actually deliver.
  • Supports hotwords for recurring product and branch names in live calls.
  • Voice activity detection settings control how quickly sentences are finalised.

When to choose another model

  • It is aimed at narrowband phone audio; wideband microphones suit Paraformer Realtime v2 better.
  • Streaming results cannot be revised from later audio, so use Paraformer 8K v2 for the final archive transcript.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Speech to textWebSocket
    WS/v1/realtime
    // npm install ws
    import WebSocket from 'ws';
    
    const socket = new WebSocket("wss://api.tokenlab.sh/v1/realtime?model=paraformer-realtime-8k-v2", {
      headers: { Authorization: "Bearer sk-xxx" }
    });
    
    socket.on('open', () => {
      // Send the input events supported by this model with socket.send(...).
      // https://tokenlab.sh/docs/en/api-reference/realtime/connect
    });
    socket.on('message', (data) => console.log(data.toString()));
    socket.on('error', (error) => console.error(error));
    socket.on('close', (code, reason) => console.log(code, reason.toString()));
    
    // Close the session when finished.
    process.on('SIGINT', () => socket.close());

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Realtime / Speech to text

per second
Official price
$0.00003529
Official
$0.00003529
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Paraformer Realtime 8K v2 in Console with a prompt ready to edit or send.

Help me try paraformer-realtime-8k-v2 with a short audio request at /v1/realtime. Show the result and cost.

Use cases

  • Agent assist

    Feed live call text to a language model that suggests answers or surfaces policy snippets during the conversation.

  • Live call monitoring

    Flag compliance phrases or escalation keywords while the call is still active.

  • In-call notes

    Build a running note for the agent that is complete by the time the call ends.

Prompt examples

Stream this live 8 kHz Mandarin support call and show text as each sentence finishes.

Transcribe an ongoing sales call live, with hotwords for our three plan names.

Stream the call audio and mark the moment the customer says 'cancel'.

FAQ

How does it differ from Paraformer Realtime v2?

The 8K version is tuned for narrowband telephone audio sampled at 8 kHz, while Paraformer Realtime v2 handles general wideband live audio such as microphones. Match the model to where your audio comes from.

How does it differ from Paraformer 8K v2?

Both target telephone audio, but this one streams and returns text during the call, whereas Paraformer 8K v2 transcribes a finished recording from a file.

How is audio sent?

Through a streaming WebSocket session, using Alibaba's DashScope SDK or the raw protocol. You send audio chunks and receive recognition events as the call continues.

Does it only work for Chinese?

Alibaba's description centres on Chinese telephone speech. Verify any other language on your own call samples before depending on it, since narrowband audio makes recognition harder for every language.

How much does Paraformer Realtime 8K v2 cost?

On TokenLab, Paraformer Realtime 8K v2 costs $0.00003529 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Paraformer Realtime 8K v2 use?

Use https://api.tokenlab.sh/v1/realtime for Paraformer Realtime 8K v2. The request example below shows the matching code shape.

Which operations does Paraformer Realtime 8K v2 support?

Paraformer Realtime 8K v2 supports Speech to text. Select an operation above to see its endpoint and request example.

Compare Paraformer Realtime 8K v2

Sources

Reviewed Oct 2, 2026

More from Paraformer

Related models