Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Alibaba: Paraformer 8K v2

Paraformer 8K v2 is built for Chinese telephone recordings sampled at 8 kHz. It is a focused choice for transcribing call archives where the audio comes from telephony rather than a studio-quality microphone.
Compare models
paraformer-8k-v2
AvailableAlibabaSpeech to TextSync / Async
Price
From $0.000012
Modalities
Audio

About Paraformer 8K v2

Paraformer 8K v2 is Alibaba's file-transcription model tuned for 8 kHz telephone recordings, mainly Chinese call audio. It is the narrowband sibling of Paraformer v2: use it when the source is a phone line or call-centre archive, and the general version for wider-band meeting audio. Timestamps are always on, and diarization and hotwords are available.

Where it works well

  • Built for 8 kHz telephony, so call-centre recordings need no upsampling before transcription.
  • Timestamps are always returned, which suits quality review that jumps to a position in a call.
  • Speaker diarization can separate agent and customer in a two-party call.
  • Custom hotwords help with product names, plan names, and branch terms that recur in calls.

When to choose another model

  • It targets narrowband phone audio; wideband meeting or studio recordings are better served by Paraformer v2.
  • It is an offline file model, so live in-call captions need Paraformer Realtime 8K v2.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Speech to textTokenLab endpoint
    POST/v1/audio/transcriptions
    curl -X POST "https://api.tokenlab.sh/v1/audio/transcriptions" \
      -H "Authorization: Bearer sk-xxx" \
      -F file="@audio.mp3" \
      -F operation="stt" \
      -F response_format="json" \
      -F model="paraformer-8k-v2" \
      -F language="en"

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Speech to text

per second
Official price
$0.00001176
Official
$0.00001176
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Paraformer 8K v2 in Console with a prompt ready to edit or send.

Help me try paraformer-8k-v2 with a short audio request at /v1/audio/transcriptions. Show the result and cost.

Use cases

  • Call-centre archives

    Transcribe stored customer calls to feed quality scoring, compliance review, and keyword search.

  • Agent and customer split

    Use diarization to produce a two-sided transcript, then summarise each side's requests separately.

  • Dispute evidence

    Keep timestamps beside the text so a reviewer can find and replay the exact moment of a disputed statement.

Prompt examples

Transcribe this 8 kHz Mandarin call recording and separate the agent from the customer.

Transcribe these stored loan-collection calls using hotwords for our product names.

Transcribe this phone interview and return timestamps for every sentence.

FAQ

What is the 8k in paraformer-8k-v2?

It refers to 8 kHz sampling, the rate of ordinary telephone audio. The model is tuned for that narrowband material, mainly Chinese call recordings, rather than for wideband sources.

Should I use it for meeting recordings?

Usually not. Meeting audio recorded on laptops or room microphones is wider band, and Paraformer v2 is the matching model. Use the 8K version when the source really is a phone call.

Are timestamps optional on this model?

No. Alibaba states that timestamps are permanently enabled for the Paraformer file models and cannot be turned off, so every transcript this model returns carries them with the text.

Can it transcribe a live call?

Not as a stream. This is the offline file model. For live phone audio use Paraformer Realtime 8K v2, which accepts a stream of narrowband audio.

How much does Paraformer 8K v2 cost?

On TokenLab, Paraformer 8K v2 costs $0.00001176 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Paraformer 8K v2 use?

Use https://api.tokenlab.sh/v1/audio/transcriptions for Paraformer 8K v2. The request example below shows the matching code shape.

Which operations does Paraformer 8K v2 support?

Paraformer 8K v2 supports Speech to text. Select an operation above to see its endpoint and request example.

Compare Paraformer 8K v2

Sources

Reviewed Oct 2, 2026

More from Paraformer

Related models