Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Alibaba: Paraformer v2

Paraformer v2 transcribes recorded audio for offline processing. It is useful for meeting archives and longer conversation recordings that need a reusable written transcript after the session has ended.
Compare models
paraformer-v2
AvailableAlibabaSpeech to TextSync / Async
Price
From $0.000012
Modalities
Audio

About Paraformer v2

Paraformer v2 is Alibaba's file-based speech recognition model for offline transcripts of meetings and long conversations. It is called asynchronously on recordings up to 2 GB and 12 hours. Compared with Paraformer 8K v2 it is meant for general wideband audio, and compared with the real-time models it trades immediacy for a complete transcript with diarization and hotwords.

Where it works well

  • Handles very long recordings, up to 12 hours or 2 GB per file per Alibaba's documentation.
  • Timestamps are always included, so no extra option is needed to get them.
  • Speaker diarization separates participants in meetings and group discussions.
  • Sensitive-word filtering can mask listed terms in the transcript text.

When to choose another model

  • Results arrive after the job completes; live captions need Paraformer Realtime v2.
  • For 8 kHz phone recordings the 8K variant matches the audio better.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Speech to textTokenLab endpoint
    POST/v1/audio/transcriptions
    curl -X POST "https://api.tokenlab.sh/v1/audio/transcriptions" \
      -H "Authorization: Bearer sk-xxx" \
      -F file="@audio.mp3" \
      -F operation="stt" \
      -F response_format="json" \
      -F model="paraformer-v2" \
      -F language="en"

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Speech to text

per second
Official price
$0.00001176
Official
$0.00001176
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Paraformer v2 in Console with a prompt ready to edit or send.

Help me try paraformer-v2 with a short audio request at /v1/audio/transcriptions. Show the result and cost.

Use cases

  • Meeting archives

    Process saved meeting recordings into searchable text with speaker labels and timestamps.

  • Lecture transcripts

    Transcribe multi-hour lectures once, then index or summarise the text for students.

  • Content moderation prep

    Mask configured sensitive words in transcripts before they are shared with a wider team.

Prompt examples

Transcribe this 3-hour all-hands recording with speaker labels and timestamps.

Transcribe this lecture video's audio and mask the listed sensitive words.

Transcribe these weekly stand-up recordings using hotwords for our project names.

FAQ

What is Paraformer v2 best for?

Offline transcription of recorded meetings, interviews, and conversations. It accepts mainstream audio and video formats at any sample rate and returns the text together with timestamps.

How big can the input be?

Alibaba documents files up to 2 GB and 12 hours for its Paraformer file models. Split anything longer into parts before you submit it for transcription.

Can I turn timestamps off?

No. Alibaba states that timestamps are permanently enabled for Paraformer, so every transcript carries them and no request option is needed to get time positions.

Paraformer v2 or Fun-ASR?

Both transcribe files with diarization and hotwords. Fun-ASR is the newer line; Paraformer v2 is the established one. Run a sample of your audio through each and compare the errors that matter to you.

How much does Paraformer v2 cost?

On TokenLab, Paraformer v2 costs $0.00001176 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Paraformer v2 use?

Use https://api.tokenlab.sh/v1/audio/transcriptions for Paraformer v2. The request example below shows the matching code shape.

Which operations does Paraformer v2 support?

Paraformer v2 supports Speech to text. Select an operation above to see its endpoint and request example.

Compare Paraformer v2

Sources

Reviewed Oct 2, 2026

More from Paraformer

Related models