Alibaba: Paraformer v2
paraformer-v2- Price
- From $0.000012
- Modalities
- Audio
About Paraformer v2
Paraformer v2 is Alibaba's file-based speech recognition model for offline transcripts of meetings and long conversations. It is called asynchronously on recordings up to 2 GB and 12 hours. Compared with Paraformer 8K v2 it is meant for general wideband audio, and compared with the real-time models it trades immediacy for a complete transcript with diarization and hotwords.
Where it works well
- Handles very long recordings, up to 12 hours or 2 GB per file per Alibaba's documentation.
- Timestamps are always included, so no extra option is needed to get them.
- Speaker diarization separates participants in meetings and group discussions.
- Sensitive-word filtering can mask listed terms in the transcript text.
When to choose another model
- Results arrive after the job completes; live captions need Paraformer Realtime v2.
- For 8 kHz phone recordings the 8K variant matches the audio better.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Speech to textTokenLab endpointPOST/v1/audio/transcriptionscurl -X POST "https://api.tokenlab.sh/v1/audio/transcriptions" \ -H "Authorization: Bearer sk-xxx" \ -F file="@audio.mp3" \ -F operation="stt" \ -F response_format="json" \ -F model="paraformer-v2" \ -F language="en"
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Speech to text
per second- Official price
- $0.00001176
- Official
- $0.00001176
- Discount
- —
| Official priceper second | Officialper second | Discount | |
|---|---|---|---|
| Speech to text | $0.00001176 | $0.00001176 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Paraformer v2 in Console with a prompt ready to edit or send.
Help me try paraformer-v2 with a short audio request at /v1/audio/transcriptions. Show the result and cost.
Use cases
Meeting archives
Process saved meeting recordings into searchable text with speaker labels and timestamps.
Lecture transcripts
Transcribe multi-hour lectures once, then index or summarise the text for students.
Content moderation prep
Mask configured sensitive words in transcripts before they are shared with a wider team.
Prompt examples
Transcribe this 3-hour all-hands recording with speaker labels and timestamps.
Transcribe this lecture video's audio and mask the listed sensitive words.
Transcribe these weekly stand-up recordings using hotwords for our project names.
FAQ
What is Paraformer v2 best for?
Offline transcription of recorded meetings, interviews, and conversations. It accepts mainstream audio and video formats at any sample rate and returns the text together with timestamps.
How big can the input be?
Alibaba documents files up to 2 GB and 12 hours for its Paraformer file models. Split anything longer into parts before you submit it for transcription.
Can I turn timestamps off?
No. Alibaba states that timestamps are permanently enabled for Paraformer, so every transcript carries them and no request option is needed to get time positions.
Paraformer v2 or Fun-ASR?
Both transcribe files with diarization and hotwords. Fun-ASR is the newer line; Paraformer v2 is the established one. Run a sample of your audio through each and compare the errors that matter to you.
How much does Paraformer v2 cost?
On TokenLab, Paraformer v2 costs $0.00001176 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should Paraformer v2 use?
Use https://api.tokenlab.sh/v1/audio/transcriptions for Paraformer v2. The request example below shows the matching code shape.
Which operations does Paraformer v2 support?
Paraformer v2 supports Speech to text. Select an operation above to see its endpoint and request example.
Compare Paraformer v2
Sources
- Model Studio: recording file recognition models
- Model Studio: speech-to-text models for real-time and file transcription
Reviewed Oct 2, 2026