Audio Models
AI models for speech, music, and audio processing
Audio · Alibaba Cloud
Text to speech
- TokenLab price— Discount
- —
- Official price
- Input$22.06Output$0.00/1M characters
Tool useVisionWeb search
- TokenLab price— Discount
- —
- Official price
- Input$3.97Output$15.74/1M
Tool useVisionWeb search
- TokenLab price— Discount
- —
- Official price
- Input$11.76Output$44.12/1M
Vision
- TokenLab price— Discount
- —
- Official price
- Input$1.47Output$1.47/1M
Vision
- TokenLab price— Discount
- —
- Official price
- Input$5.88Output$14.71/1M
Speech to Text
- TokenLab price— Discount
- —
- Official price
- $0.00003235/sec
Speech to Text
- TokenLab price— Discount
- —
- Official price
- $0.0000353/sec
Speech to Text
- TokenLab price— Discount
- —
- Official price
- $0.00003235/sec
Speech to Text
- TokenLab price— Discount
- —
- Official price
- $0.00004853/sec
Text to speech
- TokenLab price— Discount
- —
- Official price
- Input$11.76Output$0.00/1M characters
Text to speech
- TokenLab price— Discount
- —
- Official price
- Input$14.71Output$0.00/1M characters
All model series
3 seriesChoosing the right audio model
The right model balances task fit, output quality, latency, and price.
Selection signals
- Match the model to the input and output you actually need.
- Compare models that use the same pricing unit.
- A small test reveals quality and latency differences that a catalog cannot.
FAQ
What makes a good audio model?
A strong fit matches the task, quality bar, latency target, and budget. A small evaluation set makes quality and reliability differences easier to see.
Can I switch models later?
Yes. Keep the public model ID and request format explicit, and compare alternatives on the same tasks.