Audio Models
AI models for speech, music, and audio processing
Audio
Text to speech
- Verified price— Discount
- —
- Official price
- Input$100.00/1M charactersOutput$0.00/1M characters
Text to speech
- Verified price— Discount
- —
- Official price
- Input$22.06/1M charactersOutput$0.00/1M characters
Text to speech
- Verified price— Discount
- —
- Official price
- Input$40.00/1M charactersOutput$0.00/1M characters
Text to speech
- Verified price— Discount
- —
- Official price
- Input$15.00/1M charactersOutput$0.00/1M characters
Text to speech
- Verified price— Discount
- —
- Official price
- Input$15.00/1M charactersOutput$0.00/1M characters
Text to speech
- Verified price— Discount
- —
- Official price
- Input$100.00/1M charactersOutput$0.00/1M characters
Speech to Text
- Verified price— Discount
- —
- Official price
- $0.000035/sec
Text to speech
- Verified price— Discount
- —
- Official price
- Input$15.00/1M charactersOutput$0.00/1M characters
Text to speech
- Verified price— Discount
- —
- Official price
- Input$15.00/1M charactersOutput$0.00/1M characters
Text to speech
- Verified price— Discount
- —
- Official price
- Input$10.00/1M charactersOutput$0.00/1M characters
FastText to speech
- Verified price— Discount
- —
- Official price
- Input$60.00/1M charactersOutput$0.00/1M characters
All model series
6 seriesKling
ImageVideo+2
Auto pricing units vary by model
15 modelsOpen series
MiniMax
ChatImage+4
Auto pricing units vary by model
11 modelsOpen series
Grok Voice
AudioText to speech+1
Auto pricing units vary by model
6 modelsOpen series
Qwen Omni
ChatAudio+3
Auto from $0.15 · per 1M tokens
4 modelsOpen series
Stable Audio
AudioMusic
Auto from $0.0206 · per request
4 modelsOpen series
Qwen TTS
AudioText to speech+1
Auto from $10.00 · per 1M characters
3 modelsOpen series
Match the audio API to the task
Speech generation, transcription, and real-time conversation have different inputs, outputs, and latency requirements. An audio category label does not mean every model supports all three.
Selection signals
- Choose the required operation first, then verify accepted audio formats, languages, and length limits.
- Compare the relevant units, which may include audio duration, text characters, or tokens.
- Evaluate accents, background noise, pronunciation, and turn-taking on representative recordings.