Audio Models

AI models for speech, music, and audio processing

Audio · xAI

grok-voice-latest
Available
TokenLab price— Discount
—
Official price
$0.001333/sec
grok-voice-stt
Available
Speech to Text
TokenLab price/ Official price— Discount
—
grok-voice-think-fast-2.0
Available
TokenLab price— Discount
—
Official price
$0.001333/sec
grok-voice-tts
Available
Text to speech
TokenLab price/ Official price— Discount
—

All model series

1 series

Choosing the right audio model

The right model balances task fit, output quality, latency, and price.

Selection signals

  • Match the model to the input and output you actually need.
  • Compare models that use the same pricing unit.
  • A small test reveals quality and latency differences that a catalog cannot.

FAQ

What makes a good audio model?

A strong fit matches the task, quality bar, latency target, and budget. A small evaluation set makes quality and reliability differences easier to see.

Can I switch models later?

Yes. Keep the public model ID and request format explicit, and compare alternatives on the same tasks.