xAI: Grok Voice TTS
grok-voice-tts- Input / Output
- $15.00 / $0.00
- Modalities
- Audio
About Grok Voice TTS
Grok Voice TTS is xAI's text-to-speech model, rendering written text as spoken audio in a chosen voice and language. Inline tags such as pauses, laughter, and whispers shape delivery. Use it for narration, spoken notifications, and localized audio answers where the text is already written and no live conversation is needed.
Where it works well
- Inline speech tags such as [pause] and [laugh], plus wrapping tags like whisper, give delivery control inside the text itself.
- Built-in voices all speak the supported languages, so one voice identity can carry across localized versions of the same content.
- Custom voices can be cloned from a short reference clip for brand or character voices.
- Character-level timestamps can be returned for syncing captions or lip movement to the audio.
When to choose another model
- It speaks fixed text; a caller who interrupts or replies needs a conversational voice model instead.
- Output quality for long scripts depends on punctuation and tags in the source text, so unedited machine text can sound flat.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to speechTokenLab endpointPOST/v1/audio/speechcurl -X POST "https://api.tokenlab.sh/v1/audio/speech" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "language_code": "en", "operation": "tts", "response_format": "mp3", "voice": "eve", "model": "grok-voice-tts", "input": "Hello from TokenLab." }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Pricing
per 1M characters- Official
- Input$15.00/1M charactersOutput$0.00/1M characters
- Official price
- Input$15.00/1M charactersOutput$0.00/1M characters
| Official priceper 1M characters | Officialper 1M characters | Discount | |
|---|---|---|---|
| Input | $15.00 | $15.00 | — |
| Output | $0.00 | $0.00 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Grok Voice TTS in Console with a prompt ready to edit or send.
Help me try grok-voice-tts with a short audio request at /v1/audio/speech. Show the result and cost.
Use cases
Narrated articles and lessons
Turn a finished article or course script into an audio track, using pauses to separate sections.
Spoken alerts
Read out order updates or safety notices in the user's language through a telephony codec.
Localized product videos
Voice the same script in several languages with one chosen voice, then use timestamps to line captions up.
Prompt examples
Read this paragraph in a warm, steady voice and add a short [pause] after each sentence.
Speak the following notice in French, then repeat it in German, using the same voice.
<whisper>Your package has shipped.</whisper> [pause] It should arrive on Thursday.
FAQ
What does Grok Voice TTS do?
It converts text into speech audio. You choose a voice, optionally a language code, and an output format, and it returns audio for the text, including any expressive tags you wrote inline.
Which audio formats can it produce?
MP3, WAV, and raw PCM, plus telephony codecs mu-law and A-law, at sample rates from 8 kHz to 48 kHz. Telephony codecs fit phone systems; MP3 suits web and mobile playback.
How many languages can it speak?
xAI documents twenty languages selected by BCP-47 code, with automatic language detection as an alternative. Each built-in voice can speak all of the supported languages, so switching language does not require a new voice.
Can I stream audio while text is still arriving?
Yes. A WebSocket endpoint generates audio in real time as text is sent, which helps when another model is still writing the reply that will be spoken.
Can I fix mispronounced words?
Yes. Pronunciation replacements let you map a written term to how it should sound, so brand names and acronyms come out right without editing the displayed text.
How much does Grok Voice TTS cost?
On TokenLab, Grok Voice TTS costs Input $15.00 / Output $0.00 per 1M characters. The pricing table above shows the full breakdown.
Which endpoint should Grok Voice TTS use?
Use https://api.tokenlab.sh/v1/audio/speech for Grok Voice TTS. The request example below shows the matching code shape.
Which operations does Grok Voice TTS support?
Grok Voice TTS supports Text to speech. Select an operation above to see its endpoint and request example.
Compare Grok Voice TTS
Sources
Reviewed Oct 2, 2026