Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

xAI: Grok TTS

Grok TTS generates speech from text using xAI's expressive voices. The published interface returns MP3 audio and supports choosing a voice, making it useful for spoken responses, narration, and multilingual voice output.
Compare models
grok-tts
AvailablexAIText to speechSync
Input / Output
$15.00 / $0.00
Modalities
Audio

About Grok TTS

Grok TTS is xAI's text-to-speech model, turning written text into spoken audio in expressive voices. The interface returns MP3 audio and lets you choose a voice. xAI documents inline speech tags for laughter, whispers and pauses, so scripts can carry delivery cues. It fits spoken replies, narration and multilingual voice output.

Where it works well

  • Offers a large roster of expressive voices, so you can pick one that matches the product or character.
  • Returns MP3 audio, which plays in browsers and apps without conversion.
  • xAI documents inline tags such as laughter, whispers and pauses for shaping delivery inside the text.
  • All built-in voices work across the supported languages, so one voice can speak several languages.
  • Plain text goes in and a finished audio file comes out, which keeps long scripts simple to automate.

When to choose another model

  • Output is MP3 only through this interface, so telephony formats need conversion.
  • Voice cloning and live streaming delivery are separate features and are not part of this request.
  • Long scripts need to be split into chunks, and you should listen to pronunciation of names and acronyms.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to speechTokenLab endpoint
    POST/v1/audio/speech
    curl -X POST "https://api.tokenlab.sh/v1/audio/speech" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "operation": "tts",
      "response_format": "mp3",
      "model": "grok-tts",
      "input": "Hello from TokenLab.",
      "voice": "alloy"
    }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Text to speech

per 1M characters
Official price
Input $15.00 / Output $0.00
Official
Input $15.00 / Output $0.00
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Grok TTS in Console with a prompt ready to edit or send.

Help me try grok-tts with a short audio request at /v1/audio/speech. Show the result and cost.

Use cases

  • Voice assistants

    Speak replies from a chat model aloud, choosing a voice that fits the product.

  • Narration and audiobooks

    Convert articles, lessons or chapters into audio files, adding pause tags where pacing needs a break.

  • Multilingual announcements

    Produce the same message in several languages using one voice for a consistent brand sound.

Prompt examples

Welcome back! [pause] Your order has shipped and should arrive on Thursday.

Read this product description in a calm, friendly tone, whispering the final line.

Narrate the following paragraph in French at an even pace, with a short pause after each sentence.

FAQ

What is Grok TTS?

It is xAI's text-to-speech model. You send text, choose a voice, and receive spoken audio as an MP3 file. It is meant for spoken replies, narration and multilingual voice output from written text.

Can I choose a voice?

Yes. xAI provides a roster of expressive voices, with eve as the default. Voice is a request option, so you can change it per call to suit the script.

Does it support expressive speech?

xAI documents inline speech tags for effects such as laughter, whispers and pauses. Place them in the text where you want the effect and listen to the result.

Which languages can it speak?

xAI states that its built-in voices work across all available languages. Test your language with a short passage first, paying attention to names and numbers.

What audio format does it return?

MP3, which plays directly in browsers and mobile apps. If your platform needs another format, such as telephony audio, convert the file after generation. Any standard audio tool can do the conversion.

How much does Grok TTS cost?

On TokenLab, Grok TTS costs Input $15.00 / Output $0.00 per 1M characters. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Grok TTS use?

Use https://api.tokenlab.sh/v1/audio/speech for Grok TTS. The request example below shows the matching code shape.

Which operations does Grok TTS support?

Grok TTS supports Text to speech. Select an operation above to see its endpoint and request example.

Compare Grok TTS

Sources

Reviewed Oct 2, 2026

More from Grok Voice

Related models