Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

xAI: Grok Voice Think Fast 2.0

Grok Voice Think Fast 2.0 is a voice model for live conversation. It is a candidate for spoken interfaces in which the assistant must keep up with an ongoing exchange rather than simply read a prepared script.
Compare models
grok-voice-think-fast-2.0
AvailablexAIAudioSync
Price
From $0.001333
Modalities
Audio

About Grok Voice Think Fast 2.0

Grok Voice Think Fast 2.0 is xAI's real-time voice model for live, two-way conversation, built for assistants that must keep pace with a speaker rather than read a prepared script. It is the versioned model behind the grok-voice-latest alias. Use this exact ID when a product needs a fixed model version for its spoken replies.

Where it works well

  • It is the fixed, versioned release behind xAI's voice agent alias, so behavior stays tied to one model until you move off it.
  • Server-side voice activity detection manages turn-taking, so the client streams microphone audio without its own pause logic.
  • Per-response instructions let a session change the system prompt for a single reply, for example to switch tone for a refund explanation.
  • Pronunciation replacements correct how names and product terms are spoken without altering the transcript.

When to choose another model

  • It is a conversational speech-to-speech model, so batch transcription of files and narration of long texts are better served by the dedicated speech models.
  • Keeping a connection open for live audio takes a WebSocket client; a plain request and response call will not drive it.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    AudioChat Completions API
    POST/v1/chat/completions
    curl https://api.tokenlab.sh/v1/chat/completions \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer sk-xxx" \
      -d '{
        "model": "grok-voice-think-fast-2.0",
        "messages": [
          {"role": "user", "content": "Hello!"}
        ]
      }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Audio Duration

per second
Official price
$0.001333
Official
$0.001333
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Use cases

  • Interactive voice response replacement

    Replace menu-driven phone trees with a caller who simply states the problem, while the agent transfers the call or resolves it.

  • Language practice partner

    Hold a spoken dialogue in a target language, correcting mistakes aloud and adjusting speed when the learner asks.

  • Hands-free field assistant

    Let a technician ask questions about a manual by voice and hear answers while both hands stay on the equipment.

Prompt examples

You are a calm support agent for an internet provider. Ask for the account postcode, then walk the caller through restarting the router.

Pronounce our brand 'Zyntra' as ZIN-tra, keep a friendly tone, and never interrupt the caller mid-sentence.

Act as a Portuguese tutor. Speak slowly, repeat my sentence corrected, and ask one follow-up question each turn.

FAQ

What is the difference between grok-voice-think-fast-2.0 and grok-voice-latest?

grok-voice-think-fast-2.0 is a specific model version, while grok-voice-latest is an alias that currently points to it. The alias will move to newer releases; the dated name stays on the version you tested.

Is Grok Voice Think Fast 2.0 a text-to-speech model?

No. It listens and answers in a live session, taking audio in and producing spoken replies. For turning prepared text into audio, use Grok Voice TTS; for transcribing recordings, use Grok Voice STT.

How does it decide when I have finished speaking?

Sessions can enable server-side voice activity detection, which detects the end of a user turn and starts the reply. You can also manage turns yourself if your client already has push-to-talk.

Which voices can it use?

A session can choose among xAI's built-in voices, with eve as the default, or a custom voice cloned from a short reference recording. The built-in voices speak all supported languages.

How much does Grok Voice Think Fast 2.0 cost?

On TokenLab, Grok Voice Think Fast 2.0 costs $0.001333 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Grok Voice Think Fast 2.0 use?

Use https://api.tokenlab.sh/v1/chat/completions for Grok Voice Think Fast 2.0. The request example below shows the matching code shape.

Compare Grok Voice Think Fast 2.0

Sources

Reviewed Oct 2, 2026

More from Grok Voice

Related models