Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

xAI: Grok Voice Latest

Grok Voice Latest provides a conversational voice session with selectable voices and tool interaction. It is useful for spoken assistants whose responses need to remain connected to actions and information from the surrounding application.
Compare models
grok-voice-latest
AvailablexAIAudioSyncNew
Price
From $0.001333
Released
Apr 23, 2026
Modalities
Audio

About Grok Voice Latest

Grok Voice Latest is xAI's floating alias for its current real-time speech-to-speech model, used to hold a spoken conversation over a WebSocket session. At the time of review it points to grok-voice-think-fast-2.0. Choose it when a voice assistant should pick up xAI's improvements without a code change; pin the versioned model when behavior must stay fixed.

Where it works well

  • It is the alias xAI recommends for voice agents that should receive model upgrades automatically, with no change to the session configuration.
  • A session accepts built-in or custom voices, server-side voice activity detection for turn-taking, and language hints for transcription.
  • Agents can call web and X search, file search over document collections, remote MCP servers, and custom function tools.
  • Sessions can resume after a dropped connection by reusing a conversation ID, which suits phone and mobile clients.

When to choose another model

  • Because the alias moves, voice, pacing, or wording can shift after an upgrade; pin the versioned model for regression-tested products.
  • It handles live audio sessions only; transcribing stored recordings or rendering prepared text needs the speech-to-text or text-to-speech models.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    RealtimeWebSocket
    WS/v1/realtime
    // npm install ws
    import WebSocket from 'ws';
    
    const socket = new WebSocket("wss://api.tokenlab.sh/v1/realtime?model=grok-voice-latest", {
      headers: { Authorization: "Bearer sk-xxx" }
    });
    
    socket.on('open', () => {
      // Send the input events supported by this model with socket.send(...).
      // https://tokenlab.sh/docs/en/api-reference/realtime/connect
    });
    socket.on('message', (data) => console.log(data.toString()));
    socket.on('error', (error) => console.error(error));
    socket.on('close', (code, reason) => console.log(code, reason.toString()));
    
    // Close the session when finished.
    process.on('SIGINT', () => socket.close());

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Audio Duration

per second
Official price
$0.001333
Official
$0.001333
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Grok Voice Latest in Console with a prompt ready to edit or send.

Help me try grok-voice-latest with a short audio request at /v1/realtime. Show the result and cost.

Use cases

  • Phone and call-center agents

    Answer inbound calls over telephony codecs, look up an account through a function tool, and hand the caller to a person when the request falls outside policy.

  • In-app voice assistant

    Let users speak to a product, with ephemeral tokens keeping the API key off the client and tools reaching the data behind the app.

  • Scripted disclosures inside a live call

    Insert a forced, fixed line for a compliance notice while the model handles the rest of the conversation.

Prompt examples

You are a booking assistant for a dental clinic. Greet the caller, ask what they need, and offer the next two open slots using the calendar tool.

Switch to Spanish when the caller does, keep answers under two sentences, and confirm the order total before ending the call.

Search our returns policy collection, answer the customer's question aloud, and say which document you used.

FAQ

Which model does grok-voice-latest run?

It is an alias that currently resolves to grok-voice-think-fast-2.0, the real-time speech-to-speech model in xAI's Grok Voice suite. xAI advises using the alias to receive improvements automatically, so the underlying version can change over time.

When should I pin a version instead of using the alias?

Pin the versioned model when you have tested prompts, voice choice, and tool behavior and need them to stay stable. Use the alias in prototypes and products where picking up upgrades without redeploying matters more than exact repeatability.

What audio formats does a session use?

Sessions carry PCM at sample rates from 8 kHz to 48 kHz, G.711 mu-law and A-law for telephony, and Opus. Audio travels as base64 JSON messages or as binary WebSocket frames.

Can the voice agent call tools?

Yes, according to xAI's voice documentation: web and X search, file search, remote MCP servers, and your own function tools defined with JSON schemas. Check that your integration passes tool definitions in the session setup.

Can I keep a conversation after a disconnect?

Yes. Session resumption caches earlier turns under a conversation ID, so a client that reconnects after a network drop can continue the same dialogue instead of starting over.

How much does Grok Voice Latest cost?

On TokenLab, Grok Voice Latest costs $0.001333 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Grok Voice Latest use?

Use https://api.tokenlab.sh/v1/realtime for Grok Voice Latest. The request example below shows the matching code shape.

Which operations does Grok Voice Latest support?

Grok Voice Latest supports Realtime. Select an operation above to see its endpoint and request example.

Compare Grok Voice Latest

Guides that use Grok Voice Latest

Sources

Reviewed Oct 2, 2026

More from Grok Voice

Related models