xAI: Grok Voice Latest
grok-voice-latest- Price
- From $0.001333
- Released
- Apr 23, 2026
- Modalities
- Audio
About Grok Voice Latest
Grok Voice Latest is xAI's floating alias for its current real-time speech-to-speech model, used to hold a spoken conversation over a WebSocket session. At the time of review it points to grok-voice-think-fast-2.0. Choose it when a voice assistant should pick up xAI's improvements without a code change; pin the versioned model when behavior must stay fixed.
Where it works well
- It is the alias xAI recommends for voice agents that should receive model upgrades automatically, with no change to the session configuration.
- A session accepts built-in or custom voices, server-side voice activity detection for turn-taking, and language hints for transcription.
- Agents can call web and X search, file search over document collections, remote MCP servers, and custom function tools.
- Sessions can resume after a dropped connection by reusing a conversation ID, which suits phone and mobile clients.
When to choose another model
- Because the alias moves, voice, pacing, or wording can shift after an upgrade; pin the versioned model for regression-tested products.
- It handles live audio sessions only; transcribing stored recordings or rendering prepared text needs the speech-to-text or text-to-speech models.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
RealtimeWebSocketWS/v1/realtime// npm install ws import WebSocket from 'ws'; const socket = new WebSocket("wss://api.tokenlab.sh/v1/realtime?model=grok-voice-latest", { headers: { Authorization: "Bearer sk-xxx" } }); socket.on('open', () => { // Send the input events supported by this model with socket.send(...). // https://tokenlab.sh/docs/en/api-reference/realtime/connect }); socket.on('message', (data) => console.log(data.toString())); socket.on('error', (error) => console.error(error)); socket.on('close', (code, reason) => console.log(code, reason.toString())); // Close the session when finished. process.on('SIGINT', () => socket.close());
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Audio Duration
per second- Official price
- $0.001333
- Official
- $0.001333
- Discount
- —
| Official priceper second | Officialper second | Discount | |
|---|---|---|---|
| Audio Duration | $0.001333 | $0.001333 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Grok Voice Latest in Console with a prompt ready to edit or send.
Help me try grok-voice-latest with a short audio request at /v1/realtime. Show the result and cost.
Use cases
Phone and call-center agents
Answer inbound calls over telephony codecs, look up an account through a function tool, and hand the caller to a person when the request falls outside policy.
In-app voice assistant
Let users speak to a product, with ephemeral tokens keeping the API key off the client and tools reaching the data behind the app.
Scripted disclosures inside a live call
Insert a forced, fixed line for a compliance notice while the model handles the rest of the conversation.
Prompt examples
You are a booking assistant for a dental clinic. Greet the caller, ask what they need, and offer the next two open slots using the calendar tool.
Switch to Spanish when the caller does, keep answers under two sentences, and confirm the order total before ending the call.
Search our returns policy collection, answer the customer's question aloud, and say which document you used.
FAQ
Which model does grok-voice-latest run?
It is an alias that currently resolves to grok-voice-think-fast-2.0, the real-time speech-to-speech model in xAI's Grok Voice suite. xAI advises using the alias to receive improvements automatically, so the underlying version can change over time.
When should I pin a version instead of using the alias?
Pin the versioned model when you have tested prompts, voice choice, and tool behavior and need them to stay stable. Use the alias in prototypes and products where picking up upgrades without redeploying matters more than exact repeatability.
What audio formats does a session use?
Sessions carry PCM at sample rates from 8 kHz to 48 kHz, G.711 mu-law and A-law for telephony, and Opus. Audio travels as base64 JSON messages or as binary WebSocket frames.
Can the voice agent call tools?
Yes, according to xAI's voice documentation: web and X search, file search, remote MCP servers, and your own function tools defined with JSON schemas. Check that your integration passes tool definitions in the session setup.
Can I keep a conversation after a disconnect?
Yes. Session resumption caches earlier turns under a conversation ID, so a client that reconnects after a network drop can continue the same dialogue instead of starting over.
How much does Grok Voice Latest cost?
On TokenLab, Grok Voice Latest costs $0.001333 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should Grok Voice Latest use?
Use https://api.tokenlab.sh/v1/realtime for Grok Voice Latest. The request example below shows the matching code shape.
Which operations does Grok Voice Latest support?
Grok Voice Latest supports Realtime. Select an operation above to see its endpoint and request example.
Compare Grok Voice Latest
Guides that use Grok Voice Latest
Sources
Reviewed Oct 2, 2026