xAI: Grok Voice STT
grok-voice-stt- Price
- From $0.000028
- Modalities
- Audio
About Grok Voice STT
Grok Voice STT is xAI's speech-to-text model, turning recorded or streamed audio into text transcripts. It accepts common audio file formats, returns word-level timestamps, can separate speakers and channels, and formats numbers and currencies by language. Use it for captions, searchable meeting records, and call transcripts that feed later text analysis.
Where it works well
- Word-level timestamps make it direct to build captions, jump-to-moment search, and subtitle files from the transcript.
- Speaker diarization and multichannel transcription separate who said what in meetings and two-sided call recordings.
- Language hints and a list of custom key terms steer recognition toward product names and jargon.
- Text formatting renders numbers, dates, and currencies the way the spoken language writes them.
When to choose another model
- It produces transcripts only; summarizing, translating, or answering from the audio needs a text model after it.
- Very noisy recordings may need the voice-detection threshold tuned or cleaner input before accuracy is acceptable.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Speech to textTokenLab endpointPOST/v1/audio/transcriptionscurl -X POST "https://api.tokenlab.sh/v1/audio/transcriptions" \ -H "Authorization: Bearer sk-xxx" \ -F file="@audio.mp3" \ -F operation="stt" \ -F response_format="json" \ -F model="grok-voice-stt" \ -F language="en"
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Pricing
per second- Official
- $0.000028per second
- Official price
- $0.000028per second
| Official priceper second | Officialper second | Discount | |
|---|---|---|---|
| Price | $0.000028per second | $0.000028per second | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Grok Voice STT in Console with a prompt ready to edit or send.
Help me try grok-voice-stt with a short audio request at /v1/audio/transcriptions. Show the result and cost.
Use cases
Meeting transcripts
Transcribe a recorded call with speakers labeled, then pass the text to a summarizer for action items.
Subtitles for video
Generate timed captions from an uploaded video's audio track using the word timings.
Call-center quality review
Transcribe two-channel recordings with agent and customer on separate channels, then scan for compliance phrases.
Prompt examples
Transcribe this 45-minute podcast episode in English with word timestamps and speakers labeled.
Transcribe the attached customer call, using the key terms 'Zyntra', 'OmniPlan', and 'SLA' to bias recognition.
Convert this Spanish voice memo to text and format the amounts as currency.
FAQ
What does Grok Voice STT do?
It converts audio into text. You send a file or stream audio, and it returns a transcript that can include word-level timestamps, speaker labels, and per-channel text, depending on the options you set.
Which audio formats can it read?
xAI lists twelve formats, including MP3, WAV, FLAC, and Opus, with a file size ceiling of 500 MB. Raw PCM, mu-law, and A-law audio need explicit encoding parameters.
Does it support live transcription?
Yes. A WebSocket endpoint streams interim results and uses Smart Turn end-of-turn detection to tell natural pauses from the end of a sentence, which cuts premature cut-offs.
How many languages does it cover?
xAI documents more than 25 languages for transcription and language-specific formatting. Pass a language hint when you know the language, so recognition does not have to guess.
Can it tell speakers apart?
Yes, diarization labels speakers within a recording, and multichannel mode handles up to eight separate channels, which suits call recordings with one party per channel.
How much does Grok Voice STT cost?
On TokenLab, Grok Voice STT costs $0.000028 per second. The pricing table above shows the full breakdown.
Which endpoint should Grok Voice STT use?
Use https://api.tokenlab.sh/v1/audio/transcriptions for Grok Voice STT. The request example below shows the matching code shape.
Which operations does Grok Voice STT support?
Grok Voice STT supports Speech to text. Select an operation above to see its endpoint and request example.
Compare Grok Voice STT
Sources
Reviewed Oct 2, 2026