Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

xAI: Grok Voice STT

Grok Voice STT converts recorded audio into text. Language hints and timestamp controls make it useful for transcripts that will feed captioning, search, or a later text-analysis step.
Compare models
grok-voice-stt
AvailablexAISpeech to TextSync
Price
From $0.000028
Modalities
Audio

About Grok Voice STT

Grok Voice STT is xAI's speech-to-text model, turning recorded or streamed audio into text transcripts. It accepts common audio file formats, returns word-level timestamps, can separate speakers and channels, and formats numbers and currencies by language. Use it for captions, searchable meeting records, and call transcripts that feed later text analysis.

Where it works well

  • Word-level timestamps make it direct to build captions, jump-to-moment search, and subtitle files from the transcript.
  • Speaker diarization and multichannel transcription separate who said what in meetings and two-sided call recordings.
  • Language hints and a list of custom key terms steer recognition toward product names and jargon.
  • Text formatting renders numbers, dates, and currencies the way the spoken language writes them.

When to choose another model

  • It produces transcripts only; summarizing, translating, or answering from the audio needs a text model after it.
  • Very noisy recordings may need the voice-detection threshold tuned or cleaner input before accuracy is acceptable.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Speech to textTokenLab endpoint
    POST/v1/audio/transcriptions
    curl -X POST "https://api.tokenlab.sh/v1/audio/transcriptions" \
      -H "Authorization: Bearer sk-xxx" \
      -F file="@audio.mp3" \
      -F operation="stt" \
      -F response_format="json" \
      -F model="grok-voice-stt" \
      -F language="en"

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Pricing

per second
Official
$0.000028per second
Official price
$0.000028per second

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Grok Voice STT in Console with a prompt ready to edit or send.

Help me try grok-voice-stt with a short audio request at /v1/audio/transcriptions. Show the result and cost.

Use cases

  • Meeting transcripts

    Transcribe a recorded call with speakers labeled, then pass the text to a summarizer for action items.

  • Subtitles for video

    Generate timed captions from an uploaded video's audio track using the word timings.

  • Call-center quality review

    Transcribe two-channel recordings with agent and customer on separate channels, then scan for compliance phrases.

Prompt examples

Transcribe this 45-minute podcast episode in English with word timestamps and speakers labeled.

Transcribe the attached customer call, using the key terms 'Zyntra', 'OmniPlan', and 'SLA' to bias recognition.

Convert this Spanish voice memo to text and format the amounts as currency.

FAQ

What does Grok Voice STT do?

It converts audio into text. You send a file or stream audio, and it returns a transcript that can include word-level timestamps, speaker labels, and per-channel text, depending on the options you set.

Which audio formats can it read?

xAI lists twelve formats, including MP3, WAV, FLAC, and Opus, with a file size ceiling of 500 MB. Raw PCM, mu-law, and A-law audio need explicit encoding parameters.

Does it support live transcription?

Yes. A WebSocket endpoint streams interim results and uses Smart Turn end-of-turn detection to tell natural pauses from the end of a sentence, which cuts premature cut-offs.

How many languages does it cover?

xAI documents more than 25 languages for transcription and language-specific formatting. Pass a language hint when you know the language, so recognition does not have to guess.

Can it tell speakers apart?

Yes, diarization labels speakers within a recording, and multichannel mode handles up to eight separate channels, which suits call recordings with one party per channel.

How much does Grok Voice STT cost?

On TokenLab, Grok Voice STT costs $0.000028 per second. The pricing table above shows the full breakdown.

Which endpoint should Grok Voice STT use?

Use https://api.tokenlab.sh/v1/audio/transcriptions for Grok Voice STT. The request example below shows the matching code shape.

Which operations does Grok Voice STT support?

Grok Voice STT supports Speech to text. Select an operation above to see its endpoint and request example.

Compare Grok Voice STT

Sources

Reviewed Oct 2, 2026

More from Grok Voice

Related models