xAI: Grok STT
grok-stt- Price
- From $0.000028
- Modalities
- Audio
About Grok STT
Grok STT is xAI's speech-to-text service for turning recorded audio into text. This model handles uploaded recordings in batch and returns the transcript as JSON. xAI documents support for 25 languages in its transcription service. Real-time streaming transcription is a separate xAI interface, so this model suits files such as meetings, calls and interviews that already exist.
Where it works well
- Transcribes uploaded audio files in one request and returns a JSON result.
- xAI documents support for 25 languages, covering many multilingual recordings.
- Fits batch jobs like meeting archives, call recordings and interview backlogs.
- Works through the standard audio transcription request format, so existing Whisper-style clients need few changes.
- Long recordings are handled as a single file upload rather than split into short chunks.
When to choose another model
- Batch only; live captioning and low-latency streaming need a separate real-time interface.
- Output is a transcript in JSON, with no summarization or translation built in.
- Accuracy depends on audio quality, so noisy or overlapping speech may need cleanup or review.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Speech to textTokenLab endpointPOST/v1/audio/transcriptionscurl -X POST "https://api.tokenlab.sh/v1/audio/transcriptions" \ -H "Authorization: Bearer sk-xxx" \ -F file="@audio.mp3" \ -F operation="stt" \ -F response_format="json" \ -F model="grok-stt" \ -F language="en"
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Speech to text
per second- Official price
- $0.00002778
- Official
- $0.00002778
- Discount
- —
| Official priceper second | Officialper second | Discount | |
|---|---|---|---|
| Speech to text | $0.00002778 | $0.00002778 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Grok STT in Console with a prompt ready to edit or send.
Help me try grok-stt with a short audio request at /v1/audio/transcriptions. Show the result and cost.
Use cases
Meeting transcripts
Upload recorded calls or meetings and store the text so teams can search them or feed them to a summarizing model.
Interview and podcast archives
Convert a backlog of recordings to text for editing, quotes and show notes.
Support call review
Transcribe customer calls in bulk so quality teams can scan for issues and topics.
Prompt examples
Transcribe this 45-minute team meeting recording in English.
Transcribe the attached customer call. The speakers talk in Spanish.
Convert this recorded lecture to text so I can search it for the section on photosynthesis.
FAQ
What is Grok STT?
It is xAI's speech-to-text service. It accepts an uploaded recording and returns the transcript as JSON, working in batch mode rather than live. Use it for recordings you already have, such as meetings, calls and interviews.
Which languages does it support?
xAI documents 25 languages for its transcription service. Test a short sample in your language first, since accuracy varies with language and recording quality. Accents also affect results.
Can Grok STT transcribe live audio?
No. It handles uploaded files only, so use it after a call or meeting has ended. xAI offers streaming transcription through a separate interface, and live captioning needs that interface instead of this model.
What format do I get back?
A JSON response containing the transcript text. You can pass that text on to a language model for summaries, action items, or translation into another language.
How is it different from Whisper?
Both transcribe uploaded audio through the same request style. They differ in language coverage and accuracy by audio type, so run both on a sample of your recordings and compare the transcripts.
How much does Grok STT cost?
On TokenLab, Grok STT costs $0.00002778 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should Grok STT use?
Use https://api.tokenlab.sh/v1/audio/transcriptions for Grok STT. The request example below shows the matching code shape.
Which operations does Grok STT support?
Grok STT supports Speech to text. Select an operation above to see its endpoint and request example.
Compare Grok STT
Sources
Reviewed Oct 2, 2026