Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

xAI: Grok STT

Grok STT is xAI's speech-to-text service for uploaded recordings. This integration exposes batch transcription with the reported audio duration; the real-time streaming service uses a separate interface and pricing contract.
Compare models
grok-stt
AvailablexAISpeech to TextSync
Price
From $0.000028
Modalities
Audio

About Grok STT

Grok STT is xAI's speech-to-text service for turning recorded audio into text. This model handles uploaded recordings in batch and returns the transcript as JSON. xAI documents support for 25 languages in its transcription service. Real-time streaming transcription is a separate xAI interface, so this model suits files such as meetings, calls and interviews that already exist.

Where it works well

  • Transcribes uploaded audio files in one request and returns a JSON result.
  • xAI documents support for 25 languages, covering many multilingual recordings.
  • Fits batch jobs like meeting archives, call recordings and interview backlogs.
  • Works through the standard audio transcription request format, so existing Whisper-style clients need few changes.
  • Long recordings are handled as a single file upload rather than split into short chunks.

When to choose another model

  • Batch only; live captioning and low-latency streaming need a separate real-time interface.
  • Output is a transcript in JSON, with no summarization or translation built in.
  • Accuracy depends on audio quality, so noisy or overlapping speech may need cleanup or review.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Speech to textTokenLab endpoint
    POST/v1/audio/transcriptions
    curl -X POST "https://api.tokenlab.sh/v1/audio/transcriptions" \
      -H "Authorization: Bearer sk-xxx" \
      -F file="@audio.mp3" \
      -F operation="stt" \
      -F response_format="json" \
      -F model="grok-stt" \
      -F language="en"

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Speech to text

per second
Official price
$0.00002778
Official
$0.00002778
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Grok STT in Console with a prompt ready to edit or send.

Help me try grok-stt with a short audio request at /v1/audio/transcriptions. Show the result and cost.

Use cases

  • Meeting transcripts

    Upload recorded calls or meetings and store the text so teams can search them or feed them to a summarizing model.

  • Interview and podcast archives

    Convert a backlog of recordings to text for editing, quotes and show notes.

  • Support call review

    Transcribe customer calls in bulk so quality teams can scan for issues and topics.

Prompt examples

Transcribe this 45-minute team meeting recording in English.

Transcribe the attached customer call. The speakers talk in Spanish.

Convert this recorded lecture to text so I can search it for the section on photosynthesis.

FAQ

What is Grok STT?

It is xAI's speech-to-text service. It accepts an uploaded recording and returns the transcript as JSON, working in batch mode rather than live. Use it for recordings you already have, such as meetings, calls and interviews.

Which languages does it support?

xAI documents 25 languages for its transcription service. Test a short sample in your language first, since accuracy varies with language and recording quality. Accents also affect results.

Can Grok STT transcribe live audio?

No. It handles uploaded files only, so use it after a call or meeting has ended. xAI offers streaming transcription through a separate interface, and live captioning needs that interface instead of this model.

What format do I get back?

A JSON response containing the transcript text. You can pass that text on to a language model for summaries, action items, or translation into another language.

How is it different from Whisper?

Both transcribe uploaded audio through the same request style. They differ in language coverage and accuracy by audio type, so run both on a sample of your recordings and compare the transcripts.

How much does Grok STT cost?

On TokenLab, Grok STT costs $0.00002778 per second. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Grok STT use?

Use https://api.tokenlab.sh/v1/audio/transcriptions for Grok STT. The request example below shows the matching code shape.

Which operations does Grok STT support?

Grok STT supports Speech to text. Select an operation above to see its endpoint and request example.

Compare Grok STT

Sources

Reviewed Oct 2, 2026

More from Grok Voice

Related models