TokenLab

Audio & Realtime

Create Transcription

Transcribes audio into the input language

POST
/v1/audio/transcriptions

Use GET /v1/models?recommended_for=stt to get current transcription models.

Request Body

Transcription can return the transcript immediately or accept an asynchronous task with HTTP 200, id, and status. For an accepted task, follow poll_url, or query /v1/tasks/{id}. For longer synchronous audio, allow a client timeout of at least 120s.

filefilerequired

Audio file to transcribe. Supported formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm. Maximum file size: 25 MB.

modelstringdefault: whisper-1

Transcription model ID, for example whisper-1 or gpt-4o-transcribe. Find current choices with GET /v1/models?recommended_for=stt.

languagestring

Language of the audio in ISO-639-1 format (e.g., en, zh, ja).

promptstring

Optional text to guide the model's style or continue a previous segment.

response_formatstringdefault: json

Output format. Whisper supports json, text, srt, verbose_json, and vtt; other models may support a different subset. Check the selected model's details.

temperaturenumber

Sampling temperature from 0 to 1, only for models that support this parameter.

timestamp_granularitiesarray

word and/or segment, when supported by the selected model. For Whisper, use response_format=verbose_json. In multipart requests, repeat timestamp_granularities[] for multiple values.

Response

The fields below describe JSON transcript or translation responses. text, srt, and vtt return raw text or subtitles rather than a JSON object. Additional fields depend on the model and output format.

textstring

The transcribed text.

For verbose_json:

taskstring

Task name when returned, such as transcribe.

languagestring

Detected language.

durationnumber

Audio duration in seconds.

segmentsarray

Transcription segments with timestamps.

wordsarray

Word-level timestamps (if requested).

An accepted asynchronous transcription task returns the following fields instead of the transcript:

idstring

Task ID.

task_idstring

Async task identifier alias when provided.

statusstring

Task status: pending, processing, completed, or failed.

poll_urlstring

Preferred polling URL when the create response supplies one.

Request

curl -X POST "https://api.tokenlab.sh/v1/audio/transcriptions" \
  -H "Authorization: Bearer sk-your-api-key" \
  -F file="@audio.mp3" \
  -F model="whisper-1" \
  -F language="en"

Response

{
  "text": "Hello, this is a test of the transcription API."
}

Translation

To translate audio to English, use the translations endpoint:

with open("audio.mp3", "rb") as audio_file:
    response = client.audio.translations.create(
        model="whisper-1",
        file=audio_file
    )

print(response.text)

Authorization

BearerAuth
AuthorizationBearer <token>

API Key authentication. Create or manage API keys in Dashboard > API > API Keys.

In: header

Headers

X-TokenLab-Delivery-Policy?string

Per-request Delivery policy. Overrides the API key and Workspace defaults. Auto tries TokenLab Verified first and may switch once to Official only before output, request acceptance, or persistent resource creation.

Value in

  • "auto"
  • "verified"
  • "official"

Request Body

multipart/form-data

Response

application/json

application/json

application/json

application/json

application/json

application/json