Audio & Realtime
Create Transcription
Transcribes audio into the input language
Use GET /v1/models?recommended_for=stt to get current transcription models.
Request Body
Transcription can return the transcript immediately or accept an asynchronous task with HTTP 200, id, and status. For an accepted task, follow poll_url, or query /v1/tasks/{id}. For longer synchronous audio, allow a client timeout of at least 120s.
Audio file to transcribe. Supported formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm. Maximum file size: 25 MB.
whisper-1Transcription model ID, for example whisper-1 or gpt-4o-transcribe. Find current choices with GET /v1/models?recommended_for=stt.
Language of the audio in ISO-639-1 format (e.g., en, zh, ja).
Optional text to guide the model's style or continue a previous segment.
jsonOutput format. Whisper supports json, text, srt, verbose_json, and vtt; other models may support a different subset. Check the selected model's details.
Sampling temperature from 0 to 1, only for models that support this parameter.
word and/or segment, when supported by the selected model. For Whisper, use response_format=verbose_json. In multipart requests, repeat timestamp_granularities[] for multiple values.
Response
The fields below describe JSON transcript or translation responses. text, srt, and vtt return raw text or subtitles rather than a JSON object. Additional fields depend on the model and output format.
The transcribed text.
For verbose_json:
Task name when returned, such as transcribe.
Detected language.
Audio duration in seconds.
Transcription segments with timestamps.
Word-level timestamps (if requested).
An accepted asynchronous transcription task returns the following fields instead of the transcript:
Task ID.
Async task identifier alias when provided.
Task status: pending, processing, completed, or failed.
Preferred polling URL when the create response supplies one.
Request
curl -X POST "https://api.tokenlab.sh/v1/audio/transcriptions" \
-H "Authorization: Bearer sk-your-api-key" \
-F file="@audio.mp3" \
-F model="whisper-1" \
-F language="en"Response
{
"text": "Hello, this is a test of the transcription API."
}Translation
To translate audio to English, use the translations endpoint:
with open("audio.mp3", "rb") as audio_file:
response = client.audio.translations.create(
model="whisper-1",
file=audio_file
)
print(response.text)Authorization
BearerAuth API Key authentication. Create or manage API keys in Dashboard > API > API Keys.
In: header
Headers
Per-request Delivery policy. Overrides the API key and Workspace defaults. Auto tries TokenLab Verified first and may switch once to Official only before output, request acceptance, or persistent resource creation.
Value in
- "auto"
- "verified"
- "official"
Request Body
multipart/form-data
Response
application/json
application/json
application/json
application/json
application/json
application/json