Media Guides
Audio and realtime
Create speech, transcribe audio, translate speech, or open a realtime session
Use the audio endpoints when the result is a file or text. Use the realtime WebSocket endpoint for a live, two-way audio experience.
Choose an endpoint
| Need | Endpoint | Use it when |
|---|---|---|
| Text to speech | POST /v1/audio/speech | You need an audio file from text. |
| Transcription | POST /v1/audio/transcriptions | You need text from an audio file. |
| Audio translation | POST /v1/audio/translations | You need translated text from an audio file. |
| Realtime WebSocket | WS /v1/realtime?model={model} | You need bidirectional streaming audio or realtime multimodal events over WebSocket. Plain GET /v1/realtime only returns metadata. |
Choose a model
Read the current model list instead of hard-coding it. Check the selected model's details for realtime support before opening a socket.
curl "https://api.tokenlab.sh/v1/models?recommended_for=tts" \
-H "Authorization: Bearer sk-your-api-key"
curl "https://api.tokenlab.sh/v1/models?recommended_for=stt" \
-H "Authorization: Bearer sk-your-api-key"Audio requests
Speech returns audio in the HTTP response. Transcription and translation can return final text or an accepted task. When the body contains a task ID and status instead of final text, save the ID and follow poll_url until completion or failure. Allow enough time for large files and keep the Request ID.
Eligible generated audio HTTP(S) result URLs may be retained as media copies for 30 days. Check media_retention.items for each item's status and expires_at; pending or failed copies are not guaranteed. Inline or binary audio responses are outside this URL-copy retention. See Data retention.
curl -X POST "https://api.tokenlab.sh/v1/audio/speech" \
-H "Authorization: Bearer sk-your-api-key" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1-hd",
"voice": "nova",
"input": "Welcome to TokenLab."
}' \
--output speech.mp3Realtime Sessions
Open a WebSocket with the model in the query string and the API key in the Authorization header. Keep the event format documented for the selected realtime model, and close the socket when the session is complete.
This guide covers the WebSocket subset only. TokenLab does not currently provide OpenAI Realtime REST client secret, translation client secret, Calls, or legacy beta session management endpoints.
For the server-side example below, install ws and set TOKENLAB_REALTIME_MODEL to a current model whose detail declares /v1/realtime. Provide TOKENLAB_API_KEY in the server environment; the name alone does not establish realtime support.
npm install wsimport WebSocket from 'ws';
const model = process.env.TOKENLAB_REALTIME_MODEL;
const apiKey = process.env.TOKENLAB_API_KEY;
if (!model || !apiKey) throw new Error('Set TOKENLAB_REALTIME_MODEL and TOKENLAB_API_KEY');
const url = new URL('wss://api.tokenlab.sh/v1/realtime');
url.searchParams.set('model', model);
const socket = new WebSocket(url, {
headers: { Authorization: `Bearer ${apiKey}` },
});
socket.on('message', (event) => console.log(event.toString()));
socket.on('error', (error) => console.error(error.message));
process.once('SIGINT', () => socket.close(1000, 'Client stopped'));Show progress safely
- Save generated audio files instead of replaying the same request on refresh.
- For transcription and translation, show upload and processing states even when the API call is synchronous.
- For realtime, handle close events and reconnect only after the user starts a new session.
- Do not put API keys, private URLs, or account secrets in audio text input.
API Reference
| Topic | Reference |
|---|---|
| Create Speech | Create Speech |
| Create Transcription | Create Transcription |
| Create Translation | Create Translation |
| Realtime WebSocket | Realtime WebSocket |
| List Models | List Models |
| Billing & Pricing | Billing & Pricing |