TokenLab

Media Guides

Audio and realtime

Create speech, transcribe audio, translate speech, or open a realtime session

Use the audio endpoints when the result is a file or text. Use the realtime WebSocket endpoint for a live, two-way audio experience.

Choose an endpoint

NeedEndpointUse it when
Text to speechPOST /v1/audio/speechYou need an audio file from text.
TranscriptionPOST /v1/audio/transcriptionsYou need text from an audio file.
Audio translationPOST /v1/audio/translationsYou need translated text from an audio file.
Realtime WebSocketWS /v1/realtime?model={model}You need bidirectional streaming audio or realtime multimodal events over WebSocket. Plain GET /v1/realtime only returns metadata.

Choose a model

Read the current model list instead of hard-coding it. Check the selected model's details for realtime support before opening a socket.

curl "https://api.tokenlab.sh/v1/models?recommended_for=tts" \
  -H "Authorization: Bearer sk-your-api-key"

curl "https://api.tokenlab.sh/v1/models?recommended_for=stt" \
  -H "Authorization: Bearer sk-your-api-key"

Audio requests

Speech returns audio in the HTTP response. Transcription and translation can return final text or an accepted task. When the body contains a task ID and status instead of final text, save the ID and follow poll_url until completion or failure. Allow enough time for large files and keep the Request ID.

Eligible generated audio HTTP(S) result URLs may be retained as media copies for 30 days. Check media_retention.items for each item's status and expires_at; pending or failed copies are not guaranteed. Inline or binary audio responses are outside this URL-copy retention. See Data retention.

curl -X POST "https://api.tokenlab.sh/v1/audio/speech" \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1-hd",
    "voice": "nova",
    "input": "Welcome to TokenLab."
  }' \
  --output speech.mp3

Realtime Sessions

Open a WebSocket with the model in the query string and the API key in the Authorization header. Keep the event format documented for the selected realtime model, and close the socket when the session is complete.

This guide covers the WebSocket subset only. TokenLab does not currently provide OpenAI Realtime REST client secret, translation client secret, Calls, or legacy beta session management endpoints.

For the server-side example below, install ws and set TOKENLAB_REALTIME_MODEL to a current model whose detail declares /v1/realtime. Provide TOKENLAB_API_KEY in the server environment; the name alone does not establish realtime support.

npm install ws
import WebSocket from 'ws';

const model = process.env.TOKENLAB_REALTIME_MODEL;
const apiKey = process.env.TOKENLAB_API_KEY;
if (!model || !apiKey) throw new Error('Set TOKENLAB_REALTIME_MODEL and TOKENLAB_API_KEY');

const url = new URL('wss://api.tokenlab.sh/v1/realtime');
url.searchParams.set('model', model);
const socket = new WebSocket(url, {
  headers: { Authorization: `Bearer ${apiKey}` },
});
socket.on('message', (event) => console.log(event.toString()));
socket.on('error', (error) => console.error(error.message));
process.once('SIGINT', () => socket.close(1000, 'Client stopped'));

Show progress safely

  • Save generated audio files instead of replaying the same request on refresh.
  • For transcription and translation, show upload and processing states even when the API call is synchronous.
  • For realtime, handle close events and reconnect only after the user starts a new session.
  • Do not put API keys, private URLs, or account secrets in audio text input.

API Reference

TopicReference
Create SpeechCreate Speech
Create TranscriptionCreate Transcription
Create TranslationCreate Translation
Realtime WebSocketRealtime WebSocket
List ModelsList Models
Billing & PricingBilling & Pricing

On this page