Audio & Realtime
Create Speech
Generates audio from the input text
Use GET /v1/models?recommended_for=tts to get current speech models.
Request Body
The response mode depends on the selected model and stream_format. For longer input, allow an HTTP client timeout of at least 120s.
Optional fields, voices, output formats, and SSE support depend on the selected model. Check GET /v1/models/{model} before sending them. Unsupported parameters can return 400 unsupported_parameter.
tts-1TTS model ID. Find current choices with GET /v1/models?recommended_for=tts.
The text to generate audio for. Maximum 4096 characters.
Voice selector. Pass a built-in voice name such as nova, a Gemini voice such as Kore, or an object like { "id": "voice-id" } for compatible custom voices.
Voice ID for models that support this parameter, such as compatible MiniMax speech models.
Optional style or delivery instructions for OpenAI-compatible TTS models that support them.
Optional speaking style prompt for Gemini TTS models.
Language code, such as en-US, for models that support it.
Audio format. Common values include mp3, opus, aac, flac, wav, and pcm; supported values vary by model family.
audioTokenLab delivery format: audio or sse. stream_format=sse is not supported for tts-1 or tts-1-hd.
Speech speed for model families that support it (0.25 to 4.0).
Sampling temperature for Gemini TTS, from 0 to 2.
Response
With stream_format=audio, the response contains audio bytes. Models that support stream_format=sse return a text/event-stream response; parse its events instead of saving the entire response as an audio file.
Request
curl -X POST "https://api.tokenlab.sh/v1/audio/speech" \
-H "Authorization: Bearer sk-your-api-key" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1-hd",
"voice": "nova",
"input": "Hello, welcome to TokenLab!"
}' \
--output speech.mp3Voice Samples
Voice availability depends on the model; the names below are examples, not a shared voice list for every TTS model.
| Voice | Description |
|---|---|
alloy | Neutral, balanced |
ash | Calm, measured |
ballad | Melodic, expressive |
coral | Warm, inviting |
echo | Warm, conversational |
fable | Expressive, narrative |
nova | Friendly, clear |
onyx | Deep, authoritative |
sage | Wise, thoughtful |
shimmer | Soft, gentle |
verse | Dynamic, versatile |
Response
<binary audio data>Important fields
HTTP response media type: the audio format, or text/event-stream for SSE.
Audio bytes for audio; an event stream for sse. Choose the parser using Content-Type and the requested mode.
MiniMax Speech 2.8 Turbo synthesizes a script in a selected voice with a speed-focused workflow, useful for responsive spoken interfaces and narration.
{
"model": "speech-2.8-turbo",
"input": "Welcome to TokenLab.",
"voice_id": "male-qn-qingse",
"speed": 1,
"response_format": "mp3",
"stream_format": "audio"
}Authorization
BearerAuth API Key authentication. Create or manage API keys in Dashboard > API > API Keys.
In: header
Headers
Per-request Delivery policy. Overrides the API key and Workspace defaults. Auto tries TokenLab Verified first and may switch once to Official only before output, request acceptance, or persistent resource creation.
Value in
- "auto"
- "verified"
- "official"
Request Body
application/json
Response
application/json
application/json
application/json
application/json
application/json
application/json