TokenLab

Audio & Realtime

Create Speech

Generates audio from the input text

POST
/v1/audio/speech

Use GET /v1/models?recommended_for=tts to get current speech models.

Request Body

The response mode depends on the selected model and stream_format. For longer input, allow an HTTP client timeout of at least 120s.

Optional fields, voices, output formats, and SSE support depend on the selected model. Check GET /v1/models/{model} before sending them. Unsupported parameters can return 400 unsupported_parameter.

modelstringdefault: tts-1

TTS model ID. Find current choices with GET /v1/models?recommended_for=tts.

inputstringrequired

The text to generate audio for. Maximum 4096 characters.

voicestring | object

Voice selector. Pass a built-in voice name such as nova, a Gemini voice such as Kore, or an object like { "id": "voice-id" } for compatible custom voices.

voice_idstring

Voice ID for models that support this parameter, such as compatible MiniMax speech models.

instructionsstring

Optional style or delivery instructions for OpenAI-compatible TTS models that support them.

promptstring

Optional speaking style prompt for Gemini TTS models.

language_codestring

Language code, such as en-US, for models that support it.

response_formatstring

Audio format. Common values include mp3, opus, aac, flac, wav, and pcm; supported values vary by model family.

stream_formatstringdefault: audio

TokenLab delivery format: audio or sse. stream_format=sse is not supported for tts-1 or tts-1-hd.

speednumber

Speech speed for model families that support it (0.25 to 4.0).

temperaturenumber

Sampling temperature for Gemini TTS, from 0 to 2.

Response

With stream_format=audio, the response contains audio bytes. Models that support stream_format=sse return a text/event-stream response; parse its events instead of saving the entire response as an audio file.

Request

curl -X POST "https://api.tokenlab.sh/v1/audio/speech" \
  -H "Authorization: Bearer sk-your-api-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1-hd",
    "voice": "nova",
    "input": "Hello, welcome to TokenLab!"
  }' \
  --output speech.mp3

Voice Samples

Voice availability depends on the model; the names below are examples, not a shared voice list for every TTS model.

VoiceDescription
alloyNeutral, balanced
ashCalm, measured
balladMelodic, expressive
coralWarm, inviting
echoWarm, conversational
fableExpressive, narrative
novaFriendly, clear
onyxDeep, authoritative
sageWise, thoughtful
shimmerSoft, gentle
verseDynamic, versatile

Response example

Response

200 OK
<binary audio data>

Important fields

Content-Typestring

HTTP response media type: the audio format, or text/event-stream for SSE.

bodybinary | SSE

Audio bytes for audio; an event stream for sse. Choose the parser using Content-Type and the requested mode.

MiniMax Speech 2.8 Turbo synthesizes a script in a selected voice with a speed-focused workflow, useful for responsive spoken interfaces and narration.

{
  "model": "speech-2.8-turbo",
  "input": "Welcome to TokenLab.",
  "voice_id": "male-qn-qingse",
  "speed": 1,
  "response_format": "mp3",
  "stream_format": "audio"
}

Authorization

BearerAuth
AuthorizationBearer <token>

API Key authentication. Create or manage API keys in Dashboard > API > API Keys.

In: header

Headers

X-TokenLab-Delivery-Policy?string

Per-request Delivery policy. Overrides the API key and Workspace defaults. Auto tries TokenLab Verified first and may switch once to Official only before output, request acceptance, or persistent resource creation.

Value in

  • "auto"
  • "verified"
  • "official"

Request Body

application/json

Response

application/json

application/json

application/json

application/json

application/json

application/json