Alibaba: Qwen3-TTS 1.7B VoiceDesign
qwen3-tts-1-7b-voicedesign- Price
- From $0.000015
- Modalities
- Audio
About Qwen3-TTS 1.7B VoiceDesign
Qwen3-TTS 1.7B VoiceDesign is the Alibaba text-to-speech variant that creates a voice from a written description instead of a preset list. You describe emotion, tone, and prosody in natural language, and the model speaks the text in that voice. It covers ten languages and is meant for cases where no existing speaker fits the character or brand you have in mind.
Where it works well
- A voice is defined by a description, such as an elderly narrator with a slow, warm delivery, rather than chosen from a fixed list.
- Timbre, emotion, and prosody can all be steered by instruction, including strong moods like anger.
- Ten languages are covered: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian, plus dialect variants.
- Alibaba lists low synthesis latency of about 97 ms for the Qwen3-TTS family, so it fits interactive use.
When to choose another model
- The same description can yield slightly different voices on different calls, so it is a weaker choice when one exact voice must repeat for months.
- If you want a fixed named speaker, the CustomVoice variant with nine presets is more predictable.
- Quality depends on how clearly you describe the voice; vague descriptions give generic results.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to speechTokenLab endpointPOST/v1/audio/speechcurl -X POST "https://api.tokenlab.sh/v1/audio/speech" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "operation": "tts", "model": "qwen3-tts-1-7b-voicedesign", "input": "Hello from TokenLab.", "voice": "alloy" }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Text to speech
per 1M characters- Official price
- $0.000015
- Official
- $0.000015
- Discount
- —
| Official priceper 1M characters | Officialper 1M characters | Discount | |
|---|---|---|---|
| Text to speech | $0.000015 | $0.000015 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Qwen3-TTS 1.7B VoiceDesign in Console with a prompt ready to edit or send.
Help me try qwen3-tts-1-7b-voicedesign with a short audio request at /v1/audio/speech. Show the result and cost.
Use cases
Character voices for games
Describe each character, for example a gruff dwarf blacksmith or a nervous young courier, and generate lines without recruiting a voice actor first.
Brand voice exploration
Try several descriptions of a company voice, such as confident and low versus bright and quick, and compare them on the same script.
Expressive narration
Direct delivery for story sections with instructions like hushed and tense, then relieved and warm, within one project.
Prototype voices for video
Create a temporary narrator for a storyboard or explainer before the final voice is cast.
Prompt examples
Voice: a middle-aged woman, calm and low, speaking slowly like a documentary narrator. Text: The river has changed course three times in the last century.
Voice: a young man, excited and slightly out of breath, fast pace. Text: You will not believe what just came through the door.
Voice: speak in an especially angry tone, fast and clipped (用特别愤怒的语气说). Text: 这已经是第三次了,你们到底有没有在听?
FAQ
What is Qwen3-TTS VoiceDesign for?
It generates speech in a voice you specify in words. Instead of selecting a named speaker, you write a description covering emotion, tone, and delivery, and the model produces audio that follows it. It suits characters and brand voices that do not match a stock preset.
How is it different from CustomVoice?
CustomVoice offers nine fixed named speakers with optional style instructions. VoiceDesign has no fixed list and builds the voice from your description. Pick VoiceDesign to create something new and CustomVoice for a stable, repeatable speaker.
Which languages are supported?
Ten: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian, along with some dialect variants. Write the voice description in a language the model understands; the example on the model card uses Chinese.
Can I clone a real person's voice with it?
The Qwen3-TTS family describes voice cloning from a short sample, but this variant's own focus is instruction-based design. Only clone or imitate voices you have permission to use, and check the output against your own consent rules.
Will the voice be identical on every call?
Not guaranteed. A description steers the voice but does not pin one exact timbre, so repeated calls can sound slightly different. If you need exact repetition across a long project, use a fixed preset speaker instead.
How much does Qwen3-TTS 1.7B VoiceDesign cost?
On TokenLab, Qwen3-TTS 1.7B VoiceDesign costs - . The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should Qwen3-TTS 1.7B VoiceDesign use?
Use https://api.tokenlab.sh/v1/audio/speech for Qwen3-TTS 1.7B VoiceDesign. The request example below shows the matching code shape.
Which operations does Qwen3-TTS 1.7B VoiceDesign support?
Qwen3-TTS 1.7B VoiceDesign supports Text to speech. Select an operation above to see its endpoint and request example.
Compare Qwen3-TTS 1.7B VoiceDesign
Sources
Reviewed Oct 2, 2026