Alibaba: Qwen3-TTS 1.7B CustomVoice
qwen3-tts-1-7b-customvoice- Price
- From $0.000015
- Modalities
- Audio
About Qwen3-TTS 1.7B CustomVoice
Qwen3-TTS 1.7B CustomVoice is Alibaba's text-to-speech model built around a fixed set of nine preset voices. You pick a speaker by name and can add a plain-language instruction about tone, emotion, or speaking style. It speaks ten languages, including Chinese, English, Japanese, Korean, and several European languages, and Alibaba quotes end-to-end synthesis latency as low as 97 ms for interactive use.
Where it works well
- Nine named premium speakers cover different genders, ages, languages, and dialects, so a voice can stay the same across a whole product.
- A written instruction can steer emotion and manner of speaking without changing the chosen speaker.
- Ten languages are supported: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian.
- Alibaba reports end-to-end latency as low as 97 ms, which suits voice assistants and live narration.
When to choose another model
- The voice list is fixed; to invent a new voice from a description, the VoiceDesign variant is the better fit.
- Results are best when a speaker reads its own native language, so an English preset reading Japanese may sound accented.
- Noisy or unusual input text, such as heavy markup or mixed symbols, can reduce output quality.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to speechTokenLab endpointPOST/v1/audio/speechcurl -X POST "https://api.tokenlab.sh/v1/audio/speech" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "operation": "tts", "model": "qwen3-tts-1-7b-customvoice", "input": "Hello from TokenLab.", "voice": "alloy" }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Text to speech
per 1M characters- Official price
- $0.000015
- Official
- $0.000015
- Discount
- —
| Official priceper 1M characters | Officialper 1M characters | Discount | |
|---|---|---|---|
| Text to speech | $0.000015 | $0.000015 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Qwen3-TTS 1.7B CustomVoice in Console with a prompt ready to edit or send.
Help me try qwen3-tts-1-7b-customvoice with a short audio request at /v1/audio/speech. Show the result and cost.
Use cases
Voice assistants
Give an in-app assistant one steady named voice and low response delay so replies start playing quickly after the text is ready.
Audiobook and article narration
Read long text with a chosen speaker and add an instruction such as calm and warm to keep the delivery consistent across chapters.
Localised product videos
Use the Japanese, Korean, or English presets to voice the same script per market while the brand keeps a recognisable sound.
Game and character lines
Pick distinct presets for different characters and add emotion instructions for angry, tired, or excited lines.
Prompt examples
Speaker: Ryan. Instruction: calm, friendly, slightly slow. Text: Welcome back. Your order has shipped and should arrive on Thursday.
Speaker: Ono_Anna. Instruction: cheerful announcer. Text: 本日はご来店いただき、ありがとうございます。
Speaker: Vivian. Instruction: whisper, as if telling a secret. Text: 你听说了吗?明天会有一个大消息。
FAQ
Which voices does Qwen3-TTS CustomVoice offer?
Nine presets: Vivian, Serena, Uncle_Fu, Dylan, and Eric for Chinese variants, Ryan and Aiden for English, Ono_Anna for Japanese, and Sohee for Korean. You choose one by name for each request.
Can I control emotion and speaking style?
Yes. You can add a natural-language instruction, for example angry, gentle, or fast and excited, and the model adjusts tone and manner while keeping the selected speaker's timbre.
How is CustomVoice different from VoiceDesign?
CustomVoice uses nine fixed speakers that you select. VoiceDesign has no preset list; you describe the voice you want in words, such as age, mood, and prosody, and the model creates it. Choose CustomVoice for consistency and VoiceDesign for new voices.
Which languages does it speak?
Ten: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian. The model card advises using each speaker in its native language for the best quality.
Is it fast enough for live conversation?
Alibaba states end-to-end synthesis latency as low as 97 ms, aimed at real-time use. Actual delay in your app also depends on network distance and text length, so measure it with your own traffic.
How much does Qwen3-TTS 1.7B CustomVoice cost?
On TokenLab, Qwen3-TTS 1.7B CustomVoice costs - . The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should Qwen3-TTS 1.7B CustomVoice use?
Use https://api.tokenlab.sh/v1/audio/speech for Qwen3-TTS 1.7B CustomVoice. The request example below shows the matching code shape.
Which operations does Qwen3-TTS 1.7B CustomVoice support?
Qwen3-TTS 1.7B CustomVoice supports Text to speech. Select an operation above to see its endpoint and request example.
Compare Qwen3-TTS 1.7B CustomVoice
Sources
Reviewed Oct 2, 2026