Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Alibaba: Qwen3-TTS 1.7B CustomVoice

Qwen3-TTS 1.7B CustomVoice is a text-to-speech model from Alibaba that offers nine premium preset timbres across various combinations of gender, age, language, and dialect. It provides precise style control over target voices through user instructions, supports voice cloning from a 3-second sample, and generates speech in 10+ languages with latency as low as 97ms.
Compare models
qwen3-tts-1-7b-customvoice
AvailableAlibabaText to speechSync
Price
From $0.000015
Modalities
Audio

About Qwen3-TTS 1.7B CustomVoice

Qwen3-TTS 1.7B CustomVoice is Alibaba's text-to-speech model built around a fixed set of nine preset voices. You pick a speaker by name and can add a plain-language instruction about tone, emotion, or speaking style. It speaks ten languages, including Chinese, English, Japanese, Korean, and several European languages, and Alibaba quotes end-to-end synthesis latency as low as 97 ms for interactive use.

Where it works well

  • Nine named premium speakers cover different genders, ages, languages, and dialects, so a voice can stay the same across a whole product.
  • A written instruction can steer emotion and manner of speaking without changing the chosen speaker.
  • Ten languages are supported: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian.
  • Alibaba reports end-to-end latency as low as 97 ms, which suits voice assistants and live narration.

When to choose another model

  • The voice list is fixed; to invent a new voice from a description, the VoiceDesign variant is the better fit.
  • Results are best when a speaker reads its own native language, so an English preset reading Japanese may sound accented.
  • Noisy or unusual input text, such as heavy markup or mixed symbols, can reduce output quality.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to speechTokenLab endpoint
    POST/v1/audio/speech
    curl -X POST "https://api.tokenlab.sh/v1/audio/speech" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "operation": "tts",
      "model": "qwen3-tts-1-7b-customvoice",
      "input": "Hello from TokenLab.",
      "voice": "alloy"
    }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Text to speech

per 1M characters
Official price
$0.000015
Official
$0.000015
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Qwen3-TTS 1.7B CustomVoice in Console with a prompt ready to edit or send.

Help me try qwen3-tts-1-7b-customvoice with a short audio request at /v1/audio/speech. Show the result and cost.

Use cases

  • Voice assistants

    Give an in-app assistant one steady named voice and low response delay so replies start playing quickly after the text is ready.

  • Audiobook and article narration

    Read long text with a chosen speaker and add an instruction such as calm and warm to keep the delivery consistent across chapters.

  • Localised product videos

    Use the Japanese, Korean, or English presets to voice the same script per market while the brand keeps a recognisable sound.

  • Game and character lines

    Pick distinct presets for different characters and add emotion instructions for angry, tired, or excited lines.

Prompt examples

Speaker: Ryan. Instruction: calm, friendly, slightly slow. Text: Welcome back. Your order has shipped and should arrive on Thursday.

Speaker: Ono_Anna. Instruction: cheerful announcer. Text: 本日はご来店いただき、ありがとうございます。

Speaker: Vivian. Instruction: whisper, as if telling a secret. Text: 你听说了吗?明天会有一个大消息。

FAQ

Which voices does Qwen3-TTS CustomVoice offer?

Nine presets: Vivian, Serena, Uncle_Fu, Dylan, and Eric for Chinese variants, Ryan and Aiden for English, Ono_Anna for Japanese, and Sohee for Korean. You choose one by name for each request.

Can I control emotion and speaking style?

Yes. You can add a natural-language instruction, for example angry, gentle, or fast and excited, and the model adjusts tone and manner while keeping the selected speaker's timbre.

How is CustomVoice different from VoiceDesign?

CustomVoice uses nine fixed speakers that you select. VoiceDesign has no preset list; you describe the voice you want in words, such as age, mood, and prosody, and the model creates it. Choose CustomVoice for consistency and VoiceDesign for new voices.

Which languages does it speak?

Ten: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian. The model card advises using each speaker in its native language for the best quality.

Is it fast enough for live conversation?

Alibaba states end-to-end synthesis latency as low as 97 ms, aimed at real-time use. Actual delay in your app also depends on network distance and text length, so measure it with your own traffic.

How much does Qwen3-TTS 1.7B CustomVoice cost?

On TokenLab, Qwen3-TTS 1.7B CustomVoice costs - . The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Qwen3-TTS 1.7B CustomVoice use?

Use https://api.tokenlab.sh/v1/audio/speech for Qwen3-TTS 1.7B CustomVoice. The request example below shows the matching code shape.

Which operations does Qwen3-TTS 1.7B CustomVoice support?

Qwen3-TTS 1.7B CustomVoice supports Text to speech. Select an operation above to see its endpoint and request example.

Compare Qwen3-TTS 1.7B CustomVoice

Sources

Reviewed Oct 2, 2026

More from Qwen TTS

Related models