Alibaba: CosyVoice v3.5 Plus
cosyvoice-v3.5-plus- Input / Output
- $22.06 / $0.00
- Modalities
- Audio
About CosyVoice v3.5 Plus
CosyVoice v3.5 Plus is Alibaba's speech synthesis model for voices you create yourself. It has no stock voice list: you clone a voice from a sample or design one from a text description, then synthesize with it over a WebSocket connection. It speaks Mandarin, English, and several other languages, plus many Chinese regional dialects, and it takes natural-language instructions to steer delivery, which Qwen3-TTS-Flash does not.
Where it works well
- Works with cloned voices from a recording or voices designed from a written description.
- Accepts instructions to control tone and speaking style.
- Speaks Mandarin, Cantonese, and many Chinese regional dialects as well as English, French, German, Japanese, Korean, Russian, Portuguese, Thai, Indonesian, and Vietnamese.
- Streams audio over WebSocket, so playback can start before synthesis finishes.
When to choose another model
- It has no built-in system voices; you must create a cloned or designed voice first.
- It requires a WebSocket client rather than a simple HTTP request.
- Voice cloning needs a clean reference sample and consent to use that voice.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to speechTokenLab endpointPOST/v1/audio/speechcurl -X POST "https://api.tokenlab.sh/v1/audio/speech" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "operation": "tts", "response_format": "mp3", "speed": 1, "model": "cosyvoice-v3.5-plus", "input": "Hello from TokenLab.", "voice": "alloy" }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Cosy Tts Number
per 1M characters- Official price
- Input $22.06 / Output $0.00
- Official
- Input $22.06 / Output $0.00
- Discount
- —
| Official priceper 1M characters | Officialper 1M characters | Discount | |
|---|---|---|---|
| Cosy Tts Number | Input $22.06 / Output $0.00 | Input $22.06 / Output $0.00 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open CosyVoice v3.5 Plus in Console with a prompt ready to edit or send.
Help me try cosyvoice-v3.5-plus with a short audio request at /v1/audio/speech. Show the result and cost.
Use cases
Branded voice
Clone an approved speaker once and use the voice for product tours, IVR prompts, and announcements.
Character voices
Design a voice from a description, such as a warm older storyteller, for audiobooks or games.
Regional-dialect narration
Produce spoken content in Sichuan, Shanghai, or Cantonese speech for local audiences.
Prompt examples
Read this announcement in a brisk, upbeat tone, with a short pause after the date.
Speak the following paragraph slowly and warmly, as if reading a bedtime story.
Say this in Cantonese with a relaxed, conversational pace.
FAQ
Does CosyVoice v3.5 Plus have preset voices?
No. Alibaba lists no system voices for this model. You need a cloned voice from a sample or a designed voice from a text description, then pass that voice when synthesizing.
What languages and dialects does it support?
Mandarin, Cantonese, and many Chinese regional dialects, plus English, French, German, Japanese, Korean, Russian, Portuguese, Thai, Indonesian, and Vietnamese. Dialect support is a main difference from the Qwen3-TTS-Flash models.
Can I control how it sounds?
Yes. The model supports instruction control, so you can describe tone, emotion, or pacing in natural language. Voice choice sets who speaks; the instruction shapes how the line is delivered.
How is it called?
Through a WebSocket API that streams audio back as it is generated. This suits long scripts and live playback, but you will need a client that handles the connection.
When should I choose Qwen3-TTS-Flash instead?
When you want ready-made voices and a simple HTTP call and do not need cloning or instructions. Choose CosyVoice when voice identity or delivery style is part of your product.
How much does CosyVoice v3.5 Plus cost?
On TokenLab, CosyVoice v3.5 Plus costs Input $22.06 / Output $0.00 per 1M characters. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should CosyVoice v3.5 Plus use?
Use https://api.tokenlab.sh/v1/audio/speech for CosyVoice v3.5 Plus. The request example below shows the matching code shape.
Which operations does CosyVoice v3.5 Plus support?
CosyVoice v3.5 Plus supports Text to speech. Select an operation above to see its endpoint and request example.
Compare CosyVoice v3.5 Plus
Sources
Reviewed Oct 2, 2026