Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Alibaba: CosyVoice v3.5 Plus

CosyVoice v3.5 Plus turns written scripts into expressive speech. Use it for narrated articles, conversational voiceovers, and spoken product guidance, with voice and speaking speed chosen to fit the audience.
Compare models
cosyvoice-v3.5-plus
AvailableAlibabaText to speechSync
Input / Output
$22.06 / $0.00
Modalities
Audio

About CosyVoice v3.5 Plus

CosyVoice v3.5 Plus is Alibaba's speech synthesis model for voices you create yourself. It has no stock voice list: you clone a voice from a sample or design one from a text description, then synthesize with it over a WebSocket connection. It speaks Mandarin, English, and several other languages, plus many Chinese regional dialects, and it takes natural-language instructions to steer delivery, which Qwen3-TTS-Flash does not.

Where it works well

  • Works with cloned voices from a recording or voices designed from a written description.
  • Accepts instructions to control tone and speaking style.
  • Speaks Mandarin, Cantonese, and many Chinese regional dialects as well as English, French, German, Japanese, Korean, Russian, Portuguese, Thai, Indonesian, and Vietnamese.
  • Streams audio over WebSocket, so playback can start before synthesis finishes.

When to choose another model

  • It has no built-in system voices; you must create a cloned or designed voice first.
  • It requires a WebSocket client rather than a simple HTTP request.
  • Voice cloning needs a clean reference sample and consent to use that voice.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to speechTokenLab endpoint
    POST/v1/audio/speech
    curl -X POST "https://api.tokenlab.sh/v1/audio/speech" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "operation": "tts",
      "response_format": "mp3",
      "speed": 1,
      "model": "cosyvoice-v3.5-plus",
      "input": "Hello from TokenLab.",
      "voice": "alloy"
    }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Cosy Tts Number

per 1M characters
Official price
Input $22.06 / Output $0.00
Official
Input $22.06 / Output $0.00
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open CosyVoice v3.5 Plus in Console with a prompt ready to edit or send.

Help me try cosyvoice-v3.5-plus with a short audio request at /v1/audio/speech. Show the result and cost.

Use cases

  • Branded voice

    Clone an approved speaker once and use the voice for product tours, IVR prompts, and announcements.

  • Character voices

    Design a voice from a description, such as a warm older storyteller, for audiobooks or games.

  • Regional-dialect narration

    Produce spoken content in Sichuan, Shanghai, or Cantonese speech for local audiences.

Prompt examples

Read this announcement in a brisk, upbeat tone, with a short pause after the date.

Speak the following paragraph slowly and warmly, as if reading a bedtime story.

Say this in Cantonese with a relaxed, conversational pace.

FAQ

Does CosyVoice v3.5 Plus have preset voices?

No. Alibaba lists no system voices for this model. You need a cloned voice from a sample or a designed voice from a text description, then pass that voice when synthesizing.

What languages and dialects does it support?

Mandarin, Cantonese, and many Chinese regional dialects, plus English, French, German, Japanese, Korean, Russian, Portuguese, Thai, Indonesian, and Vietnamese. Dialect support is a main difference from the Qwen3-TTS-Flash models.

Can I control how it sounds?

Yes. The model supports instruction control, so you can describe tone, emotion, or pacing in natural language. Voice choice sets who speaks; the instruction shapes how the line is delivered.

How is it called?

Through a WebSocket API that streams audio back as it is generated. This suits long scripts and live playback, but you will need a client that handles the connection.

When should I choose Qwen3-TTS-Flash instead?

When you want ready-made voices and a simple HTTP call and do not need cloning or instructions. Choose CosyVoice when voice identity or delivery style is part of your product.

How much does CosyVoice v3.5 Plus cost?

On TokenLab, CosyVoice v3.5 Plus costs Input $22.06 / Output $0.00 per 1M characters. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should CosyVoice v3.5 Plus use?

Use https://api.tokenlab.sh/v1/audio/speech for CosyVoice v3.5 Plus. The request example below shows the matching code shape.

Which operations does CosyVoice v3.5 Plus support?

CosyVoice v3.5 Plus supports Text to speech. Select an operation above to see its endpoint and request example.

Compare CosyVoice v3.5 Plus

Sources

Reviewed Oct 2, 2026

Related models