Alibaba: Sambert Zhiyuan v1
sambert-zhiyuan-v1- Input / Output
- $14.71 / $0.00
- Modalities
- Audio
About Sambert Zhiyuan v1
Sambert Zhiyuan v1 is a single named voice in Alibaba's Sambert speech synthesis family, intended for general-purpose Chinese narration. Zhiyuan is a calm, scholarly female voice that also reads English, and the model supports word-level timestamps and a default 16 kHz sample rate. It is an older, fixed-voice engine: it has no cloning, dialect list, or style control like CosyVoice v3.5 Plus or Qwen3-TTS-Flash.
Where it works well
- Delivers one consistent voice, so repeated content sounds like the same narrator.
- Supports timestamps, which helps sync subtitles or highlight text as it is read.
- Reads Chinese and English text.
- Fits calm, explanatory content such as articles and announcements.
When to choose another model
- It is one fixed voice with no cloning or design, so it cannot represent a custom brand voice.
- It is an older Sambert-generation model without instruction control for emotion or style.
- The default 16 kHz output is lower fidelity than newer synthesis models offer.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to speechTokenLab endpointPOST/v1/audio/speechcurl -X POST "https://api.tokenlab.sh/v1/audio/speech" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "operation": "tts", "response_format": "mp3", "speed": 1, "model": "sambert-zhiyuan-v1", "input": "Hello from TokenLab.", "voice": "alloy" }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Tts Text Number
per 1M characters- Official price
- Input $14.71 / Output $0.00
- Official
- Input $14.71 / Output $0.00
- Discount
- —
| Official priceper 1M characters | Officialper 1M characters | Discount | |
|---|---|---|---|
| Tts Text Number | Input $14.71 / Output $0.00 | Input $14.71 / Output $0.00 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Sambert Zhiyuan v1 in Console with a prompt ready to edit or send.
Help me try sambert-zhiyuan-v1 with a short audio request at /v1/audio/speech. Show the result and cost.
Use cases
Article reading
Read long-form Chinese articles with a steady voice that stays the same from page to page.
Subtitled explainers
Use the timestamps to sync on-screen text with narration in short educational videos.
Announcements
Generate repeated notices in a consistent tone for kiosks, apps, and public-address style messages.
Prompt examples
Read this notice in Mandarin in a calm, even pace.
Narrate the following paragraph, then give me word timestamps for subtitles.
Read this bilingual product intro, Chinese first and English second.
FAQ
What is Sambert Zhiyuan v1?
A fixed female voice named Zhiyuan from Alibaba's Sambert text-to-speech family. Alibaba describes it as a general-purpose Chinese voice with a calm, scholarly tone, and it also handles English text.
Does it support timestamps?
Yes. Alibaba lists timestamp output for this voice, which lets you align subtitles or highlight words while the audio plays. This is useful in reading and explainer apps.
Can I customize the voice?
No. Zhiyuan is a fixed voice. For cloning, voice design, or style instructions, use CosyVoice v3.5 Plus; for ready-made multilingual voices over HTTP, use Qwen3-TTS-Flash.
What sample rate does it output?
The default is 16 kHz, per Alibaba's documentation. That suits speech for apps and prompts, though newer synthesis models may offer richer audio for music-like or broadcast use.
When is Sambert still a good choice?
When you want a stable, predictable Chinese narrator and timestamps, and do not need dialects, cloning, or emotional control. For most new projects, Qwen3-TTS-Flash has broader language coverage.
How much does Sambert Zhiyuan v1 cost?
On TokenLab, Sambert Zhiyuan v1 costs Input $14.71 / Output $0.00 per 1M characters. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should Sambert Zhiyuan v1 use?
Use https://api.tokenlab.sh/v1/audio/speech for Sambert Zhiyuan v1. The request example below shows the matching code shape.
Which operations does Sambert Zhiyuan v1 support?
Sambert Zhiyuan v1 supports Text to speech. Select an operation above to see its endpoint and request example.
Compare Sambert Zhiyuan v1
Sources
Reviewed Oct 2, 2026