Alibaba: Text Embedding v4
text-embedding-v4- Input / Output
- $0.07 / —
- Context
- 8K
- Modalities
- Embedding
About Text Embedding v4
Text Embedding v4 is Alibaba's established multilingual text embedding model, with over 100 languages and output sizes from 64 to 2048 dimensions. It takes 10 items per batch, and can return sparse vectors through the DashScope API. It sits below Qwen3.7 Text Embedding in input length but offers smaller vectors.
Where it works well
- Offers small output sizes down to 64 and 128 dimensions for tight index budgets.
- Sparse vectors via the DashScope API enable hybrid keyword-plus-semantic search.
- Covers Chinese, English, Spanish, French, and over 100 other languages.
- Supports the instruct parameter to describe the retrieval task.
When to choose another model
- Inputs per item are short, so long documents must be chunked before embedding.
- Batches hold 10 items, fewer than the 20 allowed by Qwen3.7 Text Embedding.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text EmbeddingTokenLab endpointPOST/v1/embeddingscurl -X POST "https://api.tokenlab.sh/v1/embeddings" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "model": "text-embedding-v4", "input": "The quick brown fox jumps over the lazy dog." }'
Pricing
The Verified price applies to Verified, which costs less on most models. The Official price is the model maker's published price and applies to the more reliable Official route. Auto bills the route that completes the request.
Default
per 1M tokens- Official price
- Input $0.07 / Text Input $0.07
- Verified price
- —
- Discount
- —
| Official priceper 1M tokens | Verified priceper 1M tokens | Discount | |
|---|---|---|---|
| Default | Input $0.07 / Text Input $0.07 | — | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Text Embedding v4 in Console with a prompt ready to edit or send.
Help me try text-embedding-v4 with text-embedding at /v1/embeddings. Show the result and cost.
Use cases
Semantic site search
Embed pages in chunks and match user questions to the closest passages.
Clustering feedback
Group thousands of reviews or support tickets by topic using compact vectors and a clustering step.
Hybrid retrieval
Combine dense and sparse output so exact terms like part numbers still match.
Prompt examples
Embed these 2,000 support tickets at 256 dimensions and group them by topic.
Embed this FAQ chunk with sparse output for hybrid search.
Embed the Spanish query 'cómo cancelo mi suscripción' to find matching help articles.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What dimensions does text-embedding-v4 support?
2048, 1536, 1024 (default), 768, 512, 256, 128, and 64. Lower sizes reduce storage and speed up search, at some cost in precision. Start at the default and step down while testing recall.
How long can an input be?
Alibaba documents a per-item limit well below Qwen3.7 and 10 items per batch. Split longer documents into passages before embedding them. Check Alibaba's page for the current exact limit.
Does it return sparse vectors?
Yes, through the DashScope API, which returns sparse vectors alongside dense ones so you can run hybrid keyword-plus-semantic search over the same index. Sparse output is only available on that API.
Should I move to Qwen3.7?
Only if you need much longer inputs or larger vectors. If compact 64 or 128-dimension vectors suit you, v4 remains the option that offers them.
How much does Text Embedding v4 cost?
Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
Which endpoint should Text Embedding v4 use?
Use https://api.tokenlab.sh/v1/embeddings for Text Embedding v4. The request example below shows the matching code shape.
Which operations does Text Embedding v4 support?
Text Embedding v4 supports Text Embedding. Select an operation above to see its endpoint and request example.
Compare Text Embedding v4
Sources
Reviewed Oct 2, 2026