Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Alibaba: Text Embedding v4

Text Embedding v4 converts multilingual passages into vectors with adjustable dimensions. It is useful for search and clustering when an application wants to tune the balance between semantic representation and index footprint.
Compare models
text-embedding-v4
AvailableAlibabaEmbeddingSync
Input / Output
$0.07 / —
Context
8K
Modalities
Embedding

About Text Embedding v4

Text Embedding v4 is Alibaba's established multilingual text embedding model, with over 100 languages and output sizes from 64 to 2048 dimensions. It takes 10 items per batch, and can return sparse vectors through the DashScope API. It sits below Qwen3.7 Text Embedding in input length but offers smaller vectors.

Where it works well

  • Offers small output sizes down to 64 and 128 dimensions for tight index budgets.
  • Sparse vectors via the DashScope API enable hybrid keyword-plus-semantic search.
  • Covers Chinese, English, Spanish, French, and over 100 other languages.
  • Supports the instruct parameter to describe the retrieval task.

When to choose another model

  • Inputs per item are short, so long documents must be chunked before embedding.
  • Batches hold 10 items, fewer than the 20 allowed by Qwen3.7 Text Embedding.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text EmbeddingTokenLab endpoint
    POST/v1/embeddings
    curl -X POST "https://api.tokenlab.sh/v1/embeddings" \
      -H "Authorization: Bearer sk-xxx" \
      -H "Content-Type: application/json" \
      -d '{
      "model": "text-embedding-v4",
      "input": "The quick brown fox jumps over the lazy dog."
    }'

Pricing

The Verified price applies to Verified, which costs less on most models. The Official price is the model maker's published price and applies to the more reliable Official route. Auto bills the route that completes the request.

Default

per 1M tokens
Official price
Input $0.07 / Text Input $0.07
Verified price
—
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Text Embedding v4 in Console with a prompt ready to edit or send.

Help me try text-embedding-v4 with text-embedding at /v1/embeddings. Show the result and cost.

Use cases

  • Semantic site search

    Embed pages in chunks and match user questions to the closest passages.

  • Clustering feedback

    Group thousands of reviews or support tickets by topic using compact vectors and a clustering step.

  • Hybrid retrieval

    Combine dense and sparse output so exact terms like part numbers still match.

Prompt examples

Embed these 2,000 support tickets at 256 dimensions and group them by topic.

Embed this FAQ chunk with sparse output for hybrid search.

Embed the Spanish query 'cómo cancelo mi suscripción' to find matching help articles.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

What dimensions does text-embedding-v4 support?

2048, 1536, 1024 (default), 768, 512, 256, 128, and 64. Lower sizes reduce storage and speed up search, at some cost in precision. Start at the default and step down while testing recall.

How long can an input be?

Alibaba documents a per-item limit well below Qwen3.7 and 10 items per batch. Split longer documents into passages before embedding them. Check Alibaba's page for the current exact limit.

Does it return sparse vectors?

Yes, through the DashScope API, which returns sparse vectors alongside dense ones so you can run hybrid keyword-plus-semantic search over the same index. Sparse output is only available on that API.

Should I move to Qwen3.7?

Only if you need much longer inputs or larger vectors. If compact 64 or 128-dimension vectors suit you, v4 remains the option that offers them.

How much does Text Embedding v4 cost?

Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Text Embedding v4 use?

Use https://api.tokenlab.sh/v1/embeddings for Text Embedding v4. The request example below shows the matching code shape.

Which operations does Text Embedding v4 support?

Text Embedding v4 supports Text Embedding. Select an operation above to see its endpoint and request example.

Compare Text Embedding v4

Sources

Reviewed Oct 2, 2026

More from Qwen Embedding

Related models