Google: Gemini Embedding 2
gemini-embedding-2- Input / Output
- $0.20 / $0.00
- Context
- 8K
- Released
- Sep 30, 2026
- Modalities
- Embedding
About Gemini Embedding 2
Gemini Embedding 2 is Google's newer embedding model in the Gemini API. It turns text into vectors that capture meaning, for search, clustering and recommendations, with dimensions adjustable from 128 to 3,072. It accepts longer text than gemini-embedding-001 and replaces that model's task type parameter with instructions written into the prompt.
Where it works well
- Text input can run several times longer than gemini-embedding-001 allows, so fewer chunks are needed.
- Truncated vectors are normalized automatically, which removes a step that the older model leaves to you.
- Task instructions go in the prompt, so one request can say what the vector is for in ordinary words.
- Dimensions from 128 to 3,072 let you trade storage and search speed against precision.
When to choose another model
- Vectors cannot be compared with gemini-embedding-001 output, so migrating an index means re-embedding everything.
- Vectors are returned as float values only, so a compact base64 encoding is not available.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
EmbeddingTokenLab endpointPOST/v1/embeddingscurl -X POST "https://api.tokenlab.sh/v1/embeddings" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "dimensions": 768, "encoding_format": "float", "input": "Semantic search", "model": "gemini-embedding-2" }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Pricing
per 1M tokens- Official
- Input$0.20per 1M tokensOutput$0.00per 1M tokens
- Official price
- Input$0.20per 1M tokensOutput$0.00per 1M tokens
| Official priceper 1M tokens | Officialper 1M tokens | Discount | |
|---|---|---|---|
| Input | $0.20 | $0.20 | — |
| Output | $0.00 | $0.00 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Gemini Embedding 2 in Console with a prompt ready to edit or send.
Help me try gemini-embedding-2 with embedding at /v1/embeddings. Show the result and cost.
Use cases
Longer passage retrieval
Embed larger text chunks to keep paragraphs and their context together in knowledge-base search.
Instruction-guided similarity
Add a short instruction such as finding code snippets or finding FAQs, and let the prompt steer what counts as similar.
Support ticket matching
Match a new ticket to earlier ones that describe the same problem in different words.
Topic clustering
Group reviews or articles by meaning to see themes without writing keyword lists.
Prompt examples
task: question answering | query: how do I rotate an API key without downtime
title: Refund policy | text: Customers may return unused items within 30 days of delivery for a full refund to the original payment method.
Embed this support ticket so it can be matched with similar past tickets: customer cannot reset password after changing email address.
Cost calculator
FAQ
What does Gemini Embedding 2 add over Gemini Embedding 001?
Longer text input, prompt-based instructions in place of the task type parameter, and automatic normalization of shortened vectors. Both models turn text into vectors, but 001 has a far shorter input cap.
Can I reuse embeddings made with gemini-embedding-001?
No. Google states the two vector spaces are incompatible, so you cannot compare vectors from one model against vectors from the other. When you move to Embedding 2, re-embed both your documents and your queries.
What embedding sizes are available?
From 128 to 3,072 dimensions, with 768, 1,536 and 3,072 recommended. Google says smaller sizes score close to larger ones. Embedding 2 normalizes truncated vectors for you, so you can store and search them directly.
How does task selection work?
There is no task type parameter. You put the intent into the prompt, for example marking text as a query or as a document to retrieve. Use the same convention on both sides of your index so queries and documents line up.
What does it return?
Float vectors, one per input text, at the dimension you choose. Store them in any vector index and compare them with cosine similarity or dot product.
How much does Gemini Embedding 2 cost?
On TokenLab, Gemini Embedding 2 costs Input $0.20 / Output $0.00 per 1M tokens. The pricing table above shows the full breakdown.
Which endpoint should Gemini Embedding 2 use?
Use https://api.tokenlab.sh/v1/embeddings for Gemini Embedding 2. The request example below shows the matching code shape.
Which operations does Gemini Embedding 2 support?
Gemini Embedding 2 supports Embedding. Select an operation above to see its endpoint and request example.
Compare Gemini Embedding 2
Sources
Reviewed Oct 2, 2026