Google: Gemini Embedding 001
gemini-embedding-001- Input / Output-50%
- $0.15 / $0.00$0.075 / $0.00
- Context
- 2K
- Modalities
- Embedding
About Gemini Embedding 001
Gemini Embedding 001 is Google's text-only embedding model in the Gemini API. It turns text into vectors that capture meaning, so you can match a question to passages, cluster related documents or power recommendations without relying on shared keywords. Output size is adjustable from 128 to 3,072 dimensions, and a task type parameter tells the model what the vectors will be used for.
Where it works well
- A task type parameter, such as retrieval document or semantic similarity, lets you tune vectors to the job instead of using one general embedding.
- Output dimensions can be reduced from 3,072 to as few as 128, and Google notes smaller sizes score close to the full size.
- Each input string gets its own embedding, which suits batch indexing where every document needs its own vector.
- Vectors come back as float values or base64, so they can be stored compactly.
When to choose another model
- Input is text only and capped well below the newer model, so long documents must be split into chunks before embedding.
- Vectors below 3,072 dimensions need to be normalized by you before similarity search.
- Its vector space is not compatible with gemini-embedding-2, so switching means re-embedding the whole collection.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
EmbeddingTokenLab endpointPOST/v1/embeddingscurl -X POST "https://api.tokenlab.sh/v1/embeddings" \ -H "Authorization: Bearer sk-xxx" \ -H "Content-Type: application/json" \ -d '{ "dimensions": 768, "encoding_format": "float", "input": "Semantic search", "model": "gemini-embedding-001" }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Pricing
per 1M tokens- TokenLab price
- Input$0.075per 1M tokensOutput$0.00per 1M tokens
- Official price
- Input$0.15per 1M tokensOutput$0.00per 1M tokens
Discount: -50%
| Official priceper 1M tokens | TokenLab priceper 1M tokens | Discount | |
|---|---|---|---|
| Input | $0.15 | $0.075 | -50% |
| Output | $0.00 | $0.00 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Gemini Embedding 001 in Console with a prompt ready to edit or send.
Help me try gemini-embedding-001 with embedding at /v1/embeddings. Show the result and cost.
Use cases
Retrieval for question answering
Embed document chunks once, embed each question with the query task type, and fetch the closest chunks to feed a chat model.
Duplicate and near-duplicate detection
Compare vectors of support tickets or product descriptions to find items that say the same thing in different words.
Topic clustering
Group feedback or articles by meaning to see themes without writing keyword lists.
Content recommendations
Suggest related articles by finding the nearest vectors to what a reader just finished.
Prompt examples
Embed these 500 help-center paragraphs for retrieval and store the vectors so users can search them in plain language.
Find the five tickets most similar in meaning to: customer cannot reset password after changing email address.
Cluster the last month of app reviews by embedding each review and grouping the nearest neighbors.
Cost calculator
FAQ
What is Gemini Embedding 001 used for?
Semantic search and retrieval, clustering, classification and recommendations over text. It converts text into a vector so similar meanings land close together. Use the task type parameter to say whether you are embedding documents, queries or comparing similarity.
How is Gemini Embedding 001 different from Gemini Embedding 2?
001 handles text only, with a much shorter input cap, and uses a task type parameter. Embedding 2 takes longer text and expects task instructions in the prompt instead. Their vectors cannot be compared with each other.
What dimension should I choose?
Google recommends 768, 1,536 or 3,072, and allows anything from 128 to 3,072. Lower sizes save storage and search time with scores close to the full size. Normalize the vector yourself when you use anything under 3,072.
How long can the input be?
Each input has a length cap. For longer documents, split them into passages that fit and embed each one. Chunk boundaries at paragraphs or headings usually give better retrieval than fixed-length cuts.
How do I call it?
Use the standard embeddings request format. Send the model name and the input text or a list of strings, then read the vectors from the response. Output can be requested as float values or base64.
How much does Gemini Embedding 001 cost?
On TokenLab, Gemini Embedding 001 costs Input $0.075 / Output $0.00 per 1M tokens. The pricing table above shows the full breakdown.
Which endpoint should Gemini Embedding 001 use?
Use https://api.tokenlab.sh/v1/embeddings for Gemini Embedding 001. The request example below shows the matching code shape.
Which operations does Gemini Embedding 001 support?
Gemini Embedding 001 supports Embedding. Select an operation above to see its endpoint and request example.
Compare Gemini Embedding 001
Sources
Reviewed Oct 2, 2026