Google: Gemini 3.5 Flash-Lite
gemini-3.5-flash-lite- Input / Output-50%
- $0.30 / $2.50$0.15 / $1.25
- Context
- 1M
- Released
- Jul 21, 2026
- Max output
- 64K
- Modalities
- VisionChat
- Capabilities
- Tool usePrompt Cache
About Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is Google's fastest and most cost-effective model of the 3.5 generation, released in July 2026. It targets high-volume work where short response times matter: document processing, extraction, and small repeated agent tasks. It reads text and images, calls tools, returns schema-bound JSON, and caches context, and Google now points new projects at it.
Where it works well
- Designed for short response times on work that repeats thousands of times.
- Handles text and images together for document and screenshot processing.
- Schema-bound JSON keeps extraction output consistent.
- Function calling lets it run small agent steps such as lookups or routing.
When to choose another model
- Long coding sessions and complex enterprise workflows belong to 3.8 Flash or 3.1 Pro.
- It will miss details that a fuller Flash model catches on dense or ambiguous documents.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textGemini APIPOST/v1beta/models/gemini-3.5-flash-lite:generateContentAPI format:curl "https://api.tokenlab.sh/v1beta/models/gemini-3.5-flash-lite:generateContent" \ -H "Content-Type: application/json" \ -H "x-goog-api-key: sk-xxx" \ -d '{ "contents": [ {"role": "user", "parts": [{"text": "Hello!"}]} ] }'
Pricing
TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.
Default spec
per 1M tokens- Official price
- Input $0.30 / Output $2.50 / Cache read $0.03
- TokenLab price
- Input $0.15 / Output $1.25 / Cache read $0.015
- Discount
- -50%
Cache read
- Official price
- $0.03
- TokenLab price
- $0.015
- Discount
- -50%
| Official priceper 1M tokens | TokenLab priceper 1M tokens | Discount | |
|---|---|---|---|
| Default spec | Input $0.30 / Output $2.50 / Cache read $0.03 | Input $0.15 / Output $1.25 / Cache read $0.015 | -50% |
| Prompt cache pricing | |||
| Cache read | $0.03 | $0.015 | -50% |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
- 30-day success rate
- 99.7%
- 7-day median latency
- 1.9 sn=346
- 7-day P95 latency
- 7 sn=346
- 30-day requests
- 100+
- Last active
- 17 hours ago
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Gemini 3.5 Flash-Lite in Console with a prompt ready to edit or send.
Help me try gemini-3.5-flash-lite with a short message at /v1beta/models/gemini-3.5-flash-lite:generateContent. Show the reply, latency, and cost.
Use cases
Best for- Vision
Document pipelines
Convert contracts, invoices, and reports into structured fields at scale, writing one validated JSON record for every incoming file.
Queue triage
Sort incoming tickets or messages by topic and urgency before a human or larger model sees them.
Small agents
Run narrow tools-in-a-loop jobs such as order lookups or calendar changes, where each step is small and the loop repeats often.
Visual checks
Confirm that uploaded photos meet simple rules, like a visible label or correct orientation.
Prompt examples
Extract every line item from this invoice as JSON with description, quantity, and unit amount.
Rate the urgency of this ticket as low, medium, or high and give the single sentence that justifies it.
Check whether this product photo shows the packaging label clearly; answer yes or no and say why.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What is Gemini 3.5 Flash-Lite for?
Fast, repeated, simple work: extraction, classification, routing, and light multimodal checks. It is the economical end of the 3.5 family, built for response time over depth.
Is it newer than Gemini 3.1 Flash-Lite?
Yes. It is the later generation, and Google lists it as the recommended lightweight choice for new projects. The older 3.1 Flash-Lite remains available for existing integrations.
Can it read images?
Yes, images can accompany text in a request, and the answer comes back as plain text or as structured JSON, which suits photo checks and scanned-document extraction.
When is it too small?
When the task needs sustained reasoning, careful long-document analysis, or multi-hour coding work. Step up to Gemini 3.8 Flash or 3.1 Pro in those cases.
How much does Gemini 3.5 Flash-Lite cost?
On TokenLab, Gemini 3.5 Flash-Lite costs Input $0.15 / Output $1.25 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of Gemini 3.5 Flash-Lite?
Gemini 3.5 Flash-Lite accepts up to 1,048,576 tokens of context and returns up to 65,536 tokens in one response.
Which endpoint should Gemini 3.5 Flash-Lite use?
Use https://api.tokenlab.sh/v1beta/models/gemini-3.5-flash-lite:generateContent for Gemini 3.5 Flash-Lite. The request example below shows the matching code shape.
Which operations does Gemini 3.5 Flash-Lite support?
Gemini 3.5 Flash-Lite supports Text to text. Select an operation above to see its endpoint and request example.
Compare Gemini 3.5 Flash-Lite
Sources
Reviewed Oct 2, 2026