Google: Gemini 3.6 Flash
gemini-3.6-flash- Input / Output
- $0.75 / $3.75
- Context
- 1M
- Released
- Jul 21, 2026
- Max output
- 64K
- Modalities
- VisionChat
- Capabilities
- Tool usePrompt Cache
About Gemini 3.6 Flash
Gemini 3.6 Flash is a stable Flash-tier Gemini model that Google describes as balancing speed and multimodal capability. It was released in July 2026, between 3.5 Flash and 3.7 Flash. Google frames it around efficient coding loops, tool orchestration, and visual reasoning, so it suits agents that inspect a result, revise, and decide the next action.
Where it works well
- Fits loops in which an agent runs code, reads the output, and revises.
- Tool orchestration across several calls is a stated focus.
- Visual reasoning over screenshots and diagrams works in the same request as text.
- Schema-bound JSON and caching support repeatable production use.
When to choose another model
- 3.7 and 3.8 Flash are newer and are pitched at more complex coding and longer autonomous runs.
- It has no dedicated thinking mode, so deep deliberation is better done on a Pro model.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textGemini APIPOST/v1beta/models/gemini-3.6-flash:generateContentAPI format:curl "https://api.tokenlab.sh/v1beta/models/gemini-3.6-flash:generateContent" \ -H "Content-Type: application/json" \ -H "x-goog-api-key: sk-xxx" \ -d '{ "contents": [ {"role": "user", "parts": [{"text": "Hello!"}]} ] }'
Pricing
TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.
Default spec
per 1M tokens- Official price
- Input $0.75 / Output $3.75 / Cache read $0.075
- TokenLab price
- Input $0.75 / Output $3.75 / Cache read $0.075
- Discount
- —
Cache read
- Official price
- $0.075
- TokenLab price
- $0.075
- Discount
- —
Charged per call, only when the model uses the tool.
Web search
- Official price
- $0.014/search
- TokenLab price
- $0.007/search
- Discount
- -50%
| Official priceper 1M tokens | TokenLab priceper 1M tokens | Discount | |
|---|---|---|---|
| Default spec | Input $0.75 / Output $3.75 / Cache read $0.075 | Input $0.75 / Output $3.75 / Cache read $0.075 | — |
| Prompt cache pricing | |||
| Cache read | $0.075 | $0.075 | — |
| Tool feesCharged per call, only when the model uses the tool. | |||
| Web search | $0.014/search | $0.007/search | -50% |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
- 30-day success rate
- 99.7%
- 30-day requests
- 100+
- Last active
- 17 hours ago
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Gemini 3.6 Flash in Console with a prompt ready to edit or send.
Help me try gemini-3.6-flash with a short message at /v1beta/models/gemini-3.6-flash:generateContent. Show the reply, latency, and cost.
Use cases
Best for- Vision
Code-fix loops
Let an agent run tests, read failures, patch files, and repeat until they pass.
UI inspection
Compare a screenshot against a written spec and list each spacing, color, or layout mismatch that a reviewer should fix.
Tool-heavy assistants
Coordinate search, calendar, and CRM calls inside one conversation, passing results from each tool into the next decision.
Multimodal Q&A
Answer questions that need both a chart image and the surrounding report text.
Prompt examples
Run the tests, read the failure output, and change only the function that breaks the second test.
Compare this screenshot with the design spec and list each spacing or color mismatch.
Check my calendar and the CRM, then draft a reply that proposes two meeting times.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What is Gemini 3.6 Flash best at?
Iterative agent work: coding loops, tool orchestration, and reasoning over images. It balances speed against multimodal ability, which makes it a reasonable default for agents that do not need the newest Flash.
How is it different from 3.5 Flash?
3.6 is the later release and is described around coding loops and tool use, while 3.5 Flash is the baseline for routine high-volume work. Expect better agent behavior from 3.6.
Should I move to 3.8 Flash?
If you run long software tasks or complex enterprise flows, 3.8 Flash is the one Google positions for them. For simpler agents 3.6 may be enough; run both on your own traces.
Does it accept images?
Yes. Images and text can be combined in a single request, and the model answers in text or in JSON that follows a schema you provide. It does not generate images itself.
How much does Gemini 3.6 Flash cost?
On TokenLab, Gemini 3.6 Flash costs Input $0.75 / Output $3.75 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of Gemini 3.6 Flash?
Gemini 3.6 Flash accepts up to 1,048,576 tokens of context and returns up to 65,536 tokens in one response.
Which endpoint should Gemini 3.6 Flash use?
Use https://api.tokenlab.sh/v1beta/models/gemini-3.6-flash:generateContent for Gemini 3.6 Flash. The request example below shows the matching code shape.
Which operations does Gemini 3.6 Flash support?
Gemini 3.6 Flash supports Text to text. Select an operation above to see its endpoint and request example.
Compare Gemini 3.6 Flash
Sources
Reviewed Oct 2, 2026