Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Google: Gemini 3.7 Flash

High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning
Compare models
gemini-3.7-flash
AvailableGoogleChatCached input 90% less
Input / Output-50%
$0.75 / $3.75$0.375 / $1.88
Context
1M
Released
Aug 13, 2026
Max output
64K
Modalities
VisionChat
Capabilities
Tool usePrompt Cache

About Gemini 3.7 Flash

Gemini 3.7 Flash is a stable Flash-tier model from Google, released in August 2026, that Google describes as built for complex coding, agentic workflows, and reliable multi-step execution. It reads text and images, supports thinking, calls tools, returns schema-bound JSON, and caches context. It is a step above 3.6 Flash and one step below 3.8 Flash.

Where it works well

  • Reasoning can run before the answer, which helps on coding and multi-step tasks.
  • Reliable multi-step execution is the stated design goal, useful when an agent must not lose its place.
  • Tool calling and schema-bound JSON make results easy to feed into other systems.
  • Image input handles screenshots, diagrams, and scanned documents.

When to choose another model

  • Gemini 3.8 Flash is the newer model that Google presents as its most capable Flash tier.
  • Thinking adds latency, so simple lookups are better served by Flash-Lite.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textGemini API
    POST/v1beta/models/gemini-3.7-flash:generateContent
    API format:
    curl "https://api.tokenlab.sh/v1beta/models/gemini-3.7-flash:generateContent" \
      -H "Content-Type: application/json" \
      -H "x-goog-api-key: sk-xxx" \
      -d '{
        "contents": [
          {"role": "user", "parts": [{"text": "Hello!"}]}
        ]
      }'

Pricing

The Verified price applies to Verified, which costs less on most models. The Official price is the model maker's published price and applies to the more reliable Official route. Auto bills the route that completes the request.

Promotional Standard

per 1M tokens
Official price
Input $0.75 / Output $3.75 / Cache read $0.075
Verified price
Input $0.375 / Output $1.88 / Cache read $0.0375
Discount
-50%
Prompt cache pricing

Cache read

Official price
$0.075
Verified price
$0.0375
Discount
-50%
Tool fees

Charged per call, only when the model uses the tool.

Web search

Official price
$0.014/search
Verified price
$0.007/search
Discount
-50%

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
30-day success rate
99.8%
7-day median latency
13.2 sn=271
7-day P95 latency
53 sn=271
30-day requests
1K+
Last active
1 hour ago

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Gemini 3.7 Flash in Console with a prompt ready to edit or send.

Help me try gemini-3.7-flash with a short message at /v1beta/models/gemini-3.7-flash:generateContent. Show the reply, latency, and cost.

Use cases

Best for
  • Vision
  • Feature implementation

    Have an agent read a ticket, edit several files, and run tests.

  • Workflow automation

    Chain approvals, lookups, and notifications where each step depends on the last.

  • Code review

    Read a diff, flag risky changes, and explain each finding so that a human reviewer can decide quickly what to accept.

  • Data cleanup

    Reason through inconsistent records and propose normalization rules before applying them, so the cleanup can be audited first.

Prompt examples

Read this ticket and the linked files, implement the change, and list which tests you ran.

Review this diff for concurrency bugs and rank the findings by severity.

These two customer records look like duplicates. Explain which fields conflict and what rule would merge them safely.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

What is Gemini 3.7 Flash designed for?

Complex coding, agent workflows, and multi-step execution that has to stay reliable from first step to last. It is a thinking-capable Flash model with tool and image support.

Does it support thinking?

Yes. It can reason before replying, which improves results on code and planning. Leave that effort low for short interactive turns to keep responses quick.

How does it differ from 3.6 Flash?

3.7 Flash is the later model and adds reasoning, with a stronger emphasis on dependable multi-step runs. 3.6 Flash is framed around coding loops and visual reasoning without thinking.

Is 3.8 Flash a better choice?

Often, for long-horizon software engineering. 3.7 Flash remains a sound choice when you have already tuned prompts for it and its results meet your bar.

How much does Gemini 3.7 Flash cost?

On TokenLab, Gemini 3.7 Flash costs Input $0.375 / Output $1.88 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of Gemini 3.7 Flash?

Gemini 3.7 Flash accepts up to 1,048,576 tokens of context and returns up to 65,536 tokens in one response.

Which endpoint should Gemini 3.7 Flash use?

Use https://api.tokenlab.sh/v1beta/models/gemini-3.7-flash:generateContent for Gemini 3.7 Flash. The request example below shows the matching code shape.

Which operations does Gemini 3.7 Flash support?

Gemini 3.7 Flash supports Text to text. Select an operation above to see its endpoint and request example.

Compare Gemini 3.7 Flash

Sources

Reviewed Oct 2, 2026

More from Gemini 3

Related models