Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Google: Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite targets high-volume work with short response times. Consider it for document processing, extraction, and smaller agent tasks that run repeatedly and need efficient handling of text and visual input.
Compare models
gemini-3.5-flash-lite
AvailableGoogleChatCached input 90% less
Input / Output-50%
$0.30 / $2.50$0.15 / $1.25
Context
1M
Released
Jul 21, 2026
Max output
64K
Modalities
VisionChat
Capabilities
Tool usePrompt Cache

About Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite is Google's fastest and most cost-effective model of the 3.5 generation, released in July 2026. It targets high-volume work where short response times matter: document processing, extraction, and small repeated agent tasks. It reads text and images, calls tools, returns schema-bound JSON, and caches context, and Google now points new projects at it.

Where it works well

  • Designed for short response times on work that repeats thousands of times.
  • Handles text and images together for document and screenshot processing.
  • Schema-bound JSON keeps extraction output consistent.
  • Function calling lets it run small agent steps such as lookups or routing.

When to choose another model

  • Long coding sessions and complex enterprise workflows belong to 3.8 Flash or 3.1 Pro.
  • It will miss details that a fuller Flash model catches on dense or ambiguous documents.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textGemini API
    POST/v1beta/models/gemini-3.5-flash-lite:generateContent
    API format:
    curl "https://api.tokenlab.sh/v1beta/models/gemini-3.5-flash-lite:generateContent" \
      -H "Content-Type: application/json" \
      -H "x-goog-api-key: sk-xxx" \
      -d '{
        "contents": [
          {"role": "user", "parts": [{"text": "Hello!"}]}
        ]
      }'

Pricing

TokenLab price applies to Verified, which costs less on most models. Official price is the model maker's published price and applies to the more reliable Official route. Auto charges the route that completes the request.

Default spec

per 1M tokens
Official price
Input $0.30 / Output $2.50 / Cache read $0.03
TokenLab price
Input $0.15 / Output $1.25 / Cache read $0.015
Discount
-50%
Prompt cache pricing

Cache read

Official price
$0.03
TokenLab price
$0.015
Discount
-50%

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
30-day success rate
99.7%
7-day median latency
1.9 sn=346
7-day P95 latency
7 sn=346
30-day requests
100+
Last active
17 hours ago

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Gemini 3.5 Flash-Lite in Console with a prompt ready to edit or send.

Help me try gemini-3.5-flash-lite with a short message at /v1beta/models/gemini-3.5-flash-lite:generateContent. Show the reply, latency, and cost.

Use cases

Best for
  • Vision
  • Document pipelines

    Convert contracts, invoices, and reports into structured fields at scale, writing one validated JSON record for every incoming file.

  • Queue triage

    Sort incoming tickets or messages by topic and urgency before a human or larger model sees them.

  • Small agents

    Run narrow tools-in-a-loop jobs such as order lookups or calendar changes, where each step is small and the loop repeats often.

  • Visual checks

    Confirm that uploaded photos meet simple rules, like a visible label or correct orientation.

Prompt examples

Extract every line item from this invoice as JSON with description, quantity, and unit amount.

Rate the urgency of this ticket as low, medium, or high and give the single sentence that justifies it.

Check whether this product photo shows the packaging label clearly; answer yes or no and say why.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

What is Gemini 3.5 Flash-Lite for?

Fast, repeated, simple work: extraction, classification, routing, and light multimodal checks. It is the economical end of the 3.5 family, built for response time over depth.

Is it newer than Gemini 3.1 Flash-Lite?

Yes. It is the later generation, and Google lists it as the recommended lightweight choice for new projects. The older 3.1 Flash-Lite remains available for existing integrations.

Can it read images?

Yes, images can accompany text in a request, and the answer comes back as plain text or as structured JSON, which suits photo checks and scanned-document extraction.

When is it too small?

When the task needs sustained reasoning, careful long-document analysis, or multi-hour coding work. Step up to Gemini 3.8 Flash or 3.1 Pro in those cases.

How much does Gemini 3.5 Flash-Lite cost?

On TokenLab, Gemini 3.5 Flash-Lite costs Input $0.15 / Output $1.25 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite accepts up to 1,048,576 tokens of context and returns up to 65,536 tokens in one response.

Which endpoint should Gemini 3.5 Flash-Lite use?

Use https://api.tokenlab.sh/v1beta/models/gemini-3.5-flash-lite:generateContent for Gemini 3.5 Flash-Lite. The request example below shows the matching code shape.

Which operations does Gemini 3.5 Flash-Lite support?

Gemini 3.5 Flash-Lite supports Text to text. Select an operation above to see its endpoint and request example.

Compare Gemini 3.5 Flash-Lite

Sources

Reviewed Oct 2, 2026

More from Gemini 3

Related models