Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

Google: Gemma 4 31B

Gemma 4 31B is Google's dense open-weight model for text conversation, coding and multilingual tasks. This TokenLab model exposes the verified text Chat Completions contract.
Compare models
gemma-4-31b
AvailableGoogleChat
Input / Output
$0.102 / $0.297
Modalities
Chat

About Gemma 4 31B

Gemma 4 31B is the largest dense model in Google's open-weight Gemma 4 family, which launched in March 2026 in sizes from E2B to 31B. Google describes the family as built for reasoning, agentic workflows, coding, and multimodal understanding. Unlike the proprietary Gemini models, its weights are published, so the same model can also be run on your own hardware.

Where it works well

  • Open weights let you move the same model between hosted and self-run setups.
  • Dense 31B size targets reasoning and coding quality that smaller Gemma 4 sizes give up.
  • Google's card reports strong results on reasoning and coding benchmarks such as MMLU Pro and LiveCodeBench v6.
  • Multilingual coverage across more than 140 languages.

When to choose another model

  • It works with text only and has no tool calling, so agent workflows are better served by a Gemini model.
  • It is less capable than Gemini 3.x Pro and Flash models on long autonomous tasks.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textChat Completions API
    POST/v1/chat/completions
    curl https://api.tokenlab.sh/v1/chat/completions \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer sk-xxx" \
      -d '{
        "model": "gemma-4-31b",
        "messages": [
          {"role": "user", "content": "Hello!"}
        ]
      }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Default spec

per 1M tokens
Official price
Input $0.102 / Output $0.297 / Cache read $0.012
Official
Input $0.102 / Output $0.297 / Cache read $0.012
Discount
—
Prompt cache pricing

Cache read

Official price
$0.012
Official
$0.012
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Gemma 4 31B in Console with a prompt ready to edit or send.

Help me try gemma-4-31b with a short message at /v1/chat/completions. Show the reply, latency, and cost.

Use cases

  • Self-hostable assistant

    Prototype against the hosted model, then move to your own GPUs if data rules require it.

  • Multilingual chat

    Answer users in many languages from one model, which keeps support tooling simpler than running separate translators.

  • Code explanation

    Review snippets and write functions in settings where an open-weight model is a requirement for policy or hosting reasons.

  • Fine-tuning base

    Compare a hosted baseline before adapting the open weights to your domain.

Prompt examples

Translate this paragraph into Spanish, German, and Japanese and note any idiom you had to rephrase.

Explain what this SQL query does, then rewrite it so it avoids the correlated subquery.

Write a short, polite reply declining this meeting request and proposing next week.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

Is Gemma 4 31B open source?

Google publishes the weights under the Gemma terms, which makes it open-weight. Read the license on the model card before commercial use, because it is Gemma's own license.

How is Gemma different from Gemini?

Gemini models are Google's hosted proprietary line. Gemma models are smaller open-weight models built from related research, which you can download, fine-tune, and run yourself.

Does it support tool calling?

No. The hosted model handles text chat only, with text input and output and no tool calling. For agents that call functions, use a Gemini model instead.

What size is the model?

About 31 billion parameters, dense, the largest of the five Gemma 4 sizes: E2B, E4B, 12B, 26B A4B, and 31B. Smaller sizes target phones and laptops.

How much does Gemma 4 31B cost?

On TokenLab, Gemma 4 31B costs Input $0.102 / Output $0.297 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

Which endpoint should Gemma 4 31B use?

Use https://api.tokenlab.sh/v1/chat/completions for Gemma 4 31B. The request example below shows the matching code shape.

Which operations does Gemma 4 31B support?

Gemma 4 31B supports Text to text. Select an operation above to see its endpoint and request example.

Compare Gemma 4 31B

Sources

Reviewed Oct 2, 2026

Related models