Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

xAI: Grok 4.20

Grok 4.20 combines reasoning, visual input, and tool use in a long-context conversation. It is useful for investigations that bring together extensive text, images, and intermediate results before arriving at an answer.
Compare models
grok-4.20
AvailablexAIChatCached input 84% less
Input / Output-50%
$1.25 / $2.50$0.625 / $1.25
Context
1M
Max output
128K
Modalities
VisionChat
Capabilities
Tool usePrompt CacheReasoning

About Grok 4.20

Grok 4.20 is xAI's reasoning-enabled model in the 4.20 family, built for tool-calling workflows over very long inputs. It reads text and images and writes text, and it thinks before answering, which separates it from the non-reasoning and multi-agent siblings that share its family name. Choose it when an answer depends on stitching together many documents, screenshots and intermediate tool results.

Where it works well

  • Reasons through a problem before replying, so multi-step questions get worked answers instead of a first guess.
  • Accepts images alongside text, which lets one conversation mix screenshots, charts and written material.
  • Calls tools and returns structured JSON, so it can sit inside an agent loop or an extraction pipeline.
  • Handles a very long conversation history, useful for investigations that accumulate evidence over many turns.
  • Prompt caching lowers the cost of re-sending a large shared prefix across requests.

When to choose another model

  • The thinking step adds delay; the non-reasoning 4.20 variant answers faster when the task is plain drafting or extraction.
  • Newer Grok releases such as 4.7 target coding and agentic work more directly and may do better on difficult software tasks.
  • Output is text only, so it cannot produce images, audio or video.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textResponses API
    POST/v1/responses
    API format:
    curl https://api.tokenlab.sh/v1/responses \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer sk-xxx" \
      -d '{
        "model": "grok-4.20",
        "input": "Hello!"
      }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Input tokens <= 200K

per 1M tokens
Official price
Input $1.25 / Output $2.50 / Cache read $0.20
TokenLab price
Input $0.625 / Output $1.25 / Cache read $0.10
Discount
-50%

Input tokens > 200K

per 1M tokens
Official price
Input $2.50 / Output $5.00 / Cache read $0.40
TokenLab price
Input $1.25 / Output $2.50 / Cache read $0.20
Discount
-50%
Prompt cache pricing

Cache read

Official price
$0.20
TokenLab price
$0.10
Discount
-50%

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Grok 4.20 in Console with a prompt ready to edit or send.

Help me try grok-4.20 with a short message at /v1/responses. Show the reply, latency, and cost.

Use cases

Best for
  • Reasoning
  • Vision
  • Evidence-based investigations

    Feed it reports, exports and screenshots collected during an incident or audit, and ask for a reasoned timeline plus the open questions that remain.

  • Agent loops with tools

    Give it search, database or ticketing functions and let it decide which to call, read the results, and continue until the question is answered.

  • Structured extraction from messy inputs

    Turn scanned forms, invoices or chat logs into a fixed JSON schema, with the reasoning step helping on ambiguous entries.

Prompt examples

Three log exports and a screenshot of the dashboard follow. Build a timeline of the outage and flag anything the logs contradict.

Extract vendor, date, line items and tax from this invoice image as JSON matching the schema below. Mark any field you are unsure about.

You can call search_orders and get_customer. A customer says two items are missing from order 8812. Work out what to check and in what order.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

What is Grok 4.20?

Grok 4.20 is a chat model from xAI that accepts text and images and answers in text. This is the reasoning version of the family: it works through a problem before replying. xAI offers the same generation in non-reasoning and multi-agent forms.

How does grok-4.20 differ from grok-4.20-non-reasoning?

The base ID spends time reasoning before it answers, while the non-reasoning ID replies directly. Use the base ID for questions that need several steps of analysis, and the non-reasoning one for drafting, extraction and chat where speed matters more than depth.

Can Grok 4.20 read images?

Yes. You can send images with your text, for example screenshots, charts or photographed documents, and ask questions about them in the same conversation. Replies are text; the model does not generate or edit images.

Does Grok 4.20 support tool calling and JSON output?

Yes. It supports function calling and structured JSON responses, so you can wire it to your own functions or force an answer into a schema. That makes it usable for extraction jobs and for agents that act on external systems.

Can I use my existing OpenAI-style client with Grok 4.20?

Yes. Existing OpenAI-style client code works by changing the model name, and tool definitions and JSON schemas use the same fields you already send today.

How much does Grok 4.20 cost?

On TokenLab, Grok 4.20 costs Input $0.625 / Output $1.25 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of Grok 4.20?

Grok 4.20 accepts up to 1,000,000 tokens of context and returns up to 131,072 tokens in one response.

Which endpoint should Grok 4.20 use?

Use https://api.tokenlab.sh/v1/responses for Grok 4.20. The request example below shows the matching code shape.

Which operations does Grok 4.20 support?

Grok 4.20 supports Text to text. Select an operation above to see its endpoint and request example.

Compare Grok 4.20

Sources

Reviewed Oct 2, 2026

More from Grok 4

Related models