Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

xAI: Grok 4.20 Non-Reasoning

Grok 4.20 Non-Reasoning provides the direct-response variant of the 4.20 family. It is useful for extraction, drafting, and tool-assisted conversation where a separate extended reasoning phase is not the focus.
Compare models
grok-4.20-non-reasoning
AvailablexAIChatCached input 84% less
Input / Output-50%
$1.25 / $2.50$0.625 / $1.25
Context
1M
Max output
128K
Modalities
VisionChat
Capabilities
Tool usePrompt Cache

About Grok 4.20 Non-Reasoning

Grok 4.20 Non-Reasoning is the direct-answer version of xAI's 4.20 model. It skips the extended thinking phase, so it replies quickly to text and image prompts and suits drafting, extraction and tool-assisted chat. xAI describes it as combining strict prompt adherence with a low hallucination rate. Pick the reasoning or multi-agent variants when a problem needs deliberate analysis.

Where it works well

  • Replies without a separate thinking phase, which keeps latency low for interactive chat and high-volume jobs.
  • xAI highlights strict prompt adherence and a low hallucination rate for this model, useful when output must follow a template.
  • Supports function calling and structured JSON output for extraction pipelines and tool-driven assistants.
  • Accepts images with text, so it can read a receipt or screenshot and answer in one step.
  • Prompt caching helps when many requests reuse the same long instructions or reference text.

When to choose another model

  • With no reasoning step, it is weaker on multi-stage logic or math than grok-4.20 or grok-4.7.
  • It does not split work across agents; open-ended research is better served by the multi-agent variant.
  • Output is text only, with no image, audio or video generation.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textResponses API
    POST/v1/responses
    API format:
    curl https://api.tokenlab.sh/v1/responses \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer sk-xxx" \
      -d '{
        "model": "grok-4.20-non-reasoning",
        "input": "Hello!"
      }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Default spec

per 1M tokens
Official price
Input $1.25 / Output $2.50 / Cache read $0.20
TokenLab price
Input $0.625 / Output $1.25 / Cache read $0.10
Discount
-50%
Prompt cache pricing

Cache read

Official price
$0.20
TokenLab price
$0.10
Discount
-50%

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Grok 4.20 Non-Reasoning in Console with a prompt ready to edit or send.

Help me try grok-4.20-non-reasoning with a short message at /v1/responses. Show the reply, latency, and cost.

Use cases

Best for
  • Vision
  • Customer-facing chat

    Answer product and account questions in a consistent tone, calling lookup functions where needed, without making the customer wait for a long thinking phase.

  • Bulk document extraction

    Run thousands of contracts, forms or emails through a fixed JSON schema, relying on strict instruction following to keep fields consistent.

  • Drafting and rewriting

    Produce first drafts of emails, summaries and release notes in a requested format, then refine them with short follow-up instructions.

Prompt examples

Summarize this support thread in three bullet points and one suggested reply. Keep the reply under 80 words and do not promise a refund.

Pull the party names, effective date and termination notice period from this contract into the JSON schema below. Return null for anything missing.

Rewrite this changelog for non-technical customers. Keep the headings, drop internal ticket numbers, and use plain sentences.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

What is Grok 4.20 Non-Reasoning?

It is the version of xAI's Grok 4.20 that answers directly instead of thinking first. It reads text and images, writes text, and supports function calling and structured outputs. It is aimed at fast, instruction-heavy work such as extraction, drafting and tool-assisted chat.

When should I use non-reasoning instead of grok-4.20?

Use it when the task is clear and the answer is mostly recall, rewriting or formatting. It responds sooner than the reasoning ID. Switch to grok-4.20 when the answer needs several steps of analysis, such as debugging or planning.

Does Grok 4.20 Non-Reasoning follow JSON schemas?

Yes. It supports structured outputs, so you can supply a schema and receive a response that fits it. Combined with xAI's emphasis on strict prompt adherence, this makes it a reasonable default for extraction and classification jobs.

Can it accept image input?

Yes, images can be sent alongside text, for instance a photographed form or a UI screenshot. The reply is text. The model cannot create or edit images.

Does it work with an existing OpenAI-style client?

Yes. Existing OpenAI-style client code works by changing only the model name, with tool definitions and JSON schemas sent in the usual fields as before.

How much does Grok 4.20 Non-Reasoning cost?

On TokenLab, Grok 4.20 Non-Reasoning costs Input $0.625 / Output $1.25 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of Grok 4.20 Non-Reasoning?

Grok 4.20 Non-Reasoning accepts up to 1,000,000 tokens of context and returns up to 131,072 tokens in one response.

Which endpoint should Grok 4.20 Non-Reasoning use?

Use https://api.tokenlab.sh/v1/responses for Grok 4.20 Non-Reasoning. The request example below shows the matching code shape.

Which operations does Grok 4.20 Non-Reasoning support?

Grok 4.20 Non-Reasoning supports Text to text. Select an operation above to see its endpoint and request example.

Compare Grok 4.20 Non-Reasoning

Sources

Reviewed Oct 2, 2026

More from Grok 4

Related models