Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

xAI: Grok 4.6

Grok 4.6 handles text and image input with tool calling and a large context window. It can bring reference documents and screenshots into the same working conversation for analysis, writing, or software assistance.
Compare models
grok-4.6
AvailablexAIChatCached input 75% less
Input / Output-50%
$2.00 / $6.00$1.00 / $3.00
Context
500K
Released
Aug 12, 2026
Max output
128K
Modalities
VisionChat
Capabilities
Tool usePrompt Cache

About Grok 4.6

Grok 4.6 is xAI's model for coding, agentic work and knowledge-intensive tasks, released between Grok 4.5 and 4.7. It reads text and images, writes text, and offers reasoning effort from low to xhigh, defaulting to high. Teams use it to put reference documents and screenshots into a single working conversation for analysis, writing or software help.

Where it works well

  • Handles documents and screenshots in the same conversation, so a bug report with an image needs no separate preprocessing.
  • Reasoning effort has four levels, low through xhigh, with high on by default for harder questions.
  • Function calling and structured outputs make it workable as the engine of a tool-using assistant.
  • Aimed at knowledge work as well as code, which suits analysis and long-form writing tasks.
  • Prompt caching helps when a large reference set is reused across a session.

When to choose another model

  • Grok 4.7 followed it as xAI's flagship, so 4.6 mainly suits prompts already validated on it.
  • It is a text-output model; creating images or audio needs a different model.
  • Heavy reasoning settings lengthen responses, so latency-sensitive chat may do better with a lower effort or a non-reasoning model.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textResponses API
    POST/v1/responses
    API format:
    curl https://api.tokenlab.sh/v1/responses \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer sk-xxx" \
      -d '{
        "model": "grok-4.6",
        "input": "Hello!"
      }'

Pricing

Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.

Input tokens <= 200K

per 1M tokens
Official price
Input $2.00 / Output $6.00 / Cache read $0.50
TokenLab price
Input $1.00 / Output $3.00 / Cache read $0.25
Discount
-50%

Input tokens 200K-500K

per 1M tokens
Official price
Input $4.00 / Output $12.00 / Cache read $1.00
TokenLab price
Input $2.00 / Output $6.00 / Cache read $0.50
Discount
-50%
Prompt cache pricing

Cache read

Official price
$0.50
TokenLab price
$0.25
Discount
-50%

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open Grok 4.6 in Console with a prompt ready to edit or send.

Help me try grok-4.6 with a short message at /v1/responses. Show the reply, latency, and cost.

Use cases

Best for
  • Vision
  • Document analysis with figures

    Upload a report together with its charts as images and ask for inconsistencies between the narrative and the numbers shown.

  • Software assistance

    Diagnose a failing build from logs and a screenshot of the error, propose a patch, and explain the change in plain language.

  • Knowledge-work drafting

    Turn a stack of meeting notes and reference material into a structured briefing, keeping claims tied to the supplied sources.

Prompt examples

Below are the quarterly report text and the revenue chart. Tell me where the written claims disagree with the chart.

This CI log and screenshot show the same failure. Explain the likely cause and give a minimal patch.

Combine the attached notes into a one-page brief for a new hire. Mark any statement that rests on a single note.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

What is Grok 4.6 used for?

xAI built it for coding, agentic tasks and knowledge work. It accepts text and images, calls tools, returns structured output, and reasons at an adjustable effort. Typical uses are code assistance and analysis of documents with figures.

What are the reasoning settings on Grok 4.6?

Reasoning effort can be low, medium, high or xhigh, and high is the default. Choose a lower level for quick turns and a higher one for debugging, planning or analysis where accuracy matters most.

Does Grok 4.6 take image input?

Yes. Text and images are accepted, and the reply is text. That lets one request hold a document plus screenshots of the thing you are asking about.

How does Grok 4.6 differ from Grok 4.7?

Both share the same input types, tool support and effort levels. Grok 4.7 is the later release and xAI's recommended frontier model. Keep 4.6 where you have tuned prompts and tested results.

Can I get JSON from Grok 4.6?

Yes. Structured outputs are supported, so you can pass a schema and receive a response that matches it. This works alongside function calling in tool-based workflows.

How much does Grok 4.6 cost?

On TokenLab, Grok 4.6 costs Input $1.00 / Output $3.00 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of Grok 4.6?

Grok 4.6 accepts up to 500,000 tokens of context and returns up to 131,072 tokens in one response.

Which endpoint should Grok 4.6 use?

Use https://api.tokenlab.sh/v1/responses for Grok 4.6. The request example below shows the matching code shape.

Which operations does Grok 4.6 support?

Grok 4.6 supports Text to text. Select an operation above to see its endpoint and request example.

Compare Grok 4.6

Sources

Reviewed Oct 2, 2026

More from Grok 4

Related models