xAI: Grok 4.6
grok-4.6- Input / Output-50%
- $2.00 / $6.00$1.00 / $3.00
- Context
- 500K
- Released
- Aug 12, 2026
- Max output
- 128K
- Modalities
- VisionChat
- Capabilities
- Tool usePrompt Cache
About Grok 4.6
Grok 4.6 is xAI's model for coding, agentic work and knowledge-intensive tasks, released between Grok 4.5 and 4.7. It reads text and images, writes text, and offers reasoning effort from low to xhigh, defaulting to high. Teams use it to put reference documents and screenshots into a single working conversation for analysis, writing or software help.
Where it works well
- Handles documents and screenshots in the same conversation, so a bug report with an image needs no separate preprocessing.
- Reasoning effort has four levels, low through xhigh, with high on by default for harder questions.
- Function calling and structured outputs make it workable as the engine of a tool-using assistant.
- Aimed at knowledge work as well as code, which suits analysis and long-form writing tasks.
- Prompt caching helps when a large reference set is reused across a session.
When to choose another model
- Grok 4.7 followed it as xAI's flagship, so 4.6 mainly suits prompts already validated on it.
- It is a text-output model; creating images or audio needs a different model.
- Heavy reasoning settings lengthen responses, so latency-sensitive chat may do better with a lower effort or a non-reasoning model.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textResponses APIPOST/v1/responsesAPI format:curl https://api.tokenlab.sh/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxx" \ -d '{ "model": "grok-4.6", "input": "Hello!" }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Input tokens <= 200K
per 1M tokens- Official price
- Input $2.00 / Output $6.00 / Cache read $0.50
- TokenLab price
- Input $1.00 / Output $3.00 / Cache read $0.25
- Discount
- -50%
Input tokens 200K-500K
per 1M tokens- Official price
- Input $4.00 / Output $12.00 / Cache read $1.00
- TokenLab price
- Input $2.00 / Output $6.00 / Cache read $0.50
- Discount
- -50%
Cache read
- Official price
- $0.50
- TokenLab price
- $0.25
- Discount
- -50%
| Official priceper 1M tokens | TokenLab priceper 1M tokens | Discount | |
|---|---|---|---|
| Input tokens <= 200K | Input $2.00 / Output $6.00 / Cache read $0.50 | Input $1.00 / Output $3.00 / Cache read $0.25 | -50% |
| Input tokens 200K-500K | Input $4.00 / Output $12.00 / Cache read $1.00 | Input $2.00 / Output $6.00 / Cache read $0.50 | -50% |
| Prompt cache pricing | |||
| Cache read | $0.50 | $0.25 | -50% |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Grok 4.6 in Console with a prompt ready to edit or send.
Help me try grok-4.6 with a short message at /v1/responses. Show the reply, latency, and cost.
Use cases
Best for- Vision
Document analysis with figures
Upload a report together with its charts as images and ask for inconsistencies between the narrative and the numbers shown.
Software assistance
Diagnose a failing build from logs and a screenshot of the error, propose a patch, and explain the change in plain language.
Knowledge-work drafting
Turn a stack of meeting notes and reference material into a structured briefing, keeping claims tied to the supplied sources.
Prompt examples
Below are the quarterly report text and the revenue chart. Tell me where the written claims disagree with the chart.
This CI log and screenshot show the same failure. Explain the likely cause and give a minimal patch.
Combine the attached notes into a one-page brief for a new hire. Mark any statement that rests on a single note.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What is Grok 4.6 used for?
xAI built it for coding, agentic tasks and knowledge work. It accepts text and images, calls tools, returns structured output, and reasons at an adjustable effort. Typical uses are code assistance and analysis of documents with figures.
What are the reasoning settings on Grok 4.6?
Reasoning effort can be low, medium, high or xhigh, and high is the default. Choose a lower level for quick turns and a higher one for debugging, planning or analysis where accuracy matters most.
Does Grok 4.6 take image input?
Yes. Text and images are accepted, and the reply is text. That lets one request hold a document plus screenshots of the thing you are asking about.
How does Grok 4.6 differ from Grok 4.7?
Both share the same input types, tool support and effort levels. Grok 4.7 is the later release and xAI's recommended frontier model. Keep 4.6 where you have tuned prompts and tested results.
Can I get JSON from Grok 4.6?
Yes. Structured outputs are supported, so you can pass a schema and receive a response that matches it. This works alongside function calling in tool-based workflows.
How much does Grok 4.6 cost?
On TokenLab, Grok 4.6 costs Input $1.00 / Output $3.00 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of Grok 4.6?
Grok 4.6 accepts up to 500,000 tokens of context and returns up to 131,072 tokens in one response.
Which endpoint should Grok 4.6 use?
Use https://api.tokenlab.sh/v1/responses for Grok 4.6. The request example below shows the matching code shape.
Which operations does Grok 4.6 support?
Grok 4.6 supports Text to text. Select an operation above to see its endpoint and request example.
Compare Grok 4.6
Sources
Reviewed Oct 2, 2026