xAI: Grok 4.20 Multi-Agent
grok-4.20-multi-agent- Input / Output
- $1.25 / $2.50
- Context
- 1M
- Max output
- 128K
- Modalities
- VisionChat
- Capabilities
- Tool usePrompt CacheReasoning
About Grok 4.20 Multi-Agent
Grok 4.20 Multi-Agent is the xAI variant of Grok 4.20 that answers a hard question by running several agents in parallel. Each agent researches a different angle, and a leader agent merges their findings into one reply with supporting evidence. It suits research-style questions where breadth matters more than speed, unlike the single-pass reasoning and non-reasoning 4.20 models.
Where it works well
- Splits a question across several agents that explore different angles at once, then merges the results into a single answer.
- A leader agent writes the final reply, so you receive one coherent response rather than a pile of partial findings.
- xAI documents settings of about four agents for quick research and sixteen for deeper analysis, tied to the effort level.
- Suited to questions that need gathered sources and cross-checking rather than a single recall step.
- Reads images as well as text when the question includes screenshots or figures.
When to choose another model
- Running several agents takes longer and uses more tokens than one pass; a plain 4.20 model is better for simple lookups.
- Only the leader's final answer is visible, so you cannot inspect what each sub-agent concluded.
- It is a research-oriented variant; for a conventional coding or tool-calling agent, grok-4.7 is the more direct choice.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textResponses APIPOST/v1/responsesAPI format:curl https://api.tokenlab.sh/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxx" \ -d '{ "model": "grok-4.20-multi-agent", "input": "Hello!" }'
Pricing
Official price is the model maker's public baseline. TokenLab price is what you pay for this model on TokenLab.
Default spec
per 1M tokens- Official price
- Input $1.25 / Output $2.50 / Cache read $0.20
- Official
- Input $1.25 / Output $2.50 / Cache read $0.20
- Discount
- —
Cache read
- Official price
- $0.20
- Official
- $0.20
- Discount
- —
| Official priceper 1M tokens | Officialper 1M tokens | Discount | |
|---|---|---|---|
| Default spec | Input $1.25 / Output $2.50 / Cache read $0.20 | Input $1.25 / Output $2.50 / Cache read $0.20 | — |
| Prompt cache pricing | |||
| Cache read | $0.20 | $0.20 | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Grok 4.20 Multi-Agent in Console with a prompt ready to edit or send.
Help me try grok-4.20-multi-agent with a short message at /v1/responses. Show the reply, latency, and cost.
Use cases
Best for- Reasoning
- Vision
Market and competitor research
Ask for a sourced overview of a product category, with each agent covering different vendors or angles, and receive a single synthesized briefing.
Literature and policy reviews
Pose a question that spans many papers or regulations and let parallel agents gather the relevant points before the leader reconciles disagreements.
Hard open-ended analysis
Hand over a question with no single obvious method, such as why a metric moved, and get a consolidated view built from several lines of analysis.
Prompt examples
Compare the main approaches to vector search indexing for a ten-million-document corpus. Cover accuracy, memory use and operational complexity, and recommend one.
What are the open legal questions around training models on licensed images? Give the competing positions and the evidence for each.
Our signups fell 18 percent last month. Propose five independent explanations and say what data would confirm or rule out each.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What does the multi-agent version of Grok 4.20 do?
It launches several agents that work on a question in parallel. They gather and analyse information from different angles, then a leader agent combines the discussion into one final answer. You send a normal prompt and receive a single reply.
How many agents does grok-4.20-multi-agent use?
xAI describes two setups: four agents for quick, focused research and sixteen for deep, complex analysis. The count follows the reasoning effort you request, with low effort mapping to the smaller team and high effort to the larger one.
Is grok-4.20-multi-agent slower than grok-4.20?
Generally yes. Several agents explore the question before the leader writes the reply, so latency and token use rise with the agent count. For short questions or chat, the standard or non-reasoning 4.20 models respond sooner.
Can I see what each agent concluded?
No. xAI exposes only the leader agent's output, and the sub-agents' working state is not returned. Ask the model to cite the evidence in its final answer if you need to verify how a conclusion was reached.
Does it accept images?
Yes. Like the other 4.20 models it accepts images with text, so a research question can include a chart or screenshot. The response comes back as text.
How much does Grok 4.20 Multi-Agent cost?
On TokenLab, Grok 4.20 Multi-Agent costs Input $1.25 / Output $2.50 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of Grok 4.20 Multi-Agent?
Grok 4.20 Multi-Agent accepts up to 1,000,000 tokens of context and returns up to 131,072 tokens in one response.
Which endpoint should Grok 4.20 Multi-Agent use?
Use https://api.tokenlab.sh/v1/responses for Grok 4.20 Multi-Agent. The request example below shows the matching code shape.
Which operations does Grok 4.20 Multi-Agent support?
Grok 4.20 Multi-Agent supports Text to text. Select an operation above to see its endpoint and request example.
Compare Grok 4.20 Multi-Agent
Sources
Reviewed Oct 2, 2026