MiniMax: MiniMax-M3
minimax-m3- Input / Output
- $0.30 / $1.20
- Context
- 1M
- Max output
- 512K
- Modalities
- Chat
- Capabilities
- Tool usePrompt Cache
About MiniMax-M3
MiniMax-M3 is MiniMax's June 2026 open-weight model for coding, agentic work and long-context tasks. It uses MiniMax Sparse Attention to keep very long inputs affordable, reasons before answering, calls tools and caches prompts. It is the company's frontier M-series model and the successor to M2.7.
Where it works well
- MiniMax Sparse Attention cuts per-token compute on very long inputs compared with the previous generation.
- Targets autonomous task decomposition, tool invocation and multi-step reasoning for agents.
- Open weights are published by MiniMax, so behavior can be checked against a self-hosted copy.
- Reasoning, tool use and prompt caching can all be switched on for long agent loops.
When to choose another model
- Text input only, so charts and screenshots cannot be read directly.
- Reasoning and a very long input make requests slower than the highspeed M2.x models.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textResponses APIPOST/v1/responsesAPI format:curl https://api.tokenlab.sh/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxx" \ -d '{ "model": "minimax-m3", "input": "Hello!" }'
Pricing
The Verified price applies to Verified, which costs less on most models. The Official price is the model maker's published price and applies to the more reliable Official route. Auto bills the route that completes the request.
Input tokens <= 512K
per 1M tokens- Official price
- Input $0.30 / Output $1.20 / Cache read $0.06
- Verified price
- —
- Discount
- —
Input tokens > 512K
per 1M tokens- Official price
- Input $0.60 / Output $2.40 / Cache read $0.12
- Verified price
- —
- Discount
- —
Cache read
- Official price
- $0.06
- Verified price
- —
- Discount
- —
| Official priceper 1M tokens | Verified priceper 1M tokens | Discount | |
|---|---|---|---|
| Input tokens <= 512K | Input $0.30 / Output $1.20 / Cache read $0.06 | — | — |
| Input tokens > 512K | Input $0.60 / Output $2.40 / Cache read $0.12 | — | — |
| Prompt cache pricing | |||
| Cache read | $0.06 | — | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
- 30-day success rate
- 100.0%
- 30-day requests
- 100+
- Last active
- yesterday
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open MiniMax-M3 in Console with a prompt ready to edit or send.
Help me try minimax-m3 with a short message at /v1/responses. Show the reply, latency, and cost.
Use cases
Whole-repository work
Load a large codebase and have the model trace a change across modules before editing.
Planning agents
Give an agent a set of tools and a goal, and let it decompose the task and call tools step by step.
Long-document analysis
Review lengthy transcripts, specs or filings in one pass rather than chunking them.
Prompt examples
Here is the full repository. Where is the retry logic implemented, and what breaks if I change its backoff?
Here is the incident timeline and the service logs. Write the steps to reproduce the bug and name the likely faulty module.
Plan how to split this monolith into three services, and name the first tool call you would make for each.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What makes MiniMax-M3 different from M2.7?
MiniMax presents M3 as a new generation built around sparse attention for very long context and aimed at coding and agentic tasks, where M2.7 targets earlier agent harnesses.
Is MiniMax-M3 open weight?
MiniMax publishes M3 as an open-weight model, so you can self-host it if your data rules require that. Read the license on the model card before doing so.
Does it accept images and video?
No. MiniMax-M3 takes text input only, so use a model with image input when you need to read screenshots, charts, diagrams, or video frames directly.
Can it use tools?
Yes. It calls tools, and MiniMax aims it at agent planning, where a task is broken into steps and a tool is invoked at each one. Pass your tool definitions with the request and let the model decide when to call them.
When should I use M3 over a highspeed M2.x model?
Pick M3 for very long inputs or harder agent planning. Pick a highspeed M2.x variant when the delay of each turn matters more than depth of reasoning.
How much does MiniMax-M3 cost?
On TokenLab, MiniMax-M3 costs Input $0.30 / Output $1.20 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of MiniMax-M3?
MiniMax-M3 accepts up to 1,048,576 tokens of context and returns up to 524,288 tokens in one response.
Which endpoint should MiniMax-M3 use?
Use https://api.tokenlab.sh/v1/responses for MiniMax-M3. The request example below shows the matching code shape.
Which operations does MiniMax-M3 support?
MiniMax-M3 supports Text to text. Select an operation above to see its endpoint and request example.
Compare MiniMax-M3
Sources
Reviewed Oct 2, 2026