MiniMax: MiniMax-M2.1-highspeed
minimax-m2.1-highspeed- Input / Output
- $0.60 / $2.40
- Context
- 205K
- Max output
- 128K
- Modalities
- Chat
- Capabilities
- Tool usePrompt Cache
About MiniMax-M2.1-highspeed
MiniMax-M2.1-highspeed is the speed-focused version of MiniMax's M2.1 model for coding and tool-assisted tasks. It suits interactive agent loops in which an application alternates between generated instructions, tool results and the next response. It keeps M2.1's reasoning, tool use, JSON mode and schema output, trading nothing in features for faster turns.
Where it works well
- Meant for loops where the model answers, a tool runs, and the model answers again, so each turn's delay adds up.
- Keeps M2.1's JSON mode and schema-constrained output.
- Reasoning is available when a step needs it.
- Tool use and prompt caching are both supported.
When to choose another model
- Text only; screenshots and other images cannot be sent.
- M2.5-highspeed and M2.7-highspeed are later fast variants and may fit new builds better.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textResponses APIPOST/v1/responsesAPI format:curl https://api.tokenlab.sh/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxx" \ -d '{ "model": "minimax-m2.1-highspeed", "input": "Hello!" }'
Pricing
The Verified price applies to Verified, which costs less on most models. The Official price is the model maker's published price and applies to the more reliable Official route. Auto bills the route that completes the request.
Default spec
per 1M tokens- Official price
- Input $0.60 / Output $2.40 / Cache read $0.03 / Cache write $0.375
- Verified price
- —
- Discount
- —
Cache read
- Official price
- $0.03
- Verified price
- —
- Discount
- —
Cache write
- Official price
- $0.375
- Verified price
- —
- Discount
- —
| Official priceper 1M tokens | Verified priceper 1M tokens | Discount | |
|---|---|---|---|
| Default spec | Input $0.60 / Output $2.40 / Cache read $0.03 / Cache write $0.375 | — | — |
| Prompt cache pricing | |||
| Cache read | $0.03 | — | — |
| Cache write | $0.375 | — | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open MiniMax-M2.1-highspeed in Console with a prompt ready to edit or send.
Help me try minimax-m2.1-highspeed with a short message at /v1/responses. Show the reply, latency, and cost.
Use cases
Interactive coding assistants
Power an in-editor helper that must reply quickly to each instruction while it reads and edits files.
Agent loops with many short steps
Run plan, call tool, read result cycles where model latency dominates total task time.
Structured responses in live apps
Return schema-valid JSON to a front end that waits on the reply.
Prompt examples
Open utils/date.ts, find the timezone bug, and patch it. Then run the tests and report only failures.
Given these tool results, decide the next single action and output it as JSON with fields action and args.
Rename the function across the files listed and show a diff for each.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What is the highspeed version of MiniMax-M2.1?
It is the same M2.1 generation presented as a speed-focused option for coding and tool-assisted tasks. Pick it when the delay between steps of an interactive agent matters.
Does it lose features compared with M2.1?
No. It keeps the M2.1 feature set of reasoning, tool use, prompt caching, JSON mode and response schema, so the change is turn speed. Quality on your own tasks is still worth testing.
Can I use it for structured output?
Yes. JSON mode and a response schema are both marked as supported, so you can require output in a fixed shape and parse it in your application.
Is there a newer fast MiniMax model?
Yes. MiniMax later shipped M2.5-highspeed and M2.7-highspeed as faster options. Compare them on your own tasks if you are starting a new project rather than keeping an existing integration.
Does it support images?
No. It has no image input, so it only reads and writes text. Choose MiniMax M3 when a prompt needs image or video understanding. Text prompts and tool results are all it takes in.
How much does MiniMax-M2.1-highspeed cost?
On TokenLab, MiniMax-M2.1-highspeed costs Input $0.60 / Output $2.40 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of MiniMax-M2.1-highspeed?
MiniMax-M2.1-highspeed accepts up to 204,800 tokens of context and returns up to 131,072 tokens in one response.
Which endpoint should MiniMax-M2.1-highspeed use?
Use https://api.tokenlab.sh/v1/responses for MiniMax-M2.1-highspeed. The request example below shows the matching code shape.
Which operations does MiniMax-M2.1-highspeed support?
MiniMax-M2.1-highspeed supports Text to text. Select an operation above to see its endpoint and request example.
Compare MiniMax-M2.1-highspeed
Sources
Reviewed Oct 2, 2026