MiniMax: MiniMax-M2.5-highspeed
minimax-m2.5-highspeed- Input / Output
- $0.60 / $2.40
- Context
- 205K
- Max output
- 64K
- Modalities
- Chat
- Capabilities
- Tool usePrompt Cache
About MiniMax-M2.5-highspeed
MiniMax-M2.5-highspeed is the faster serving option of MiniMax-M2.5, the February 2026 model tuned for programming, tool calling, search and office productivity. MiniMax states that M2.5 and its highspeed version give identical results, with the highspeed one running faster. It reasons, calls tools, and caches prompts, aimed at low-latency coding and agent workflows.
Where it works well
- MiniMax says the highspeed version matches M2.5's results while generating faster.
- M2.5 targets coding, tool calling and search, plus office productivity work.
- Reasoning, tool use and prompt caching are all available.
- M2.5 weights are published openly, so behavior can be checked against a self-hosted copy.
When to choose another model
- Text input only, so diagrams and photos must be described in words.
- No response-schema mode, unlike M2.1; structured output relies on prompting and validation in your code.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textResponses APIPOST/v1/responsesAPI format:curl https://api.tokenlab.sh/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxx" \ -d '{ "model": "minimax-m2.5-highspeed", "input": "Hello!" }'
Pricing
The Verified price applies to Verified, which costs less on most models. The Official price is the model maker's published price and applies to the more reliable Official route. Auto bills the route that completes the request.
Default spec
per 1M tokens- Official price
- Input $0.60 / Output $2.40 / Cache read $0.03 / Cache write $0.375
- Verified price
- —
- Discount
- —
Cache read
- Official price
- $0.03
- Verified price
- —
- Discount
- —
Cache write
- Official price
- $0.375
- Verified price
- —
- Discount
- —
| Official priceper 1M tokens | Verified priceper 1M tokens | Discount | |
|---|---|---|---|
| Default spec | Input $0.60 / Output $2.40 / Cache read $0.03 / Cache write $0.375 | — | — |
| Prompt cache pricing | |||
| Cache read | $0.03 | — | — |
| Cache write | $0.375 | — | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open MiniMax-M2.5-highspeed in Console with a prompt ready to edit or send.
Help me try minimax-m2.5-highspeed with a short message at /v1/responses. Show the reply, latency, and cost.
Use cases
Low-latency coding agents
Keep an agent responsive while it reads files, runs commands and revises its plan.
Search and tool pipelines
Chain web or database tools where each model step waits on a retrieval result.
Office task automation
Draft and revise spreadsheets, reports or slide text where several quick rounds are needed.
Prompt examples
Search the repo for every use of the deprecated client, list files, and patch the simplest ones.
Summarize this quarter's tickets by theme, then draft three bullet points for a weekly update.
Call lookup_price for each SKU in the list and report any that changed.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
How does M2.5-highspeed differ from M2.5?
MiniMax describes the two as identical in capability, differing only in speed. Use the highspeed version when the delay between steps matters in an interactive loop or agent.
What was M2.5 built for?
MiniMax's release note says it reached leading results in programming, tool calling and search, and office productivity. It is a general agent model rather than a coding-only one.
Does it support structured JSON schema output?
No response schema is offered for this model, though tool use is supported. If you need strictly typed output, validate it in your code or use M2.1, which has schema output.
Is it open source?
The M2.5 weights are published on Hugging Face under MiniMaxAI. The highspeed name refers to how the model is served and how fast it answers, not to a different model.
Should I use M2.5-highspeed or M2.7-highspeed?
Try M2.7-highspeed for newer agent features such as skills and dynamic tool search. Keep M2.5-highspeed if your prompts are tuned to it and results are stable.
How much does MiniMax-M2.5-highspeed cost?
On TokenLab, MiniMax-M2.5-highspeed costs Input $0.60 / Output $2.40 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of MiniMax-M2.5-highspeed?
MiniMax-M2.5-highspeed accepts up to 204,800 tokens of context and returns up to 65,536 tokens in one response.
Which endpoint should MiniMax-M2.5-highspeed use?
Use https://api.tokenlab.sh/v1/responses for MiniMax-M2.5-highspeed. The request example below shows the matching code shape.
Which operations does MiniMax-M2.5-highspeed support?
MiniMax-M2.5-highspeed supports Text to text. Select an operation above to see its endpoint and request example.
Compare MiniMax-M2.5-highspeed
Sources
Reviewed Oct 2, 2026