Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

MiniMax: MiniMax-M2.5-highspeed

High-speed MiniMax model for low-latency coding and agent workflows
Compare models
minimax-m2.5-highspeed
AvailableMiniMaxChat
Input / Output
$0.60 / $2.40
Context
205K
Max output
64K
Modalities
Chat
Capabilities
Tool usePrompt Cache

About MiniMax-M2.5-highspeed

MiniMax-M2.5-highspeed is the faster serving option of MiniMax-M2.5, the February 2026 model tuned for programming, tool calling, search and office productivity. MiniMax states that M2.5 and its highspeed version give identical results, with the highspeed one running faster. It reasons, calls tools, and caches prompts, aimed at low-latency coding and agent workflows.

Where it works well

  • MiniMax says the highspeed version matches M2.5's results while generating faster.
  • M2.5 targets coding, tool calling and search, plus office productivity work.
  • Reasoning, tool use and prompt caching are all available.
  • M2.5 weights are published openly, so behavior can be checked against a self-hosted copy.

When to choose another model

  • Text input only, so diagrams and photos must be described in words.
  • No response-schema mode, unlike M2.1; structured output relies on prompting and validation in your code.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textResponses API
    POST/v1/responses
    API format:
    curl https://api.tokenlab.sh/v1/responses \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer sk-xxx" \
      -d '{
        "model": "minimax-m2.5-highspeed",
        "input": "Hello!"
      }'

Pricing

The Verified price applies to Verified, which costs less on most models. The Official price is the model maker's published price and applies to the more reliable Official route. Auto bills the route that completes the request.

Default spec

per 1M tokens
Official price
Input $0.60 / Output $2.40 / Cache read $0.03 / Cache write $0.375
Verified price
—
Discount
—
Prompt cache pricing

Cache read

Official price
$0.03
Verified price
—
Discount
—

Cache write

Official price
$0.375
Verified price
—
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open MiniMax-M2.5-highspeed in Console with a prompt ready to edit or send.

Help me try minimax-m2.5-highspeed with a short message at /v1/responses. Show the reply, latency, and cost.

Use cases

  • Low-latency coding agents

    Keep an agent responsive while it reads files, runs commands and revises its plan.

  • Search and tool pipelines

    Chain web or database tools where each model step waits on a retrieval result.

  • Office task automation

    Draft and revise spreadsheets, reports or slide text where several quick rounds are needed.

Prompt examples

Search the repo for every use of the deprecated client, list files, and patch the simplest ones.

Summarize this quarter's tickets by theme, then draft three bullet points for a weekly update.

Call lookup_price for each SKU in the list and report any that changed.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

How does M2.5-highspeed differ from M2.5?

MiniMax describes the two as identical in capability, differing only in speed. Use the highspeed version when the delay between steps matters in an interactive loop or agent.

What was M2.5 built for?

MiniMax's release note says it reached leading results in programming, tool calling and search, and office productivity. It is a general agent model rather than a coding-only one.

Does it support structured JSON schema output?

No response schema is offered for this model, though tool use is supported. If you need strictly typed output, validate it in your code or use M2.1, which has schema output.

Is it open source?

The M2.5 weights are published on Hugging Face under MiniMaxAI. The highspeed name refers to how the model is served and how fast it answers, not to a different model.

Should I use M2.5-highspeed or M2.7-highspeed?

Try M2.7-highspeed for newer agent features such as skills and dynamic tool search. Keep M2.5-highspeed if your prompts are tuned to it and results are stable.

How much does MiniMax-M2.5-highspeed cost?

On TokenLab, MiniMax-M2.5-highspeed costs Input $0.60 / Output $2.40 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of MiniMax-M2.5-highspeed?

MiniMax-M2.5-highspeed accepts up to 204,800 tokens of context and returns up to 65,536 tokens in one response.

Which endpoint should MiniMax-M2.5-highspeed use?

Use https://api.tokenlab.sh/v1/responses for MiniMax-M2.5-highspeed. The request example below shows the matching code shape.

Which operations does MiniMax-M2.5-highspeed support?

MiniMax-M2.5-highspeed supports Text to text. Select an operation above to see its endpoint and request example.

Compare MiniMax-M2.5-highspeed

Sources

Reviewed Oct 2, 2026

More from MiniMax

Related models