Choose Auto, TokenLab Verified, or Official for each request, with prices shown up front.See what's new

MiniMax: MiniMax-M2.1-highspeed

MiniMax M2.1 Highspeed is a speed-focused model for coding and tool-assisted tasks. It is useful for interactive agent loops in which an application alternates between generated instructions, tool results, and the next response.
Compare models
minimax-m2.1-highspeed
AvailableMiniMaxChat
Input / Output
$0.60 / $2.40
Context
205K
Max output
128K
Modalities
Chat
Capabilities
Tool usePrompt Cache

About MiniMax-M2.1-highspeed

MiniMax-M2.1-highspeed is the speed-focused version of MiniMax's M2.1 model for coding and tool-assisted tasks. It suits interactive agent loops in which an application alternates between generated instructions, tool results and the next response. It keeps M2.1's reasoning, tool use, JSON mode and schema output, trading nothing in features for faster turns.

Where it works well

  • Meant for loops where the model answers, a tool runs, and the model answers again, so each turn's delay adds up.
  • Keeps M2.1's JSON mode and schema-constrained output.
  • Reasoning is available when a step needs it.
  • Tool use and prompt caching are both supported.

When to choose another model

  • Text only; screenshots and other images cannot be sent.
  • M2.5-highspeed and M2.7-highspeed are later fast variants and may fit new builds better.

Getting started

  1. Create API key

    Create a key in Console, then use it with every model on the platform.

  2. Send your first request

    Copy the example for your language and run it against the endpoint.

    Text to textResponses API
    POST/v1/responses
    API format:
    curl https://api.tokenlab.sh/v1/responses \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer sk-xxx" \
      -d '{
        "model": "minimax-m2.1-highspeed",
        "input": "Hello!"
      }'

Pricing

The Verified price applies to Verified, which costs less on most models. The Official price is the model maker's published price and applies to the more reliable Official route. Auto bills the route that completes the request.

Default spec

per 1M tokens
Official price
Input $0.60 / Output $2.40 / Cache read $0.03 / Cache write $0.375
Verified price
—
Discount
—
Prompt cache pricing

Cache read

Official price
$0.03
Verified price
—
Discount
—

Cache write

Official price
$0.375
Verified price
—
Discount
—

Usage & activity

Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.

Usage & availability

Last 24 hours
Requests
Success rate
P95 latency
Total tokens
Model performance
Metrics appear once privacy and data volume thresholds are met.

Data is based on aggregate user requests, excluding status checks.

Open in Console

Open MiniMax-M2.1-highspeed in Console with a prompt ready to edit or send.

Help me try minimax-m2.1-highspeed with a short message at /v1/responses. Show the reply, latency, and cost.

Use cases

  • Interactive coding assistants

    Power an in-editor helper that must reply quickly to each instruction while it reads and edits files.

  • Agent loops with many short steps

    Run plan, call tool, read result cycles where model latency dominates total task time.

  • Structured responses in live apps

    Return schema-valid JSON to a front end that waits on the reply.

Prompt examples

Open utils/date.ts, find the timezone bug, and patch it. Then run the tests and report only failures.

Given these tool results, decide the next single action and output it as JSON with fields action and args.

Rename the function across the files listed and show a diff for each.

This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.

FAQ

What is the highspeed version of MiniMax-M2.1?

It is the same M2.1 generation presented as a speed-focused option for coding and tool-assisted tasks. Pick it when the delay between steps of an interactive agent matters.

Does it lose features compared with M2.1?

No. It keeps the M2.1 feature set of reasoning, tool use, prompt caching, JSON mode and response schema, so the change is turn speed. Quality on your own tasks is still worth testing.

Can I use it for structured output?

Yes. JSON mode and a response schema are both marked as supported, so you can require output in a fixed shape and parse it in your application.

Is there a newer fast MiniMax model?

Yes. MiniMax later shipped M2.5-highspeed and M2.7-highspeed as faster options. Compare them on your own tasks if you are starting a new project rather than keeping an existing integration.

Does it support images?

No. It has no image input, so it only reads and writes text. Choose MiniMax M3 when a prompt needs image or video understanding. Text prompts and tool results are all it takes in.

How much does MiniMax-M2.1-highspeed cost?

On TokenLab, MiniMax-M2.1-highspeed costs Input $0.60 / Output $2.40 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.

What are the context window and output limit of MiniMax-M2.1-highspeed?

MiniMax-M2.1-highspeed accepts up to 204,800 tokens of context and returns up to 131,072 tokens in one response.

Which endpoint should MiniMax-M2.1-highspeed use?

Use https://api.tokenlab.sh/v1/responses for MiniMax-M2.1-highspeed. The request example below shows the matching code shape.

Which operations does MiniMax-M2.1-highspeed support?

MiniMax-M2.1-highspeed supports Text to text. Select an operation above to see its endpoint and request example.

Compare MiniMax-M2.1-highspeed

Sources

Reviewed Oct 2, 2026

More from MiniMax

Related models