Back to models
LegacyChat Models

MiniMax

MiniMax M2.1 Highspeed

Legacy faster-serving M2.1 endpoint for multilingual code workflows, with 204,800 context.

FastMultilingual

At a glance

Know the model before you prompt.

MiniMax API specifications
Context window
204,800 tokens
Maximum output
API request ceiling: 204,800 tokens
Inputs → output
Text → Text
Knowledge cutoff
Not verified

API model ID: MiniMax-M2.1-highspeed

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • A faster-serving M2.1 variant for multilingual code assistance and refactoring
  • Text generation with streaming and function calls through configured OpenAI-compatible or Anthropic-compatible integrations
  • A 204,800-token context window in the current invocation guide
  • A separately named Highspeed endpoint, described by MiniMax as the same model performance with faster inference; actual latency still needs measurement

Before you choose

  • The current API schema permits a 204,800-token request ceiling and recommends 65,536 for M2.x output. The separate overview lists 128K including thinking for the original M2. These documents do not establish one guaranteed completion length across every deployment; confirm the actual endpoint limit.
  • Thinking consumes the generation budget. A request ceiling is not extra capacity on top of a full context window, and a low budget can leave little or no final answer.
  • These M2.x endpoints are text-focused. Do not carry M3 image/video support or M3.1 thinking-depth levels into them. No knowledge cutoff is established here.
  • The reviewed schemas do not establish strict JSON-schema-constrained output for these IDs. Prompted JSON still needs parsing and validation; a function call is not proof a tool executed.
  • Provider speed descriptions are approximate, not latency guarantees. Request size, thinking, tool work and service load affect completion time.

Extended thinking

Thinking is always on for these M2.x endpoints and cannot be disabled. The compatibility documentation says a disabled value does not turn it off. reasoning_split changes response formatting, not whether the model reasons. No configurable effort levels are established for these IDs; do not borrow M3.1-Flash-Preview's effort controls.

  • Use the exact case-sensitive ID MiniMax-M2.1-highspeed. Standard and Highspeed are separate model IDs; switching one does not merely toggle a local UI preference.
  • Preserve the complete assistant response, reasoning content or thinking blocks, tool calls and tool-result messages in multi-turn workflows, following the selected compatibility format.
  • Use max_completion_tokens for new OpenAI-format integrations; max_tokens is the legacy field. The Anthropic-compatible API uses its own message schema. Inspect finish_reason and usage when output is truncated.
  • Validate tool arguments and returned data, and require approval before external writes. API compatibility does not imply every OpenAI or Anthropic parameter is implemented.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Review a quick translation of code

Check semantic differences before accepting a fast port.

Review this code port from one language to another against the original behavior and requirements. Focus on collection semantics, exceptions, concurrency and numeric conversions that actually appear in the code. Cite each risk and propose a minimal test. Do not approve the port solely because it compiles or looks similar.

Workflow 02

Create a latency-aware routing policy

Decide when a fast endpoint is worth using.

Using these measured model results and task categories, propose a routing policy for interactive coding requests. Define acceptance thresholds, time budgets, fallback conditions and when to ask a person. Include actual cost per accepted task where available. Do not invent provider limits or choose routes solely from their marketing names.

Workflow 03

Prepare a safe fast-model fallback

Keep retries from changing the original task.

Design a fallback checklist for this coding workflow when the fast endpoint fails or times out. Preserve the user's scope, tool permissions and conversation evidence. Specify how to detect duplicate external actions and how to report an incomplete result. Do not retry state-changing calls automatically without a verified idempotency strategy.

Developer reference

MiniMax API pricing

These are MiniMax API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.60
Cache read$0.03
Output$2.40
Cache write$0.375
  • These rates follow MiniMax's current central pricing table. Use actual billed input, output, cache reads and cache writes; do not apply the cache-read rate to an entire request.
  • Documentation discrepancy: the older caching reference still lists $0.30 input for this Highspeed variant, while the current central pricing table lists $0.60. This guide uses the central table; confirm account billing before committing to a cost estimate.
  • Long thinking responses and multi-step tool loops can increase usage. Rates alone do not establish a task's total cost; measure representative requests on your chosen endpoint.

Common questions

A few things worth knowing.

Does Highspeed change the programming model?

MiniMax describes it as the same M2.1 performance with faster inference. It remains a distinct endpoint with its own price; test code quality and complete task time on your workload.

Is the quoted token speed a service guarantee?

No. MiniMax's approximate speed description does not cover every prompt, thinking budget, tool call or queue condition. Measure tail latency and retry behavior as well as average speed.

Why does this page retain a legacy notice?

The current model guide lists M2.1 Highspeed under Legacy Models. The notice preserves that distinction without claiming that the endpoint has shut down.

Can I disable thinking to make it faster?

Not on the documented M2.x endpoints. Thinking remains on even when disabled is supplied; response-format options do not change that behavior. Use a suitable model and measured workflow rather than an unsupported switch.

Why is the output limit qualified?

The current shared API schema gives a 204,800-token request ceiling and recommends 65,536. The original M2 overview separately says 128K including thinking. Check the serving endpoint, preserve room within context and treat neither label as a guarantee of a finished answer that long.

Can this model use images or strict JSON Schema?

Image/video support is documented for M3, not these M2.x IDs. The reviewed sources also do not establish strict schema-constrained JSON here. Use verified capabilities and validate any JSON requested through a prompt.

Is the API table my EZ Ai Assist subscription price?

No. It is a developer reference for direct MiniMax API usage. App access, plan limits and exposed controls are separate; check the EZ Ai Assist pricing page and your workspace.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider