Back to models
LegacyChat Models

MiniMax

MiniMax M2

Original legacy M2 for text, reasoning and tool workflows, with 204,800 context and qualified output limits.

AgenticReasoning

At a glance

Know the model before you prompt.

MiniMax API specifications
Context window
204,800 tokens
Maximum output
API request ceiling: 204,800 tokens
Inputs → output
Text → Text
Knowledge cutoff
Not verified

API model ID: MiniMax-M2

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • The original M2 release's coding, reasoning and agentic tool-workflow capabilities
  • Text generation with streaming and function calls through configured OpenAI-compatible or Anthropic-compatible integrations
  • A 204,800-token context window in the current invocation guide
  • Multi-turn reasoning and tool workflows when the integration preserves assistant messages and tool results

Before you choose

  • The current API schema permits a 204,800-token request ceiling and recommends 65,536 for M2.x output. The separate overview lists 128K including thinking for the original M2. These documents do not establish one guaranteed completion length across every deployment; confirm the actual endpoint limit.
  • Thinking consumes the generation budget. A request ceiling is not extra capacity on top of a full context window, and a low budget can leave little or no final answer.
  • These M2.x endpoints are text-focused. Do not carry M3 image/video support or M3.1 thinking-depth levels into them. No knowledge cutoff is established here.
  • The reviewed schemas do not establish strict JSON-schema-constrained output for these IDs. Prompted JSON still needs parsing and validation; a function call is not proof a tool executed.
  • Provider speed descriptions are approximate, not latency guarantees. Request size, thinking, tool work and service load affect completion time.

Extended thinking

Thinking is always on for these M2.x endpoints and cannot be disabled. The compatibility documentation says a disabled value does not turn it off. reasoning_split changes response formatting, not whether the model reasons. No configurable effort levels are established for these IDs; do not borrow M3.1-Flash-Preview's effort controls.

  • Use the exact case-sensitive ID MiniMax-M2. Standard and Highspeed are separate model IDs; switching one does not merely toggle a local UI preference.
  • Preserve the complete assistant response, reasoning content or thinking blocks, tool calls and tool-result messages in multi-turn workflows, following the selected compatibility format.
  • Use max_completion_tokens for new OpenAI-format integrations; max_tokens is the legacy field. The Anthropic-compatible API uses its own message schema. Inspect finish_reason and usage when output is truncated.
  • Validate tool arguments and returned data, and require approval before external writes. API compatibility does not imply every OpenAI or Anthropic parameter is implemented.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Document an existing agent workflow

Make legacy behavior reviewable before changing it.

Document this existing agent workflow using its configuration, prompts and execution logs. Map each step to its evidence, tool permissions and success condition. Identify undocumented assumptions and points where the model can exceed the user's intent. Do not infer capabilities from newer model documentation or claim a proposed step actually ran.

Workflow 02

Investigate a truncated answer

Separate a budget issue from a reasoning failure.

Review this request configuration, usage report and truncated response. Identify plausible causes involving output budget, thinking tokens, context length or stop conditions, and cite the evidence for each. State which provider limit needs confirmation. Propose a minimal diagnostic sequence without inventing missing token counts or silently raising production limits.

Workflow 03

Define a legacy-model replacement gate

Require evidence before a version change.

Create a replacement gate for this legacy model using our acceptance tests, tool traces and known failures. Define the minimum pass criteria, regression blockers, cost and latency measurements, approval owner and rollback path. Keep unknowns explicit and do not recommend a successor based only on recency or an unverified benchmark.

Developer reference

MiniMax API pricing

These are MiniMax API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.30
Cache read$0.03
Output$1.20
Cache write$0.375
  • These rates follow MiniMax's current central pricing table. Use actual billed input, output, cache reads and cache writes; do not apply the cache-read rate to an entire request.
  • Standard and Highspeed have different input/output prices. The cache-read price is also model-specific: do not substitute the M2.7 rate into an older M2.x estimate.
  • Long thinking responses and multi-step tool loops can increase usage. Rates alone do not establish a task's total cost; measure representative requests on your chosen endpoint.

Common questions

A few things worth knowing.

Which M2 model does this page cover?

The original MiniMax-M2 API ID, not M2.1, M2.5, M2.7 or M2-her. The guide keeps its identity and legacy status separate from those related names.

Why do sources disagree about maximum output?

The model overview says 128K including thinking for original M2, while the current shared API schema permits a 204,800-token request ceiling. This guide exposes the difference and asks you to verify the actual serving limit instead of silently choosing one as guaranteed capacity.

Must I migrate because it is legacy?

The reviewed sources do not establish a shutdown deadline. Confirm availability and evaluate a replacement using your own failures, contracts and rollback requirements; the label alone is not an emergency migration instruction.

Can I disable thinking to make it faster?

Not on the documented M2.x endpoints. Thinking remains on even when disabled is supplied; response-format options do not change that behavior. Use a suitable model and measured workflow rather than an unsupported switch.

Why is the output limit qualified?

The current shared API schema gives a 204,800-token request ceiling and recommends 65,536. The original M2 overview separately says 128K including thinking. Check the serving endpoint, preserve room within context and treat neither label as a guarantee of a finished answer that long.

Can this model use images or strict JSON Schema?

Image/video support is documented for M3, not these M2.x IDs. The reviewed sources also do not establish strict schema-constrained JSON here. Use verified capabilities and validate any JSON requested through a prompt.

Is the API table my EZ Ai Assist subscription price?

No. It is a developer reference for direct MiniMax API usage. App access, plan limits and exposed controls are separate; check the EZ Ai Assist pricing page and your workspace.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider