Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Multi-language programming and code refactoring emphasized by the M2.1 release
Text generation with streaming and function calls through configured OpenAI-compatible or Anthropic-compatible integrations
A 204,800-token context window in the current invocation guide
Multi-turn reasoning and tool workflows when the integration preserves assistant messages and tool results
Before you choose
The current API schema permits a 204,800-token request ceiling and recommends 65,536 for M2.x output. The separate overview lists 128K including thinking for the original M2. These documents do not establish one guaranteed completion length across every deployment; confirm the actual endpoint limit.
Thinking consumes the generation budget. A request ceiling is not extra capacity on top of a full context window, and a low budget can leave little or no final answer.
These M2.x endpoints are text-focused. Do not carry M3 image/video support or M3.1 thinking-depth levels into them. No knowledge cutoff is established here.
The reviewed schemas do not establish strict JSON-schema-constrained output for these IDs. Prompted JSON still needs parsing and validation; a function call is not proof a tool executed.
Provider speed descriptions are approximate, not latency guarantees. Request size, thinking, tool work and service load affect completion time.
Extended thinking
Thinking is always on for these M2.x endpoints and cannot be disabled. The compatibility documentation says a disabled value does not turn it off. reasoning_split changes response formatting, not whether the model reasons. No configurable effort levels are established for these IDs; do not borrow M3.1-Flash-Preview's effort controls.
Use the exact case-sensitive ID MiniMax-M2.1. Standard and Highspeed are separate model IDs; switching one does not merely toggle a local UI preference.
Preserve the complete assistant response, reasoning content or thinking blocks, tool calls and tool-result messages in multi-turn workflows, following the selected compatibility format.
Use max_completion_tokens for new OpenAI-format integrations; max_tokens is the legacy field. The Anthropic-compatible API uses its own message schema. Inspect finish_reason and usage when output is truncated.
Validate tool arguments and returned data, and require approval before external writes. API compatibility does not imply every OpenAI or Anthropic parameter is implemented.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Compare behavior across two implementations
Focus multilingual review on observable semantics.
Compare these two implementations in different programming languages against the shared specification. Identify differences in types, errors, ordering and edge-case behavior. Cite the relevant code and suggest language-appropriate tests. Do not assume the implementations are equivalent because their function names or happy-path outputs match.
Workflow 02
Plan a cross-language interface test
Protect a boundary between services.
Design contract tests for these services using the supplied request and response schemas. Cover missing fields, null handling, numeric precision, time zones and error shapes where relevant to the evidence. Separate confirmed requirements from questions. Do not invent undocumented endpoints or claim the test suite has run.
Workflow 03
Review a multilingual refactor proposal
Keep changes consistent without erasing language conventions.
Review this proposed refactor across the supplied language-specific modules. Check whether it preserves the shared contract while respecting each language's error and resource-handling conventions. Identify the smallest safe sequence and required tests. Do not rewrite unrelated modules or assume unavailable library behavior.
Developer reference
MiniMax API pricing
These are MiniMax API reference prices, not EZ Ai Assist subscription prices.
These rates follow MiniMax's current central pricing table. Use actual billed input, output, cache reads and cache writes; do not apply the cache-read rate to an entire request.
Standard and Highspeed have different input/output prices. The cache-read price is also model-specific: do not substitute the M2.7 rate into an older M2.x estimate.
Long thinking responses and multi-step tool loops can increase usage. Rates alone do not establish a task's total cost; measure representative requests on your chosen endpoint.
Common questions
A few things worth knowing.
What is M2.1's main positioning?
The provider emphasizes multi-language programming and refactoring. Treat that as a starting point for evaluating your language mix, not proof of correctness for every framework or library.
Is M2.1 a current or legacy model?
The current invocation guide places it in Legacy Models. A legacy designation does not establish a shutdown date; verify access and plan changes through your own regression suite.
How should I compare it with M2.7?
Use the same representative tasks, languages and tool schemas. Record correctness, unintended edits, token usage and task latency, and review cases where either model fails.
Can I disable thinking to make it faster?
Not on the documented M2.x endpoints. Thinking remains on even when disabled is supplied; response-format options do not change that behavior. Use a suitable model and measured workflow rather than an unsupported switch.
Why is the output limit qualified?
The current shared API schema gives a 204,800-token request ceiling and recommends 65,536. The original M2 overview separately says 128K including thinking. Check the serving endpoint, preserve room within context and treat neither label as a guarantee of a finished answer that long.
Can this model use images or strict JSON Schema?
Image/video support is documented for M3, not these M2.x IDs. The reviewed sources also do not establish strict schema-constrained JSON here. Use verified capabilities and validate any JSON requested through a prompt.
Is the API table my EZ Ai Assist subscription price?
No. It is a developer reference for direct MiniMax API usage. App access, plan limits and exposed controls are separate; check the EZ Ai Assist pricing page and your workspace.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.