Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Code generation, refactoring and complex tool-assisted work described in the M2.5 release
Text generation with streaming and function calls through configured OpenAI-compatible or Anthropic-compatible integrations
A 204,800-token context window in the current invocation guide
Multi-turn reasoning and tool workflows when the integration preserves assistant messages and tool results
Before you choose
The current API schema permits a 204,800-token request ceiling and recommends 65,536 for M2.x output. The separate overview lists 128K including thinking for the original M2. These documents do not establish one guaranteed completion length across every deployment; confirm the actual endpoint limit.
Thinking consumes the generation budget. A request ceiling is not extra capacity on top of a full context window, and a low budget can leave little or no final answer.
These M2.x endpoints are text-focused. Do not carry M3 image/video support or M3.1 thinking-depth levels into them. No knowledge cutoff is established here.
The reviewed schemas do not establish strict JSON-schema-constrained output for these IDs. Prompted JSON still needs parsing and validation; a function call is not proof a tool executed.
Provider speed descriptions are approximate, not latency guarantees. Request size, thinking, tool work and service load affect completion time.
Extended thinking
Thinking is always on for these M2.x endpoints and cannot be disabled. The compatibility documentation says a disabled value does not turn it off. reasoning_split changes response formatting, not whether the model reasons. No configurable effort levels are established for these IDs; do not borrow M3.1-Flash-Preview's effort controls.
Use the exact case-sensitive ID MiniMax-M2.5. Standard and Highspeed are separate model IDs; switching one does not merely toggle a local UI preference.
Preserve the complete assistant response, reasoning content or thinking blocks, tool calls and tool-result messages in multi-turn workflows, following the selected compatibility format.
Use max_completion_tokens for new OpenAI-format integrations; max_tokens is the legacy field. The Anthropic-compatible API uses its own message schema. Inspect finish_reason and usage when output is truncated.
Validate tool arguments and returned data, and require approval before external writes. API compatibility does not imply every OpenAI or Anthropic parameter is implemented.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Plan a migration from an existing model
Preserve known-good behavior while evaluating a replacement.
Using our current prompts, failure log and acceptance tests, create an evaluation plan for replacing an existing coding model. Separate behavior that must remain identical from areas open to improvement. Include tool-call validation, output parsing, latency, billed usage and rollback criteria. Do not recommend migration solely because a model is labeled legacy.
Workflow 02
Review a multi-step tool trace
Find where an agent workflow lost evidence or control.
Audit this tool trace against the user's stated task and allowed permissions. Identify incorrect arguments, unsupported assumptions, repeated work and unapproved state changes. Cite the relevant step for each issue and suggest a minimal prevention check. Do not treat a proposed tool call as proof the action actually completed.
Workflow 03
Create regression tests for a refactor
Protect behavior before changing implementation details.
From this before-and-after code and requirements, propose regression tests covering observable behavior, error handling and boundary cases. Explain which change each test protects. Keep test inputs realistic and distinguish suggested assertions from executed results. Avoid adding requirements that are not present in the supplied evidence.
Developer reference
MiniMax API pricing
These are MiniMax API reference prices, not EZ Ai Assist subscription prices.
These rates follow MiniMax's current central pricing table. Use actual billed input, output, cache reads and cache writes; do not apply the cache-read rate to an entire request.
Standard and Highspeed have different input/output prices. The cache-read price is also model-specific: do not substitute the M2.7 rate into an older M2.x estimate.
Long thinking responses and multi-step tool loops can increase usage. Rates alone do not establish a task's total cost; measure representative requests on your chosen endpoint.
Common questions
A few things worth knowing.
Is M2.5 still a current-model recommendation?
MiniMax now groups it under Legacy Models. This guide preserves its version-specific capabilities and pricing for evaluation and maintenance without presenting it as the newest offering.
Should launch-time prices override today's table?
No. The current pricing table is used here. Older launch material may use different names or rates; do not turn those into a current API quote without verification.
How should I assess a replacement?
Run representative tasks and known failure cases with the same rubric. Compare correctness, tool behavior, cost and end-to-end latency, and keep a tested rollback path.
Can I disable thinking to make it faster?
Not on the documented M2.x endpoints. Thinking remains on even when disabled is supplied; response-format options do not change that behavior. Use a suitable model and measured workflow rather than an unsupported switch.
Why is the output limit qualified?
The current shared API schema gives a 204,800-token request ceiling and recommends 65,536. The original M2 overview separately says 128K including thinking. Check the serving endpoint, preserve room within context and treat neither label as a guarantee of a finished answer that long.
Can this model use images or strict JSON Schema?
Image/video support is documented for M3, not these M2.x IDs. The reviewed sources also do not establish strict schema-constrained JSON here. Use verified capabilities and validate any JSON requested through a prompt.
Is the API table my EZ Ai Assist subscription price?
No. It is a developer reference for direct MiniMax API usage. App access, plan limits and exposed controls are separate; check the EZ Ai Assist pricing page and your workspace.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.