Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
A speed-oriented M2.7 endpoint for coding, refactoring and multi-step engineering assistance
Text generation with streaming and function calls through configured OpenAI-compatible or Anthropic-compatible integrations
A 204,800-token context window in the current invocation guide
A separately named Highspeed endpoint, described by MiniMax as the same model performance with faster inference; actual latency still needs measurement
Before you choose
The current API schema permits a 204,800-token request ceiling and recommends 65,536 for M2.x output. The separate overview lists 128K including thinking for the original M2. These documents do not establish one guaranteed completion length across every deployment; confirm the actual endpoint limit.
Thinking consumes the generation budget. A request ceiling is not extra capacity on top of a full context window, and a low budget can leave little or no final answer.
These M2.x endpoints are text-focused. Do not carry M3 image/video support or M3.1 thinking-depth levels into them. No knowledge cutoff is established here.
The reviewed schemas do not establish strict JSON-schema-constrained output for these IDs. Prompted JSON still needs parsing and validation; a function call is not proof a tool executed.
Provider speed descriptions are approximate, not latency guarantees. Request size, thinking, tool work and service load affect completion time.
Extended thinking
Thinking is always on for these M2.x endpoints and cannot be disabled. The compatibility documentation says a disabled value does not turn it off. reasoning_split changes response formatting, not whether the model reasons. No configurable effort levels are established for these IDs; do not borrow M3.1-Flash-Preview's effort controls.
Use the exact case-sensitive ID MiniMax-M2.7-highspeed. Standard and Highspeed are separate model IDs; switching one does not merely toggle a local UI preference.
Preserve the complete assistant response, reasoning content or thinking blocks, tool calls and tool-result messages in multi-turn workflows, following the selected compatibility format.
Use max_completion_tokens for new OpenAI-format integrations; max_tokens is the legacy field. The Anthropic-compatible API uses its own message schema. Inspect finish_reason and usage when output is truncated.
Validate tool arguments and returned data, and require approval before external writes. API compatibility does not imply every OpenAI or Anthropic parameter is implemented.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Run a focused code-review pass
Keep a fast interaction narrow and actionable.
Review this small diff for correctness only. Return the three highest-confidence issues, with file, line evidence, impact and a minimal test for each. If fewer than three are supported, say so. Do not spend the response on formatting preferences or infer execution results that are not included.
Workflow 02
Compare end-to-end response time
Measure the entire workflow rather than token speed alone.
Analyze these timings for standard and Highspeed requests: queue time, first token, generation, tool calls and retries. Compare total completion time and billed usage for successful tasks. Identify outliers and missing measurements. Do not infer model quality from speed or fill gaps with claimed provider benchmarks.
Workflow 03
Prepare an incremental refactor
Make each rapid iteration independently testable.
Break this refactor into small changes that can be reviewed independently. For each step state the behavior that must remain stable, code boundary, targeted tests and rollback condition. Keep the public interface unchanged unless the supplied requirements explicitly allow it. Do not apply changes or claim tests passed.
Developer reference
MiniMax API pricing
These are MiniMax API reference prices, not EZ Ai Assist subscription prices.
These rates follow MiniMax's current central pricing table. Use actual billed input, output, cache reads and cache writes; do not apply the cache-read rate to an entire request.
Documentation discrepancy: the older caching reference still lists $0.30 input for this Highspeed variant, while the current central pricing table lists $0.60. This guide uses the central table; confirm account billing before committing to a cost estimate.
Long thinking responses and multi-step tool loops can increase usage. Rates alone do not establish a task's total cost; measure representative requests on your chosen endpoint.
Common questions
A few things worth knowing.
Is Highspeed a different context size?
The invocation guide lists the same 204,800-token context for standard M2.7 and Highspeed. The documented distinction is faster inference, not a larger context entitlement.
What price should I use for planning?
The current central table lists $0.60 input, $0.06 cache read, $2.40 output and $0.375 cache write per million tokens. An older caching page conflicts on input price; confirm actual account billing before relying on an estimate.
Will it always finish a task faster?
No. The provider's approximate speed is not an end-to-end guarantee. Tool calls, reasoning length, retries and queueing can dominate task time; measure complete successful workflows.
Can I disable thinking to make it faster?
Not on the documented M2.x endpoints. Thinking remains on even when disabled is supplied; response-format options do not change that behavior. Use a suitable model and measured workflow rather than an unsupported switch.
Why is the output limit qualified?
The current shared API schema gives a 204,800-token request ceiling and recommends 65,536. The original M2 overview separately says 128K including thinking. Check the serving endpoint, preserve room within context and treat neither label as a guarantee of a finished answer that long.
Can this model use images or strict JSON Schema?
Image/video support is documented for M3, not these M2.x IDs. The reviewed sources also do not establish strict schema-constrained JSON here. Use verified capabilities and validate any JSON requested through a prompt.
Is the API table my EZ Ai Assist subscription price?
No. It is a developer reference for direct MiniMax API usage. App access, plan limits and exposed controls are separate; check the EZ Ai Assist pricing page and your workspace.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.