Back to models
LegacyChat Models

MiniMax

MiniMax M2.5 Highspeed

Legacy faster-serving M2.5 variant with 204,800 context and distinct input/output API rates.

FastCoding

At a glance

Know the model before you prompt.

MiniMax API specifications
Context window
204,800 tokens
Maximum output
API request ceiling: 204,800 tokens
Inputs → output
Text → Text
Knowledge cutoff
Not verified

API model ID: MiniMax-M2.5-highspeed

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • The faster-serving M2.5 variant for interactive code and tool workflows, retained in the legacy model list
  • Text generation with streaming and function calls through configured OpenAI-compatible or Anthropic-compatible integrations
  • A 204,800-token context window in the current invocation guide
  • A separately named Highspeed endpoint, described by MiniMax as the same model performance with faster inference; actual latency still needs measurement

Before you choose

  • The current API schema permits a 204,800-token request ceiling and recommends 65,536 for M2.x output. The separate overview lists 128K including thinking for the original M2. These documents do not establish one guaranteed completion length across every deployment; confirm the actual endpoint limit.
  • Thinking consumes the generation budget. A request ceiling is not extra capacity on top of a full context window, and a low budget can leave little or no final answer.
  • These M2.x endpoints are text-focused. Do not carry M3 image/video support or M3.1 thinking-depth levels into them. No knowledge cutoff is established here.
  • The reviewed schemas do not establish strict JSON-schema-constrained output for these IDs. Prompted JSON still needs parsing and validation; a function call is not proof a tool executed.
  • Provider speed descriptions are approximate, not latency guarantees. Request size, thinking, tool work and service load affect completion time.

Extended thinking

Thinking is always on for these M2.x endpoints and cannot be disabled. The compatibility documentation says a disabled value does not turn it off. reasoning_split changes response formatting, not whether the model reasons. No configurable effort levels are established for these IDs; do not borrow M3.1-Flash-Preview's effort controls.

  • Use the exact case-sensitive ID MiniMax-M2.5-highspeed. Standard and Highspeed are separate model IDs; switching one does not merely toggle a local UI preference.
  • Preserve the complete assistant response, reasoning content or thinking blocks, tool calls and tool-result messages in multi-turn workflows, following the selected compatibility format.
  • Use max_completion_tokens for new OpenAI-format integrations; max_tokens is the legacy field. The Anthropic-compatible API uses its own message schema. Inspect finish_reason and usage when output is truncated.
  • Validate tool arguments and returned data, and require approval before external writes. API compatibility does not imply every OpenAI or Anthropic parameter is implemented.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Check a low-latency assistant's failures

Look for errors hidden by fast responses.

Review these fast code-assistant responses against the source files and expected outcomes. Classify incorrect answers, incomplete patches and unsupported tool claims separately. Identify which failures need a larger context, a clearer prompt or human escalation. Do not excuse an error because the answer arrived quickly.

Workflow 02

Model a standard-versus-fast cost comparison

Use observed usage for both endpoints.

Using the supplied standard and Highspeed rates and request logs, design a comparison of cost per accepted coding task. Include input, output, cache reads, cache writes, retries and human rejection rate. Leave unmeasured values blank. State how we would decide whether a latency improvement justifies the extra expense.

Workflow 03

Audit version names in configuration

Avoid a silent model change during maintenance.

Inspect these configuration excerpts for model IDs, aliases and fallback routes. Produce a table of intended model, actual configured ID, provider and unresolved mismatch. Flag any name copied from an old announcement that lacks a verified API mapping. Do not replace IDs or deploy anything; give a safe verification sequence.

Developer reference

MiniMax API pricing

These are MiniMax API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.60
Cache read$0.03
Output$2.40
Cache write$0.375
  • These rates follow MiniMax's current central pricing table. Use actual billed input, output, cache reads and cache writes; do not apply the cache-read rate to an entire request.
  • Documentation discrepancy: the older caching reference still lists $0.30 input for this Highspeed variant, while the current central pricing table lists $0.60. This guide uses the central table; confirm account billing before committing to a cost estimate.
  • Long thinking responses and multi-step tool loops can increase usage. Rates alone do not establish a task's total cost; measure representative requests on your chosen endpoint.

Common questions

A few things worth knowing.

Is this the same API ID as M2.5?

No. Use MiniMax-M2.5-highspeed for the documented fast variant. Standard M2.5 has a separate ID and different input/output pricing.

Can I use a launch nickname as the model ID?

Use the current invocation guide's exact ID. Older announcement wording is not a verified API alias; test configuration against the provider documentation and your account.

Does legacy mean Highspeed has been retired?

The reviewed source lists it as legacy without establishing a shutdown date. Check actual endpoint availability and maintain migration tests rather than announcing an unsupported retirement.

Can I disable thinking to make it faster?

Not on the documented M2.x endpoints. Thinking remains on even when disabled is supplied; response-format options do not change that behavior. Use a suitable model and measured workflow rather than an unsupported switch.

Why is the output limit qualified?

The current shared API schema gives a 204,800-token request ceiling and recommends 65,536. The original M2 overview separately says 128K including thinking. Check the serving endpoint, preserve room within context and treat neither label as a guarantee of a finished answer that long.

Can this model use images or strict JSON Schema?

Image/video support is documented for M3, not these M2.x IDs. The reviewed sources also do not establish strict schema-constrained JSON here. Use verified capabilities and validate any JSON requested through a prompt.

Is the API table my EZ Ai Assist subscription price?

No. It is a developer reference for direct MiniMax API usage. App access, plan limits and exposed controls are separate; check the EZ Ai Assist pricing page and your workspace.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider