Back to models
NewChat Models

MiniMax

MiniMax M3

Best Overall

Multimodal coding model with text, image and video input and up to 1M context, subject to endpoint capacity.

1M contextMultimodal

At a glance

Know the model before you prompt.

MiniMax API specifications
Context window
1,000,000 tokens
Maximum output
524,288 tokens
Inputs → output
Text + Images + Video → Text
Knowledge cutoff
Not verified

API model ID: MiniMax-M3

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Text, image and video understanding in the documented M3 compatibility APIs
  • Long-context coding and agent workflows, with a published 1,000,000-token context in the invocation guide
  • Function calls and streaming through configured OpenAI-compatible and Anthropic-compatible integrations
  • Optional adaptive thinking for M3, distinct from the always-on behavior of M2.x and the separate M3.1 Flash Preview model

Before you choose

  • The product page qualifies context as up to 1M with a guaranteed minimum of 512K. Confirm the capacity available on your endpoint rather than assuming every request can use the published maximum.
  • The API schema permits a 524,288-token output ceiling and recommends 131,072. This is a request limit, not a guarantee of that much useful final text; thinking and the overall context budget must also fit.
  • Image and video input do not imply audio input or native media generation. Unsupported file types, size limits or unreadable frames can prevent useful analysis.
  • No knowledge cutoff or strict JSON-schema-constrained output support was established for M3 in the reviewed sources. Validate generated structured text and tool arguments.
  • M3 is not M3.1-Flash-Preview. The newer model's low-to-max thinking-depth controls must not be advertised for this ID.

Extended thinking

Defaults differ by API format: Anthropic-compatible M3 requests have thinking off by default; OpenAI-compatible Chat Completions have it on by default. Use thinking.type adaptive to enable it or disabled to skip it. M3 has no verified tunable effort-level list; do not borrow M3.1 controls.

  • Use the exact ID MiniMax-M3 and set thinking explicitly when migrating between compatibility APIs so the default does not silently change your workload.
  • The OpenAI-compatible API documents image_url and video_url content parts. Follow its current upload, file-size and sampling rules; do not assume a URL alone grants access to private content.
  • Preserve full assistant messages and tool history between turns. Keep tool execution under application control with argument validation and approval for consequential actions.
  • Monitor usage and truncation before increasing max_completion_tokens. Long context, long output and thinking all affect latency and cost; evaluate with real data rather than the maximum labels alone.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Connect a demo video to reported bugs

Ground a multimodal review in observable events.

Review this product demo video alongside the bug reports and acceptance criteria below. For each report, identify the relevant visible event and timestamp if available, distinguish confirmed behavior from untested assumptions, and propose a reproducible check. Do not infer hidden network responses or claim a bug is fixed. Finish with a prioritized verification list.

Workflow 02

Plan a long-context code investigation

Make a large input useful through bounded checkpoints.

Using the supplied repository excerpts, logs and issue history, create an investigation map for this failure. Identify the strongest evidence, conflicting explanations and the next minimal test for each hypothesis. Cite exact files or log entries. Work in checkpoints and do not invent code from files that were not included.

Workflow 03

Estimate a multimodal workflow budget

Account for input tiers and thinking without inventing usage.

Build a cost-estimation worksheet for our text, image and video workflow using only the usage measurements and provider rate table supplied. Separate ordinary input, cached input, output, thinking where billed and any priority tier. Flag requests crossing the published input threshold. Leave unknown token counts blank and state what must be measured before producing a total.

Developer reference

MiniMax API pricing

These are MiniMax API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Published M3 standard processing tiers · USD per 1,000,000 tokens
Token type≤ 512kinput tokens> 512kinput tokens
Input$0.30$0.60
Cache read$0.06$0.12
Output$1.20$2.40
  • The provider publishes the boundary as 512k input tokens. This table preserves that label rather than choosing an unverified 512,000 or 524,288 threshold. Check billing interpretation before estimating a boundary case.
  • These are the current displayed rates, labeled permanent 50% off by MiniMax, not the struck-through original rates or a limited-time sale invented by this guide.
  • Priority processing is separately priced at 1.5 times standard: below/at the published boundary, input $0.45, cache read $0.09 and output $1.80; above it, $0.90, $0.18 and $3.60 per million tokens.
  • No M3 cache-write rate is listed in the reviewed current table. Its omission is not a zero-price claim, and the M2.x cache-write price must not be substituted.

Common questions

A few things worth knowing.

Does every M3 request get one million tokens of context?

The invocation guide lists 1,000,000, but the product page says up to 1M with a guaranteed minimum of 512K. Confirm the actual endpoint capacity and keep the work in manageable checkpoints.

Can M3 understand images and video?

Yes, the compatibility documentation lists both alongside text. Use supported input formats and observe file limits. This does not establish audio support or image/video generation.

Why can thinking change when I switch SDK formats?

MiniMax documents thinking off by default for M3 in the Anthropic-compatible API, and on by default in OpenAI-compatible Chat Completions. Set adaptive or disabled explicitly to preserve intended behavior.

Does M3 support low, high or max effort settings?

The current documentation reserves tunable effort levels for M3.1-Flash-Preview. M3 supports enabling or disabling adaptive thinking, but this guide does not advertise those newer depth controls for M3.

When does the higher API rate apply?

The published table uses a 512k input-token threshold, with a higher rate above it. This guide preserves the provider's label rather than inventing an exact integer interpretation. Priority is an additional 1.5-times processing tier.

Is 524,288 a guaranteed final-answer length?

No. It is the API's maximum output setting; the documentation recommends 131,072. Thinking, usable context, termination conditions and task quality all affect the actual final answer.

Does this describe my EZ Ai Assist plan?

No. Model features, input limits and API rates are provider references. Check your EZ Ai Assist workspace and plan for the models and controls actually available there.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider