Back to models
NewChat Models

Mistral AI

Mistral Small 4

Fastest

Hybrid text-and-vision model combining instruction following, reasoning and coding with 256K context.

HybridCoding

At a glance

Know the model before you prompt.

Mistral AI API specifications
Context window
256K tokens (published)
Maximum output
Not verified
Inputs → output
Text + Images → Text
Knowledge cutoff
Not verified

API model ID: mistral-small-2603

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • A hybrid instruction, reasoning and coding model with 119B total parameters and 6.5B active parameters
  • Text and image understanding for supplied documents, screenshots and mixed-input questions
  • Function calling and structured outputs listed on the model card; configure and validate them in the API integration
  • Open weights under Apache 2.0; hosted API usage and self-hosted operation have separate costs and responsibilities

Before you choose

  • The model card publishes a rounded 256k context. A model-specific maximum output and knowledge cutoff were not established in the reviewed sources; neither is inferred from another model.
  • Image understanding does not mean native image, audio or video generation. Check image readability and verify visual claims against the original.
  • A tool call requests work from your integration; it does not itself run code, search the web or authorize an external action. Validate arguments and require approval for consequential changes.
  • Open weights do not guarantee that a local deployment reproduces the hosted API's tools, performance or context settings. Review the model license and serving requirements.

Choose the reasoning effort

nonehigh

The model-specific reasoning guide documents reasoning_effort none and high. none uses minimal internal reasoning without a visible thinking chunk; high returns thinking chunks before the answer and can increase latency and token use. No default is asserted here. Do not apply every value in a generic API enum to this model.

  • Use mistral-small-2603 for the documented model reference. A -latest alias may change over time; record the resolved version and re-evaluate before changing aliases.
  • For multi-turn tool workflows, preserve the full assistant message including returned thinking chunks. Removing that history can reduce quality; follow the provider's current SDK schema.
  • Keep repeatable instructions at the start of requests to make caching useful. prompt_cache_key can improve cache-hit likelihood but does not guarantee a hit; inspect usage.prompt_tokens_details.cached_tokens.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Triage support requests consistently

Turn mixed requests into a constrained work queue.

Classify these support requests using only the categories and escalation rules I supply. Return request ID, category, evidence, urgency and one proposed next step. Mark ambiguous cases for a person instead of inventing policy. Keep customer facts separate from your suggestions and do not send replies or change ticket status.

Workflow 02

Review a small bug fix

Keep hybrid coding work tied to a clear acceptance test.

Inspect this small bug fix against the reproduction steps and expected behavior. Identify concrete regressions, missing edge cases and any unrelated changes. Cite the relevant code for each finding. Suggest the smallest correction and a short test matrix. State clearly which results are predictions because no execution output was supplied.

Workflow 03

Extract facts from a product screenshot

Separate visible evidence from inferred behavior.

From this product screenshot, list the visible fields, button labels and status messages. Compare them with the requirements below and flag missing or confusing content. Do not infer what happens after a click. End with three questions to resolve before making a design change, keeping the current visual identity intact.

Developer reference

Mistral AI API pricing

These are Mistral API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.15
Cached input$0.015
Output$0.60
  • The table uses the provider's standard processing rates. Batch, priority, regional deployments and other hosts may have different prices; recheck the chosen service before budgeting.
  • Cached input is billed at 10% of the ordinary input rate for eligible cache hits. Uncached input and generated output retain their own rates; caching does not make an entire request free.
  • The cache documentation describes shared-prefix matching in 64-token blocks; prompts shorter than 64 tokens do not qualify. Measure cached usage rather than assuming repeated requests are discounted.

Common questions

A few things worth knowing.

Is Small 4 the same as Large 4?

No. This guide covers Mistral Small 4, API ID mistral-small-2603. Large 4 is a separate model; its context, price and controls should not be used as Small 4 specifications.

What does hybrid mean here?

The model combines instruction following, reasoning and coding. The provider lists 119B total parameters with 6.5B active; those architecture figures do not directly predict latency or quality on your task.

When should I try high reasoning?

Evaluate it on multi-step reasoning or coding tasks where an extra thinking pass can justify added time and tokens. The documented none setting uses minimal internal reasoning. Compare measured results rather than treating either setting as universally best.

Does 256K context also mean 256K output?

No. Context capacity and maximum generated output are different limits. The source publishes 256k context but the reviewed model card does not establish an output ceiling; this guide does not invent one.

Can it analyze images or generate new ones?

The model card supports text and vision tasks with text output. Use it to discuss visible content, and check the result against the image. That is not a claim of native image, audio or video generation.

Will repeated prompts always receive the cached rate?

No. A shared prefix and cache routing can help, but inspect reported cached tokens to confirm a hit. Only eligible cached input uses the lower rate; output is billed separately.

Are these features and rates included in my EZ Ai Assist plan?

This is a provider API reference, not the app's subscription terms or a promise that every API feature is exposed. Check the app for access and the EZ Ai Assist pricing page for your plan.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider