Back to models
NewChat Models

Mistral AI

Ministral 3 8B

Fastest

Compact 8B text-and-vision model for extraction and constrained workflows, with 256K context.

TextVision

At a glance

Know the model before you prompt.

Mistral AI API specifications
Context window
256K tokens (published)
Maximum output
Not verified
Inputs → output
Text + Images → Text
Knowledge cutoff
Not verified

API model ID: ministral-8b-2512

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • The 8B member of the Ministral 3 family, focused on efficient text and vision tasks
  • Text and image understanding for supplied documents, screenshots and mixed-input questions
  • Function calling and structured outputs listed on the model card; configure and validate them in the API integration
  • Open weights under Apache 2.0; hosted API usage and self-hosted operation have separate costs and responsibilities

Before you choose

  • The model card publishes a rounded 256k context. A model-specific maximum output and knowledge cutoff were not established in the reviewed sources; neither is inferred from another model.
  • Image understanding does not mean native image, audio or video generation. Check image readability and verify visual claims against the original.
  • A tool call requests work from your integration; it does not itself run code, search the web or authorize an external action. Validate arguments and require approval for consequential changes.
  • Open weights do not guarantee that a local deployment reproduces the hosted API's tools, performance or context settings. Review the model license and serving requirements.

Reasoning behavior

The reviewed model card does not establish a configurable reasoning-effort list or default for this API ID. Do not copy Medium 3.5 controls or settings from separately named Ministral Reasoning weights into this model. Use a bounded task, clear evidence and an explicit answer format.

  • Use ministral-8b-2512 for the documented model reference. A -latest alias may change over time; record the resolved version and re-evaluate before changing aliases.
  • The model card links Chat Completions, document Q&A, function calling and structured outputs. Agent services and built-in tools are separate integration features, not automatic access granted by a prompt.
  • Keep repeatable instructions at the start of requests to make caching useful. prompt_cache_key can improve cache-hit likelihood but does not guarantee a hit; inspect usage.prompt_tokens_details.cached_tokens.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Turn a screenshot into a field inventory

Create a compact, auditable visual extraction.

Inspect this form screenshot and return a table of field label, visible value, required marker and uncertainty. Preserve spelling exactly where readable, and use unreadable rather than guessing. Do not infer validation rules that are not shown. Finish with the fields a reviewer should inspect manually before this inventory is used.

Workflow 02

Label a set of short documents

Use a small taxonomy with explicit exceptions.

Apply this document taxonomy to the supplied excerpts. For each document return its ID, one label, a short supporting passage and an ambiguity flag. Do not create new labels. If two labels fit equally well, mark review needed and explain the conflict in one sentence. Keep the result in the table format provided.

Workflow 03

Compare compact-model candidates

Measure fit without assuming parameter count decides quality.

Given these outputs from compact models and the scoring rubric, compare extraction accuracy, instruction adherence and unsupported claims. Identify exact errors and examples where the rubric is ambiguous. Recommend a validation set and escalation threshold using the evidence supplied. Do not invent latency, token counts or hardware measurements.

Developer reference

Mistral AI API pricing

These are Mistral API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.15
Cached input$0.015
Output$0.15
  • The table uses the provider's standard processing rates. Batch, priority, regional deployments and other hosts may have different prices; recheck the chosen service before budgeting.
  • Cached input is billed at 10% of the ordinary input rate for eligible cache hits. Uncached input and generated output retain their own rates; caching does not make an entire request free.
  • The cache documentation describes shared-prefix matching in 64-token blocks; prompts shorter than 64 tokens do not qualify. Measure cached usage rather than assuming repeated requests are discounted.

Common questions

A few things worth knowing.

Why choose the 8B variant?

It is a compact text-and-vision option to evaluate for labeling, extraction and constrained assistance. Compare its errors and operating cost against the 3B and 14B variants on your own examples rather than assuming a universal winner.

Which model name belongs in an API request?

The card identifies ministral-8b-2512. The product family is Ministral 3, not Mistral 3; this guide keeps the existing Ministral catalog name and route.

Does this API ID expose reasoning-effort controls?

No model-specific list was established in the reviewed card. Separately named reasoning weights and other Mistral models are not evidence that the same controls apply here.

Does 256K context also mean 256K output?

No. Context capacity and maximum generated output are different limits. The source publishes 256k context but the reviewed model card does not establish an output ceiling; this guide does not invent one.

Can it analyze images or generate new ones?

The model card supports text and vision tasks with text output. Use it to discuss visible content, and check the result against the image. That is not a claim of native image, audio or video generation.

Will repeated prompts always receive the cached rate?

No. A shared prefix and cache routing can help, but inspect reported cached tokens to confirm a hit. Only eligible cached input uses the lower rate; output is billed separately.

Are these features and rates included in my EZ Ai Assist plan?

This is a provider API reference, not the app's subscription terms or a promise that every API feature is exposed. Check the app for access and the EZ Ai Assist pricing page for your plan.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider