Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
A 128B dense model combining instruction following, reasoning and coding in one set of weights
Text and image understanding for supplied documents, screenshots and mixed-input questions
Function calling and structured outputs listed on the model card; configure and validate them in the API integration
Open weights under a Modified MIT license; hosted API usage and self-hosted operation have separate costs and responsibilities
Before you choose
The model card publishes a rounded 256k context. A model-specific maximum output and knowledge cutoff were not established in the reviewed sources; neither is inferred from another model.
Image understanding does not mean native image, audio or video generation. Check image readability and verify visual claims against the original.
A tool call requests work from your integration; it does not itself run code, search the web or authorize an external action. Validate arguments and require approval for consequential changes.
Open weights do not guarantee that a local deployment reproduces the hosted API's tools, performance or context settings. Review the model license and serving requirements.
Choose the reasoning effort
nonehigh
The model-specific reasoning guide documents reasoning_effort none and high. none uses minimal internal reasoning without a visible thinking chunk; high returns thinking chunks before the answer and can increase latency and token use. No default is asserted here. Do not apply every value in a generic API enum to this model.
Use mistral-medium-3-5 for the documented model reference. A -latest alias may change over time; record the resolved version and re-evaluate before changing aliases.
For multi-turn tool workflows, preserve the full assistant message including returned thinking chunks. Removing that history can reduce quality; follow the provider's current SDK schema.
Keep repeatable instructions at the start of requests to make caching useful. prompt_cache_key can improve cache-hit likelihood but does not guarantee a hit; inspect usage.prompt_tokens_details.cached_tokens.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Plan a repository change with checkpoints
Define the smallest verifiable implementation before editing.
Using this repository map, issue description and acceptance criteria, propose a minimal implementation plan. Identify the files and interfaces likely to change, dependencies to inspect, and tests for each checkpoint. Separate facts in the supplied code from assumptions. Do not claim to have opened files or run tests that I have not provided.
Workflow 02
Connect a UI screenshot to its code
Ground a visual review in the supplied implementation.
Compare this UI screenshot with the component code and design requirements below. Identify visible inconsistencies and the exact code likely responsible. Explain which conclusions need an interactive check. Recommend small changes that preserve branding and layout, then give keyboard and mobile verification steps without claiming the changes are already applied.
Workflow 03
Evaluate two reasoning settings
Decide whether extra reasoning helps your actual workload.
Design a controlled evaluation of none and high reasoning for these coding tasks. Keep inputs and acceptance criteria fixed. Create a rubric for correctness, unintended edits, test quality, latency and billed usage. List the data we must collect and the conditions for choosing either setting. Do not invent measurements or assume high is always better.
Developer reference
Mistral AI API pricing
These are Mistral API reference prices, not EZ Ai Assist subscription prices.
The table uses the provider's standard processing rates. Batch, priority, regional deployments and other hosts may have different prices; recheck the chosen service before budgeting.
Cached input is billed at 10% of the ordinary input rate for eligible cache hits. Uncached input and generated output retain their own rates; caching does not make an entire request free.
The cache documentation describes shared-prefix matching in 64-token blocks; prompts shorter than 64 tokens do not qualify. Measure cached usage rather than assuming repeated requests are discounted.
Common questions
A few things worth knowing.
What distinguishes Medium 3.5?
It combines instruction following, reasoning and coding in a 128B dense model, with vision input and a published 256k context. Treat suitability for a repository task as something to evaluate, not as a guarantee of autonomous correctness.
Is it a preview or generally available?
The current model card labels it GA. The original launch article used preview language; this guide follows the current model card rather than carrying that older status forward.
Which reasoning settings are documented?
The model-specific guide describes none and high. none retains minimal internal reasoning; high returns thinking chunks before the answer. Preserve returned assistant history when using tools, and do not infer a default or extra effort levels from a generic enum.
Does 256K context also mean 256K output?
No. Context capacity and maximum generated output are different limits. The source publishes 256k context but the reviewed model card does not establish an output ceiling; this guide does not invent one.
Can it analyze images or generate new ones?
The model card supports text and vision tasks with text output. Use it to discuss visible content, and check the result against the image. That is not a claim of native image, audio or video generation.
Will repeated prompts always receive the cached rate?
No. A shared prefix and cache routing can help, but inspect reported cached tokens to confirm a hit. Only eligible cached input uses the lower rate; output is billed separately.
Are these features and rates included in my EZ Ai Assist plan?
This is a provider API reference, not the app's subscription terms or a promise that every API feature is exposed. Check the app for access and the EZ Ai Assist pricing page for your plan.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.