Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Text and image input; text output
Adaptive thinking with configurable effort
Messages API tool-use workflows with an appropriate integration
Prompt caching with 5-minute and 1-hour write options
Asynchronous Message Batches processing
Before you choose
No native audio, video, or image output; accepting images does not make this an image generator.
Active legacy status does not mean unavailable or scheduled for immediate shutdown.
Disabling thinking at xhigh/max returns an error; do not reuse a thinking-disabled request at a higher effort.
Fast mode is a separately priced first-party API research preview.
Adaptive thinking and effort
lowmediumhigh · defaultxhighmax
Adaptive thinking is on by default, with high effort. The disabled setting works at high or below, but is rejected at xhigh/max; between_tools and manual enabled/budget_tokens are not accepted. All five effort levels are supported. Effort is not a reliable response-length control: state a length requirement and leave room in max_tokens for thinking plus the answer.
Anthropic lists this model as active legacy; released July 24, 2026. Retirement is not sooner than July 24, 2027. This is an availability commitment, not a scheduled retirement.
Normal Messages output is limited to 128,000 tokens. Up to 300,000 output tokens are available only through the Message Batches API beta with the output-300k-2026-03-24 header.
The displayed cutoff is Anthropic’s reliable knowledge cutoff, published to month precision. A cutoff does not replace checking current facts or supplied evidence.
Compare against Opus 5.5 before migrating: the newer model defaults to medium effort and keeps thinking always on.
The minimum cacheable prompt length is 512 tokens. Preserve cache-eligible prefixes where your integration supports it.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Review a cross-service change proposal
Check contracts and failure handling across component boundaries.
Review this cross-service change proposal using the supplied API contracts and deployment constraints. Identify incompatible assumptions, data ownership gaps, and partial-failure cases. Reference the relevant contract for each finding. Recommend a staged test strategy and explicitly mark questions that require a service owner. Do not invent undocumented guarantees.
Workflow 02
Reconcile a financial operations workbook
Ask for auditable checks without treating estimates as verified numbers.
Inspect the supplied workbook tables and accounting definitions. Reconcile the reported totals, units, date ranges, and aggregation rules. Identify mismatches with row or cell references and distinguish arithmetic errors from missing or incompatible data. Describe repeatable validation steps and unresolved questions. Do not invent missing values or present a forecast as an actual result.
Workflow 03
Prepare a technical handoff
Preserve the difference between completed work and planned work.
Create a technical handoff from these implementation notes, test results, and open issues. Separate verified behavior, unverified assumptions, known limitations, and next actions. Link each completion claim to evidence in the supplied material. Identify operational checks and rollback information a new owner needs. Do not label a planned test or proposed fix as completed.
Developer reference
Anthropic API pricing
These are Anthropic API reference prices, not EZ Ai Assist subscription prices.
Standard rates apply across the full 1M-token context window; no separate long-context surcharge is listed for this model.
Batch processing discounts input and output by 50%. Cache writes and reads have distinct rates; calculate from actual usage rather than applying a cache-read discount to the entire request.
Thinking tokens are billed as output even when their text is hidden. Tool charges and optional US-only inference (1.1× token pricing) can also affect the total.
Fast mode is a first-party research preview at 2× Standard ($10 input / $50 output per million); it is not available with Batch or partner-operated cloud platforms. Cache and residency modifiers still apply.
Verify current provider pricing before estimating spend. These rates do not describe an EZ Ai Assist subscription.
Common questions
A few things worth knowing.
What does active legacy mean?
Opus 5 remains available in Anthropic’s API but is no longer the current Opus model. Legacy status is not an announced shutdown. Check current provider and app availability before relying on it.
What changes with Opus 5.5?
Opus 5.5 has different prices, a newer cutoff, a medium effort default, and always-on thinking. Compare your evaluation results and request configuration instead of assuming a model-ID swap preserves cost or behavior.
Can thinking be disabled on Opus 5?
Yes at low, medium, or high effort. It cannot be disabled at xhigh or max, and between_tools is not supported. This differs from Opus 5.5, where thinking is always on.
Will low effort make every response shorter?
Not reliably. State an explicit response length or output format when you need brevity. Effort changes the work devoted to a response; max_tokens also includes thinking.
Is Fast Mode included at the base rate?
No. The first-party research preview is priced at $10 input and $50 output per million tokens, twice Standard. It cannot be combined with Batch, and app availability must be checked separately.
Is July 24, 2027 its retirement date?
No. The documentation says retirement will not happen sooner than that date. It does not announce a shutdown on that date. Follow the lifecycle page for any later notice.
Can I use these prompts and output limits in the app?
The prompts are starting points for supported app inputs. The 128,000-token normal API output limit and separate 300,000-token Batch beta are provider specifications, not a promise of app capacity or plan access.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.