Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Text and image input; text output
Adaptive thinking with configurable effort
Messages API tool-use workflows with an appropriate integration
Prompt caching with 5-minute and 1-hour write options
Asynchronous Message Batches processing
Before you choose
No native audio, video, or image output; accepting images does not make this an image generator.
Thinking cannot be switched off; lowering effort is different from disabling thinking.
Fast mode is a separate first-party API research preview, not a promise of lower latency or a toggle available in EZ Ai Assist.
Adaptive thinking and effort
lowmedium · defaulthighxhighmax
Adaptive thinking is always on; disabled, manual enabled/budget_tokens, and between_tools are rejected. The API default is medium, unlike Opus 5’s high. All five effort levels are supported. Set max_tokens with enough room for thinking and the answer, and compare higher effort against a measurable quality target.
Anthropic lists this model as active; released September 22, 2026. Retirement is not sooner than September 22, 2027. This is an availability commitment, not a scheduled retirement.
Normal Messages output is limited to 128,000 tokens. Up to 300,000 output tokens are available only through the Message Batches API beta with the output-300k-2026-03-24 header.
The displayed cutoff is Anthropic’s reliable knowledge cutoff, published to month precision. A cutoff does not replace checking current facts or supplied evidence.
The minimum cacheable prompt length is 512 tokens. A cache read price does not mean every request receives a cache hit.
Use the documented thinking display setting to request summarized thinking when needed; omitted display does not eliminate thinking charges.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Investigate an intermittent production failure
Build a testable account from traces rather than a confident guess.
Analyze the incident traces, code excerpts, and timeline provided below. Identify the smallest set of hypotheses consistent with all observations, and cite the evidence that supports or weakens each. Propose a discriminating test for each unresolved hypothesis, a safe mitigation, and rollback conditions. Do not invent telemetry or recommend destructive production actions.
Workflow 02
Audit a strategic recommendation
Stress-test evidence and assumptions before accepting the conclusion.
Audit this strategy memo using only its supporting documents. Separate observed facts, forecasts, and assumptions. Identify where the recommendation depends on weak evidence, conflicting figures, or an omitted alternative. Give source references and describe what evidence would change the decision. Finish with a short review agenda, not an unsupported replacement strategy.
Workflow 03
Design a staged integration review
Turn an ambitious change into checkable milestones.
From the supplied architecture and integration requirements, propose a staged implementation review. For each stage, specify dependencies, interfaces, failure modes, acceptance checks, and a rollback trigger. Mark unspecified ownership or service behavior as unknown. Keep external actions and code execution outside this request; provide a plan that a responsible engineer can validate.
Developer reference
Anthropic API pricing
These are Anthropic API reference prices, not EZ Ai Assist subscription prices.
Standard rates apply across the full 1M-token context window; no separate long-context surcharge is listed for this model.
Batch processing discounts input and output by 50%. Cache writes and reads have distinct rates; calculate from actual usage rather than applying a cache-read discount to the entire request.
Thinking tokens are billed as output even when their text is hidden. Tool charges and optional US-only inference (1.1× token pricing) can also affect the total.
Fast mode is a first-party research preview at 2× Standard ($8 input / $40 output per million); it is not available with Batch or partner-operated cloud platforms. Cache and residency modifiers still apply.
Verify current provider pricing before estimating spend. These rates do not describe an EZ Ai Assist subscription.
Common questions
A few things worth knowing.
What effort level does Opus 5.5 use by default?
Medium. This differs from Opus 5’s high default. Low, medium, high, xhigh, and max are available, but a higher level is not automatically the best trade-off for your workload.
Can adaptive thinking be turned off?
No. Thinking is always on for this model. Lower effort can reduce reasoning work, but disabled and between_tools requests are rejected. Thinking tokens count toward output usage and max_tokens.
What does Fast Mode change?
It is a separately priced first-party API research preview for lower-latency output. Its input/output rates are twice Standard, and it is unavailable with Batch or partner-operated platforms. It does not establish an app feature or a latency guarantee.
How are cache reads different from Opus 5?
Opus 5.5 lists $0.20 per million cache-read tokens, compared with Opus 5’s $0.50. Cache-write durations have their own rates, and only eligible reused prefixes receive cache-read pricing.
Is its output limit 128K or 300K?
The normal synchronous Messages limit is 128,000. The documented Message Batches beta can return up to 300,000 with its beta header. Do not apply the beta limit to an ordinary interactive request.
Is the model scheduled to retire next September?
No scheduled shutdown is established by its not-sooner-than September 22, 2027 commitment. That is a minimum availability commitment, not an announced end date.
Should I always choose Opus over Sonnet?
Compare both on the same representative tasks and review accuracy, response time, and cost. Choose the model that meets your requirements; the guide’s API features and prices are separate from EZ Ai Assist subscription access.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.