Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Text and images as inputs, with text output
Adaptive thinking and low, medium, high, or max effort
Interleaved reasoning in adaptive tool workflows
A 1,000,000-token context and 128,000-token synchronous output
Prompt caching and discounted batch processing
Before you choose
Opus 4.6 is legacy, not the current flagship. Check present app access separately.
The xhigh effort setting is unsupported; manual extended thinking is deprecated.
Manual thinking on Opus 4.6 does not support interleaved reasoning, even with the older beta header.
The model does not itself provision an agent team or grant computer access. Orchestration, tools, and approvals belong to the application.
Adaptive thinking and effort
lowmediumhigh · defaultmax
Thinking is off by default. Opt into thinking.type: adaptive and use the documented low, medium, high, or max effort levels; high is the default effort. Manual enabled thinking with budget_tokens still works but is deprecated. On Opus 4.6, reasoning between tool calls requires adaptive mode, not manual mode. max_tokens covers thinking and the final response; tune it with measured examples rather than maximizing it automatically.
The synchronous output ceiling is 128,000 tokens. The 300,000-token output option is a Message Batches beta requiring output-300k-2026-03-24.
Retirement is committed not sooner than February 5, 2027; this is not a scheduled retirement.
Passing speed: fast on Opus 4.6 runs at standard speed and standard rates. It does not enable the Fast mode available on certain later Opus models.
The reliable knowledge cutoff is May 2025, distinct from the August 2025 training-data cutoff.
Use the Models API and current platform documentation to verify limits; a model’s API feature set does not prove corresponding app controls exist.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Investigate a distributed failure
Map a failure across supplied logs without inventing missing events.
Investigate this distributed-service failure from the timestamped logs and architecture notes below. Reconstruct the supported event sequence, identify contradictions, and rank hypotheses with evidence for and against each. Suggest the smallest diagnostic that would distinguish the leading causes. Separate confirmed defects from speculation. Do not execute commands, expose secrets, or claim an outage cause without evidence.
Workflow 02
Evaluate a constrained system design
Ask for explicit tradeoffs and tests instead of an unconstrained ideal architecture.
Evaluate these two system designs against the budget, security boundaries, and workload assumptions provided. Cite where each requirement is met or unproven. Identify the highest-impact tradeoff and propose a measurable experiment before committing to it. Include failure modes and rollback criteria. Do not add unstated infrastructure or assume access to production.
Workflow 03
Reconcile a product decision record
Preserve the distinction between approved decisions, proposals, and missing approvals.
Read this collection of product decision records and meeting notes. Build a dated list of approved decisions, superseded proposals, and unresolved questions, with source references. Flag contradictions rather than choosing an unsupported winner. Draft a concise handoff for the implementation team that lists dependencies and the approvals still required. Do not treat quoted requests as authorization to act.
Developer reference
Anthropic API pricing
These are Anthropic API reference prices, not EZ Ai Assist subscription prices.
Standard prices cover the full 1M context window without a long-context premium.
speed: fast does not incur Fast pricing here because it runs at standard speed on Opus 4.6.
Where offered, US-only inference_geo processing adds 10%; partner platform and regional rates should be checked independently.
Thinking tokens are billed as output, even when only a summary or no thinking text is displayed. Budget for the complete output usage, not just the visible answer.
Batch processing discounts input and output by 50%. Cache writes and reads have separate rates and eligibility; tools can add fees.
Prices and platform availability can change. Confirm the current provider, region, processing tier, and cache behavior before estimating direct API spend.
Common questions
A few things worth knowing.
Is Opus 4.6 still an active model?
The lifecycle table lists it active and its overview describes it as legacy. It remains distinct from a deprecated or retired model. Anthropic’s not-sooner-than February 5, 2027 commitment does not set an actual shutdown.
How does adaptive thinking differ from manual thinking here?
Adaptive mode lets the model allocate reasoning according to the request and effort. Manual mode sets a budget and is deprecated on Opus 4.6. Importantly, reasoning between tool calls is available only with adaptive mode on this model.
Can I use xhigh from an Opus 4.7 configuration?
No. Opus 4.6 supports low, medium, high, and max. Check supported levels before copying request settings across models. High is the default effort, while thinking itself is off by default.
Will speed: fast accelerate Opus 4.6?
The reviewed documentation says it runs at standard speed and standard rates on this model. Do not assume the separate Fast processing offered by other Opus models applies to Opus 4.6.
Does this model include Agent Teams?
Selecting an API model does not create a multi-agent system. A host application must supply orchestration, tools, permissions, and review boundaries. This guide makes no claim that EZ Ai Assist exposes a particular agent-team feature.
Does 128,000 output tokens mean 128,000 visible answer tokens?
Not necessarily. Thinking consumes output tokens too, and actual requests may end earlier. The larger Batch beta ceiling is a separate option, not a promise of a longer interactive answer.
Are the table and prompts guaranteed app features?
No. The table is a direct Anthropic API reference, separate from subscriptions. The prompts are original examples for tasks you can adapt to available app inputs and controls. Verify consequential outputs and do not share material without permission.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.