Back to models
LegacyChat Models

Anthropic

Claude Sonnet 4.6

Best for Coding

Anthropic’s active legacy Sonnet model for coding and analysis, with optional adaptive thinking and a 1M-token context.

1M contextCoding

At a glance

Know the model before you prompt.

Anthropic API specifications
Context window
1,000,000 tokens
Maximum output
128,000 tokens
Inputs → output
Text + Images → Text
Knowledge cutoff
August 2025

API model ID: claude-sonnet-4-6

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Text and image input with text output
  • Opt-in adaptive thinking with low, medium, high, and max effort
  • Automatic reasoning between tool calls in adaptive mode
  • A 1M-token context for supplied documents and code
  • Prompt caching and discounted asynchronous Message Batches

Before you choose

  • Legacy model: this is not the current Sonnet generation or a guarantee of app availability.
  • The xhigh effort level is not supported on Sonnet 4.6.
  • Manual extended thinking still works but is deprecated; prefer adaptive thinking for new integrations.
  • A large context is capacity, not a guarantee of complete retrieval or accurate citations. Validate evidence and tool results.

Adaptive thinking and effort

lowmediumhigh · defaultmax

Thinking is off by default. Enable thinking.type: adaptive; high is the default effort when no effort is specified. Sonnet 4.6 supports low, medium, high, and max, not xhigh. Manual enabled thinking with budget_tokens remains functional but deprecated. Adaptive mode interleaves thinking with tool calls automatically. Leave room in max_tokens for both reasoning and the final response; compare cost and accuracy on your task.

  • The standard synchronous maximum output is 128,000 tokens. Up to 300,000 output tokens is a Message Batches beta with output-300k-2026-03-24, not the ordinary Messages API limit.
  • Anthropic commits to retirement not sooner than February 17, 2027. That is not a scheduled retirement.
  • The reliable knowledge cutoff is August 2025; it is distinct from the January 2026 training-data cutoff.
  • Manual-mode interleaving with the older beta header still works but is deprecated. Adaptive thinking requires no interleaved-thinking beta header.
  • Tool integration and permissions are separate from the model’s reasoning configuration; no tool is enabled merely by selecting this model.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Draft a focused implementation plan

Convert a small feature request into reviewable changes and acceptance tests.

Using this feature request and the supplied component code, propose an implementation plan that preserves the current interface. Name the files and state transitions involved, list accessibility and error cases, and define acceptance tests. Separate requirements from assumptions and flag missing API contracts. Do not write a broad redesign or claim that any code has been tested.

Workflow 02

Answer questions from a policy packet

Keep answers grounded in the supplied documents and make missing evidence visible.

Answer the five questions below using only this numbered policy packet. For each answer, cite the document and section, state the applicable conditions, and flag any conflicting wording. If the packet does not answer a question, say what information is missing. End with questions for the policy owner. Do not infer permissions or give legal conclusions beyond the supplied text.

Workflow 03

Review a dashboard screenshot

Turn visual observations into a limited, testable improvement list.

Review this dashboard screenshot for readability and task clarity. Identify ambiguous labels, hard-to-compare values, and potential mobile-layout issues with specific visible evidence. Suggest three small changes using the current design system. Separate observations from assumptions about interaction, and give a keyboard or responsive test for each proposed change.

Developer reference

Anthropic API pricing

These are Anthropic API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$3.00
5-minute cache write$3.75
1-hour cache write$6.00
Cache read$0.30
Output$15.00
  • Standard rates apply across the full 1M context window with no long-context premium.
  • US-only inference_geo processing adds 10% where offered; partner regional and multi-region pricing has separate terms.
  • Thinking tokens are billed as output, even when only a summary or no thinking text is displayed. Budget for the complete output usage, not just the visible answer.
  • Batch processing discounts input and output by 50%. Cache writes and reads have separate rates and eligibility; tools can add fees.
  • Prices and platform availability can change. Confirm the current provider, region, processing tier, and cache behavior before estimating direct API spend.

Common questions

A few things worth knowing.

Why consider Sonnet 4.6 rather than Opus 4.6?

It has the same documented context and synchronous output ceilings at lower standard input and output rates. That does not guarantee equal quality for your workload. Compare grounded accuracy, latency, and total tokens on representative tasks.

Is Sonnet 4.6 deprecated?

The reviewed overview marks it legacy, while the lifecycle table lists it active. There is no assigned shutdown date. The not-sooner-than February 17, 2027 commitment is not a retirement announcement.

Does high effort enable thinking?

No. Thinking is off unless explicitly configured. High is the effort default; to enable adaptive thinking, set thinking.type: adaptive in a supported API integration. App controls may not expose those fields.

Does Sonnet 4.6 have xhigh effort?

No. Its documented levels are low, medium, high, and max. Do not copy the five-level configuration of a newer Sonnet or Opus model without checking compatibility.

Do I have to remove manual thinking immediately?

Existing manual extended-thinking requests still succeed, but that mode is deprecated on Sonnet 4.6. Prefer adaptive mode when updating the integration and measure the behavioral difference rather than mapping a fixed budget to effort mechanically.

How should I use the 1M context window?

Organize relevant inputs with stable document names and section identifiers. Ask for source-backed answers and verify citations. The separate output ceiling and thinking budget still matter, and irrelevant bulk can make review harder.

Do the API rates include my EZ Ai Assist plan?

No. The table describes direct Anthropic usage, including cache categories. EZ Ai Assist pricing and model access are separate. The prompts are editorial examples you should adapt and validate, not guaranteed product capabilities.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider