Back to models
LegacyChat Models

Anthropic

Claude Opus 4.8

Best for Research

Anthropic’s active legacy Opus model with optional adaptive thinking, image input, and a 1M-token context.

1M contextAgentic

At a glance

Know the model before you prompt.

Anthropic API specifications
Context window
1,000,000 tokens
Maximum output
128,000 tokens
Inputs → output
Text + Images → Text
Knowledge cutoff
January 2026

API model ID: claude-opus-4-8

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Text and image inputs with text output
  • Opt-in adaptive thinking with five effort levels
  • Reasoning between tool calls when adaptive thinking is enabled
  • A 1,000,000-token context for supplied material
  • Prompt caching and discounted Message Batches
  • Optional Claude API Fast mode under separate preview terms

Before you choose

  • Legacy does not mean retired: no scheduled shutdown is listed in the reviewed lifecycle page.
  • Manual thinking.type: enabled with budget_tokens is rejected; use adaptive thinking instead.
  • Native output is text, not generated audio or video. Tools and external actions require an integration and permissions.
  • Non-default temperature, top_p, and top_k values are rejected. Re-evaluate token usage after migrating from pre-4.7 tokenizers.

Adaptive thinking and effort

lowmediumhigh · defaultxhighmax

Thinking is off by default. Set thinking.type: adaptive to enable it; high is the effort default, not an always-on thinking setting. Anthropic recommends starting with xhigh for coding and agentic work, then measuring quality and token use. Keep sufficient max_tokens for thinking and the final response. Manual budget_tokens and between_tools are not supported; thinking can be disabled.

  • The synchronous Messages API output ceiling is 128,000 tokens. Up to 300,000 output tokens is a Message Batches beta requiring output-300k-2026-03-24, not a normal chat or app limit.
  • Anthropic lists retirement not sooner than May 28, 2027. This is a support commitment, not a scheduled retirement.
  • The tokenizer introduced in Opus 4.7 can use about 30% more tokens for the same text than earlier models. Count actual inputs when comparing cost and capacity.
  • With adaptive thinking, reasoning can interleave with tool use without the older interleaved-thinking beta header. Tools still require your integration.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Trace a repository dependency change

Ask for an evidence-backed impact analysis before modifying a widely used interface.

Trace the proposed interface change through the repository excerpts below. Identify direct callers, assumptions at each boundary, and likely compatibility risks with file references. Rank the changes that need tests, then propose a staged migration with rollback criteria. Separate observed dependencies from files we have not supplied. Do not edit files, execute tools, or assume the complete repository is present.

Workflow 02

Reconcile a long technical packet

Organize a large input around claims, dates, and a decision rather than requesting a generic summary.

Using the technical packet below, compare the three proposed architectures against our stated constraints. Build a claim-to-source table, identify contradictions, and distinguish required capabilities from preferences. Recommend a small validation experiment for the largest uncertainty. End with a decision brief that cites sections and states what would change the recommendation. Do not add external facts.

Workflow 03

Inspect a multi-screen workflow

Use ordered screenshots to distinguish visible UI evidence from untested interactions.

Review these numbered screenshots of a multi-step workflow. For each transition, describe the visible user goal, any unclear control, and the evidence for a potential failure. Prioritize three improvements that preserve the existing design. Specify an interactive test for each issue and mark anything that cannot be concluded from images alone. Do not invent screens or claim the flow was executed.

Developer reference

Anthropic API pricing

These are Anthropic API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$5.00
5-minute cache write$6.25
1-hour cache write$10.00
Cache read$0.50
Output$25.00
  • Standard rates apply across the full 1M context window; there is no long-context premium.
  • Fast mode is a Claude API research preview at 2× standard rates: $10 input and $50 output per million tokens. It is not available with Batch or partner platforms.
  • US-only inference_geo processing adds 10% to token rates where offered; partner regional pricing has separate terms.
  • Thinking tokens are billed as output, even when only a summary or no thinking text is displayed. Budget for the complete output usage, not just the visible answer.
  • Batch processing discounts input and output by 50%. Cache writes and reads have separate rates and eligibility; tools can add fees.
  • Prices and platform availability can change. Confirm the current provider, region, processing tier, and cache behavior before estimating direct API spend.

Common questions

A few things worth knowing.

Is Opus 4.8 still supported?

Yes, the reviewed documentation describes it as an active legacy model. Legacy means it is no longer the current generation, not that it has stopped working. Anthropic’s stated earliest retirement date is a commitment rather than a scheduled shutdown.

Should I migrate to Opus 5.5?

Anthropic’s overview recommends considering Opus 5.5. Compare the same representative inputs, tools, and evaluation criteria. Migration can change thinking defaults, token counts, response behavior, and price; validate those differences before switching an established workload.

Is thinking enabled automatically?

No. Opus 4.8 requires thinking.type: adaptive. Its high effort default only describes the effort parameter; it does not enable thinking. The recommended xhigh starting point for coding is also different from the API default.

Does Fast mode make the model more capable?

It is a separately priced processing option aimed at response speed, not a different context window or a guarantee of better answers. The reviewed pricing lists a research preview on the Claude API, not Batch or partner platforms. App access to this option is not established here.

Can I get a 300,000-token answer in a chat?

The documented synchronous limit is 128,000 tokens. The larger output limit is a Message Batches beta with a specific header. It should not be presented as a normal interactive-chat or EZ Ai Assist limit.

Why might a migration use more input tokens?

Opus 4.8 uses the tokenizer introduced with Opus 4.7. Anthropic reports roughly 30% more tokens for the same text than earlier tokenizers. Use measured token counts and actual cache hits instead of comparing only the price per million.

Are these examples and prices part of my plan?

The prompts are original editorial starting points, not benchmarks or promised outcomes. The prices are direct Anthropic API references, separate from EZ Ai Assist subscriptions. Adapt prompts to your app’s available inputs and verify important results.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider