The GenAI Field Guide · Trending

Claude Opus 5.5 and Sonnet 5.5

Anthropic's new mid-cycle Claude models: Opus 5.5 at a lower price than Opus 5, and Sonnet 5.5 at the old Sonnet price, both with thinking you steer rather than switch off.

Anthropic released Claude Opus 5.5 on 22 September 2026 and Claude Sonnet 5.5 on 28 September 2026. Opus 5.5 costs $4 and $20 per million input and output tokens, down from $5 and $25 on Opus 5, and Anthropic says it performs at the level of its larger Fable 5.1 on most work. Both are drop-in model IDs on the Claude API and the three big clouds, but they are not drop-in code: thinking is always on, forced tool use now returns an error, and the default effort changed, so anyone running Opus 5 or Sonnet 5 in production should read the breaking changes before switching.

What it is

The 5.5 family is a point release on the Claude 5 generation. Opus 5.5 is positioned in Anthropic's docs as the model for long-running agentic coding and knowledge work, and the models overview now tells developers to start with it for most workloads, keeping Fable 5.1 ($10 and $50 per million tokens) for the most demanding reasoning and long-horizon agent work. Sonnet 5.5 is described as the best combination of speed and intelligence, a faster and cheaper complement to Opus 5.5 for well-scoped everyday tasks. A Haiku 5.5 was announced alongside both but had not shipped as of 2 October 2026; the models overview still lists Haiku 4.5 as the fastest current model.

The headline claim is cost, not a new capability tier. Anthropic says Opus 5.5 is 20% cheaper per token than Opus 5, with cache reads 60% cheaper, and costs about 40% less to run on typical workloads because it also finishes tasks with fewer tokens. It also claims output is more than 30% faster than Opus 5. Sonnet 5.5 keeps Sonnet 5's price of $2 and $10 per million tokens, and Anthropic says it runs more than 30% faster and costs up to 30% less for most work. Coverage differs on how to state the saving: 9to5Mac reports a 20% overall decrease, which matches the per-token cut, while Anthropic's 40% figure is per task.

On Anthropic's own benchmarks, Opus 5.5 scores 66.4% on Terminal-Bench 4.0 against 55.8% for Fable 5.1 and 52.3% for Opus 5, and 1846 Elo on GDPval-AA v2.1 against 1735 and 1708. Sonnet 5.5 is reported at 70.6% on Terminal-Bench 4.0, higher than Opus 5.5's figure, 1844 on GDPval-AA and 80.1% partial credit on OSWorld 2.1 against 81.8% for Opus 5.5. All of these are vendor-run numbers and Anthropic's pages do not say whether the two models were run under the same effort settings, so treat the Sonnet over Opus coding result as a reason to test both, not as a ranking.

How it works

Both models use adaptive thinking: the model decides how much to reason before and between actions, and you steer that with the effort parameter (low, medium, high, xhigh, max) rather than a token budget. On Opus 5.5 thinking is always on; sending thinking type disabled, or a manual budget_tokens setting, returns a 400 error. Sonnet 5.5 replaces disabled with a new lowest setting, between_tools, which turns off up-front thinking but still lets the model write short progress notes between tool calls; it only works at high effort or below. Opus 5.5 defaults to medium effort where Opus 5 defaulted to high, so a request that never set effort now runs one level lower, and Anthropic says Opus 5.5 also thinks more per turn at a given level. Sonnet 5.5 keeps high as its default but its levels are recalibrated. The docs tell you to re-run an effort sweep rather than carry settings over.

Several API behaviours changed in ways that break existing code. Forced tool use (tool_choice any or a named tool) returns an error on both models; the documented replacement is auto plus strict tool use, or structured outputs. Thinking blocks are now tied to the model that produced them and to the conversation: if the system prompt, tools or an earlier message changes after a block was produced, replaying it returns a 400 error by default for accounts created on or after 31 August 2026, so conversations should be append-only. On the Claude API and Google Cloud the older computer_20251124 tool is rejected in favour of the computer_toolset_20260801 toolset. The text a model writes between tool calls now comes back inside thinking blocks that are empty at the default display setting, so a UI that streamed those notes goes silent with no error.

Safety controls show up as API behaviour too. A declined request returns HTTP 200 with stop_reason refusal and a stop_details category such as cyber, bio, frontier_llm or reasoning_extraction. Anthropic's announcement says most cybersecurity tasks on Opus 5.5 are re-routed to Opus 4.8 and that it uses Fable 5.1's biology safeguards, with a Life Sciences Verification Program for vetted organisations; for Sonnet 5.5 it says higher-risk security tasks visibly fall back to Sonnet 5. The docs describe an opt-in server-side fallback (in beta) that retries some declined categories on another model. Both models have a 1M token context window, 128K output tokens (300K on the Batch API with a beta header), text and image input, and a June 2026 knowledge cutoff, per the model pages.

How to use it

Both models are generally available to every Claude API customer and on Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. You need an Anthropic API key or a cloud account with Claude enabled. In the Claude apps, Opus 5.5 is on paid plans; Sonnet 5.5 is reported to be the model on the free plan.

  1. Change the model ID to claude-opus-5-5 or claude-sonnet-5-5 (anthropic.claude-opus-5-5 or anthropic.claude-sonnet-5-5 on Bedrock).
  2. Remove any thinking disabled or budget_tokens settings. On Opus 5.5, control cost with output_config.effort; on Sonnet 5.5, use between_tools if you need the old no-thinking behaviour.
  3. Replace tool_choice any or tool with auto plus strict tool use, and tell the model in the prompt when a tool applies.
  4. Make conversation history append-only. Change instructions or tools with mid-conversation system messages instead of editing earlier turns.
  5. Select response content blocks by type, not position, and pass thinking blocks back unchanged in tool loops. If your UI shows progress notes, set thinking.display so they are returned.
  6. Handle stop_reason refusal explicitly, and decide whether to enable server-side fallback or your own retry on another model.
  7. Set effort explicitly and run your own eval set at two or three effort levels before cutting over, comparing quality, latency and cost per task against the model you use today.

Use cases

Sources

  1. Claude Opus 5.5, Anthropic, 2026-09-22
  2. Claude Sonnet 5.5, Anthropic, 2026-09-28
  3. Claude Opus 5.5 model overview, Claude Platform docs, 2026-09-22
  4. What's new in Claude Opus 5.5, Claude Platform docs, 2026-09-22
  5. Claude Sonnet 5.5 model overview, Claude Platform docs, 2026-09-28
  6. What's new in Claude Sonnet 5.5, Claude Platform docs, 2026-09-28
  7. Models overview, Claude Platform docs, 2026-10
  8. Effort, Claude Platform docs, 2026-10
  9. Fast mode (research preview), Claude Platform docs, 2026-10
  10. Anthropic releases Opus 5.5 with lower prices and Fable-level performance, TechCrunch, 2026-09-22
  11. Anthropic upgrades Claude with new Opus 5.5 model, 9to5Mac, 2026-09-22
  12. Anthropic launches Opus 5.5, its first model since CEO Amodei called for AI slowdown, Yahoo Finance, 2026-09-22
  13. Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war, Simon Willison, 2026-09-22
  14. Claude Sonnet 5.5, Simon Willison, 2026-09-28
  15. Claude Sonnet 5.5 is available on the free plan, unlike Opus 5.5, Notebookcheck, 2026-09