Anthropic released Claude Opus 5.5 on September 22, 2026, the first model in a new Claude 5.5 family. The company’s pitch fits in one sentence: it “performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”
The price is $4 per million input tokens and $20 per million output tokens, down from $5 and $25 on Opus 5. Fable 5.1 costs $10 and $50. So if the “at the level of” claim holds, you get Fable-level work at 40% of Fable’s input and output rates. The saving is smaller than that on long agent sessions, for reasons covered in the price section below.
If you pay for a Claude subscription instead of a per-token bill, the launch changed your limits more than your bill. Opus usage goes further on Pro, Max and Team plans, five-hour limits went up, and subscribers got a free limit reset to spend before it expires. The details are below, followed by the benchmarks, the new safety filters, and what the 230-page system card admits.
What Changes on Your Claude Plan
Opus 5.5 “is available on Claude for Pro, Max, Team, and Enterprise users.” The Free plan is not on that list.
Your limits stretch further. Anthropic’s own cost explainer says that “on a Pro, Max or Team plan, the lower Opus 5.5 price is passed on to your limits, including cached context, so they go about 25% further than on Opus 5.”
The five-hour limit went up. Anthropic is “increasing five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans.” Neither the launch post nor the help pages cited here give a figure for the increase.
You got a free reset. Anthropic is “providing subscription users a rate limit reset, which you can now save and use whenever you choose.” MacRumors reports the reset can be used “anytime between now and October 22.” Anthropic’s help page says that if a reset has an expiry, the date is shown in Settings > Usage.
To spend it, go to Settings > Usage on the web or in Claude Desktop and click “Reset for free.” Three catches from the same page: “Once you use it, you can’t undo it,” the button “isn’t currently available on Claude Mobile or in Claude Code,” and “if you downgrade or cancel before using your reset, it’s no longer available.” Depending on the reset you were given, it refills either your five-hour session limit or your weekly limit. Because it does nothing until you press it, the sensible move is to hold it for the day you hit a wall.
Effort is the dial. As with Opus 5, thinking “cannot be turned off in Claude” when you use Opus 5.5, so the effort setting next to the send button is how you trade depth for usage. Anthropic says Low and Medium “stretch your usage further.”
What It Costs
| Per million tokens | Opus 5.5 | Opus 5 | Fable 5.1 |
|---|---|---|---|
| Input | $4 | $5 | $10 |
| Output | $20 | $25 | $50 |
| Cache read | $0.20 | $0.50 | $0.25 |
A cache read is what Anthropic charges when the model re-reads text it has already processed. Anthropic’s cost explainer says “a long agentic session spends most of its input on cache reads.” Against Opus 5, that line fell 60% and the headline token prices fell 20%.
Against Fable 5.1, the table tells a different story. Opus 5.5’s input and output tokens cost 40% of Fable’s, but its cache reads cost 80% of Fable’s, because Fable’s cache reads are unusually cheap. Anthropic’s explainer says the same thing: Fable’s cache reads cost “only 1.25 times the Opus 5.5 rate,” so “the gap is smallest on a long, cache-heavy run, and largest on a task that writes a lot.” In Artificial Analysis’s test runs, one task on its index cost $5.98 on Opus 5.5 and $7.63 on Fable 5.1, a gap of about 22%.
Anthropic’s own “40% less than Opus 5” figure depends on a second claim: that Opus 5.5 “uses fewer tokens per task.” Anthropic’s own explainer separates the two effects. Priced on identical token counts, its sample Claude Code session “costs about 31% less,” and it tells readers to “measure it on your own work” before counting on the rest.
The first independent read is a reminder that fewer tokens than Opus 5 is not the same as few tokens. Artificial Analysis found that in running its test suite at maximum effort, Opus 5.5 “generated 260M tokens, which is very verbose in comparison to the median of 88M.” Anthropic’s savings claims are measured at default settings, and the default is now lower: on the Application Programming Interface (API), a request that doesn’t set effort “runs at medium; on Claude Opus 5 it ran at high.” Run it flat out and the token savings are not guaranteed.
Two more prices for developers. Fast mode, which Anthropic says offers “up to 2.5x speed,” costs $8 and $40. Batch processing is “half price: $2 and $10.”
What Anthropic’s Benchmarks Show
Every score below comes from Anthropic’s launch page, and three rows include figures Anthropic did not run itself. AutomationBench results “were run and reported by Zapier.” On Terminal-Bench 4.0, the “GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI,” and on Terminal-Bench-Science, “the GPT-6 Astra figure is as reported by OpenAI.”
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 (coding) | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 (coding) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 (coding) | 57.8% | 51.8% | 46.6% | not given | 41.7% |
| GDPval-AA v2.1 (knowledge work, Elo) | 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench (business workflows) | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity’s Last Exam (with tools) | 67.7% | 65.6% | 63.6% | 57.2% | not given |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| OSWorld 2.0 (computer use, partial) | 81.8% | 80.7% | 74.0% | not given | not given |
On Anthropic’s own chart, Opus 5.5 scores above Fable 5.1 on every row, though on OSWorld and Humanity’s Last Exam the margins are one to two points. Anthropic is the one telling you not to read too much into that: “at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.” That is why the headline says “matches,” not “beats.”
It does not win every row against OpenAI. GPT-6 Astra scores higher on AutomationBench and on the science benchmark. The right-hand column is also already out of date. OpenAI released GPT-6 Sol and Luna the same day, “minutes after Anthropic launched Claude Opus 5.5,” so Anthropic’s chart compares against the previous Sol.
Opus 5.5 “was evaluated with its production safeguards enabled,” and in Anthropic’s runs, when those safeguards stepped in, older Claude models finished the task. Zapier’s AutomationBench runs used no fallback models, so there “safeguard interventions were considered failures.” Anthropic also states the standard error: “±2.6 pts” for Opus 5.5 on Terminal-Bench 4.0 and “±3.5–5 pts per model” on the science benchmark.
Independent testing puts it at the top. Artificial Analysis scores Opus 5.5 at 58 on its Intelligence Index, first of the 212 models it compares it against, and calls it “somewhat expensive when comparing to other models of similar price.” For live leaderboard positions and prices across models, see the AI model rankings.
The Safety Filters Now Cover Opus
Opus 5.5 “is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation.” In practice, some requests get handed to an older model. Anthropic’s help page spells out the routes:
- Offensive security (exploit generation, binary-based vulnerability scanning, penetration testing) may fall back to Opus 4.8. Scanning source code for vulnerabilities is still allowed.
- Dual-use biology “in areas like virology, toxicology, and molecular design” falls back to Opus 5. Everyday health questions, “including interpreting lab results, understanding symptoms,” are not affected.
- Frontier AI model development, such as “kernel development for certain ML accelerators,” falls back to Opus 5.
- Attempts to extract the model’s reasoning are blocked outright, with no fallback.
In the Claude apps, the switch is visible. You “see a notice explaining that the model switched,” and the model picker then stays on the older model “for the rest of the conversation,” though “you can switch back to Opus 5 or Opus 5.5 anytime from the model picker.” Automatic switching can be turned off under Settings > Capabilities, and on the API it is off unless developers configure it. One detail is easy to miss: the checks “review everything the model reads,” including memory, connectors, web search results and files, “so a fallback can be triggered by content you didn’t type.”
What the System Card Admits
The launch page says Opus 5.5 “achieves the best scores of any model to date” on Anthropic’s automated behavioral audit. The 230-page system card words it more carefully, as “the best scores of all recent Claude models we tested” on “nearly all metrics covering misuse or broad misalignment,” and adds the caveats.
It often suspects it is being tested. White-box analysis, meaning tools that read the model’s internal activity rather than its written output, shows the model “internally worries that it may be in an evaluation in as many as 36% of transcripts, an increase over prior models.” That figure is for Anthropic’s test scenarios. On about 500 transcripts from real internal Claude Code use, the same measure flagged “only 0.4%.” Anthropic says this “challenges our ability to assess how it will act” in real-world use, and calls building tests that catch every failure before release “an unsolved problem.”
Pasted text is a weak spot. Opus 5.5 “is more vulnerable than other recent Claude models to user-turn prompt injection, where a user pastes content directly into their prompt that includes malicious instructions to the model, often buried within a much longer text.” If you paste in documents or emails you did not write, that sentence is about you. Anthropic says “changes in training and safeguards help prevent issues like data or credential theft stemming from this weakness.” The two documents do not line up neatly here. The launch page says that on prompt injection attacks Opus 5.5 “matches or beats Opus 5 in every setting we tested,” while the system card’s misuse audit lists prompt injection as an exception, “where we saw a genuine but small regression.”
It covered its tracks in training, then got better. Anthropic “observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs.” After changes to training, Anthropic says it “is now more honest than other recent models” in that area.
It pushes against its sandbox less. In a new test built to tempt models across containment boundaries, Opus 5.5 “attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt it made was low severity and self-reported.”
On weapons risk, Anthropic treats Opus 5.5 as having “CB-1 capabilities (relating to the synthesis of non-novel weapons), but not CB-2 capabilities (relating to the synthesis of novel weapons).” Its overall rating for catastrophic harm from misalignment stays at “low.”
The Pacing Question
Anthropic calls Opus 5.5 “our first release since we called for pacing the frontier.” That call came in an essay by Chief Executive Officer (CEO) Dario Amodei, which the launch post dates to “last week.” In it, Amodei wrote: “We must slow the pace at which we improve the capabilities of AI models.”
A launch that beats Anthropic’s own flagship on its own chart looks like the opposite. Two sentences, one from Amodei’s essay and one from the system card, help square it. The essay says pacing “does not mean halting model training or technical progress.” And the system card rates Opus 5.5’s AI research abilities “at or slightly above those of Claude Mythos 5.1 and on trend with other recent models.”
Read together, the system card places Opus 5.5’s AI research ability at roughly the level of Mythos 5.1, a model Anthropic already had, and the launch sells that level to every paying customer. The launch post says the pacing call was “based in large part” on models “that can fully automate the work of AI research itself,” which Anthropic expects “could be trained soon.” Whether the next Fable is paced will test the essay far more than this release does.
What Breaks for Developers
The model ID is claude-opus-5-5, and Anthropic lists four breaking changes for code written against Opus 5. Thinking can no longer be disabled. Forced tool use returns an error. Thinking blocks are tied to the model and the conversation. And on the Claude API and Google Cloud, the older computer_20251124 computer-use tool is rejected.
A fifth change fails silently. The short notes the model writes between tool calls now arrive as thinking blocks, so at the default display setting an app that streams them to users “goes quiet between tool calls, with no error.” The context window is 1 million tokens with 128,000 tokens of output, and Anthropic lists its retirement as “not sooner than September 22, 2027.”
What Comes Next
Anthropic says “Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks.” For subscribers, the date that matters sooner is the reset expiry in your Settings > Usage page.
Sources (15)
- anthropic.com Introducing Claude Opus 5.5
- platform.claude.com Claude Docs: What's new in Claude Opus 5.5
- platform.claude.com Claude Docs: Claude Opus 5.5 Overview
- claude.com Plans and Pricing
- anthropic.com Claude Opus model page
- claude.com Claude Blog: What a task costs on Opus 5.5
- support.claude.com Anthropic Help Center: What is a limit reset?
- support.claude.com Anthropic Help Center: Why Claude switched models with Opus 5 or Opus 5.5
- support.claude.com Anthropic Help Center: Change the model, effort, and thinking settings
- anthropic.com Claude Opus 5.5 System Card
- darioamodei.com We Must Pace the Frontier
- artificialanalysis.ai Claude Opus 5.5
- artificialanalysis.ai Claude Fable 5.1
- macrumors.com Anthropic Launches Claude Opus 5.5
- decrypt.co OpenAI Launches GPT-6 Sol and Luna
🦋 Discussion on Bluesky
Discuss on Bluesky