AI

Claude Sonnet 5.5 Is Here: Anthropic's Everyday AI Model Gets Much Closer to Opus

Hasaka · SEP 28, 2026

Summary

Released September 28, 2026, Claude Sonnet 5.5 keeps Sonnet 5's API pricing ($2/$10 per million tokens) while running over 30% faster and up to 30% cheaper per task through fewer tokens and tool calls. It nearly matches Opus 5.5 on coding benchmarks, outscoring it on Terminal-Bench 4.0, while Opus remains stronger for ambiguous, open-ended work. Sonnet 5.5 also gains cybersecurity safeguards previously reserved for Anthropic's flagship models.

Anthropic has released Claude Sonnet 5.5, promising faster coding, stronger knowledge work and significantly better efficiency without increasing API prices. More interestingly, the new Sonnet is now approaching Claude Opus 5.5 on several evaluations, potentially changing which model developers should reach for first.

Anthropic released Claude Sonnet 5.5 on September 28, 2026, making it the second member of the Claude 5.5 family following Opus 5.5.

Sonnet has traditionally occupied an important position in Anthropic's lineup. It is not intended simply to be the company's most powerful model. Instead, it balances capability, speed and operating cost for the kind of AI work developers and businesses perform repeatedly.

With Sonnet 5.5, that balance has shifted considerably.

Anthropic says the model generates output more than 30% faster than Sonnet 5, can cost up to 30% less per task, and substantially improves coding, computer use and professional knowledge work. Its API token prices remain unchanged.

That combination may make Sonnet 5.5 one of Anthropic's more consequential releases for everyday AI workflows.

Sonnet 5.5 is about efficiency, not just intelligence

AI model launches are usually dominated by benchmark scores.

Anthropic is placing unusual emphasis on something arguably more useful, how much work the model can complete for the money spent.

Sonnet 5.5 retains Sonnet 5's API pricing at $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 per million tokens. But Anthropic says the new model generally requires fewer tokens and fewer tool calls to accomplish the same work, producing savings of up to 30% per task in its testing.

This distinction matters. A model with identical token pricing can still be considerably cheaper to operate if it reaches the answer in fewer steps. For developers running AI agents continuously across repositories, support systems or business processes, cost per completed task is ultimately more important than cost per token.

Claude Sonnet 5.5 vs Sonnet 5 vs Opus 5.5

The improvement becomes clearer when the three models are placed side by side.

Sonnet 5.5Sonnet 5Opus 5.5
Input / 1M tokens$2$2$4
Output / 1M tokens$10$10$20
Cache reads / 1M$0.20$0.20$0.20
Output speed30%+ faster than Sonnet 5BaselineHigher-tier model
Terminal-Bench 4.070.6%10.3%66.4%
FrontierCode 1.152.1% at Xhigh42.4%54.4%
CursorBench 4.055.5%34.1%57.8%
Humanity's Last Exam64.5%54.9%67.7%
OSWorld 2.180.1%57.0%81.8%
Best suited toEveryday coding, agents, business workflowsPrevious generationComplex, open-ended work

These are Anthropic's published evaluations and should not be interpreted as universal measures of real-world performance. Anthropic itself cautions that Opus 5.5 remains clearly stronger for complex, open-ended work requiring sustained judgment.

But the table reveals something important. The gap between Sonnet and Opus is getting remarkably small in some areas.

Coding is where the upgrade becomes obvious

Coding appears to be one of Sonnet 5.5's biggest improvements.

On CursorBench 4.0, which evaluates ambiguous multi-file coding tasks derived from real Cursor sessions, Sonnet 5.5 scores 55.5%, compared with 34.1% for Sonnet 5 and 57.8% for Opus 5.5.

On Terminal-Bench 4.0, Sonnet 5.5 reaches 70.6% in Anthropic's evaluation, compared with 66.4% for Opus 5.5. Benchmark configurations and effort levels matter, so this does not mean Sonnet is universally more capable than Opus.

Anthropic reports another important behavioural improvement, Sonnet 5.5 tends to batch tool calls together rather than repeatedly moving between reasoning and execution. That means fewer steps, fewer tokens and faster completion.

On FrontierCode, Anthropic says Sonnet 5.5 at High effort scored around 10 points above Sonnet 5 at the same setting while costing approximately one-fifteenth as much per task.

This is exactly the kind of improvement that matters when AI moves from occasionally generating code snippets to operating as an actual development agent.

It can stay with larger projects

Another important change is how the model handles sustained work.

Early testers reported Sonnet 5.5 working across tens of thousands of lines of code and continuing through multi-hour tasks. Epic Games said its testing included system design and data-flow reviews across large gameplay architectures, while Base44 tested the model across 118 real application builds. Base44 reported that Sonnet 5.5 reached its target in an average of 3.6 iterations, compared with 7.7 for Opus 5, with the fewest failed tool calls of any model the company had compared.

That points toward a larger transition happening across AI development tools. The important question is becoming less, can AI write code? And increasingly, can AI understand an existing product, work across it for hours, use tools correctly and actually finish the task?

Sonnet 5.5 appears designed around that second problem.

Sonnet is also becoming more interesting for design

For Darwin Corp, one of the more interesting details in Anthropic's announcement has little to do with traditional coding benchmarks.

Early testers specifically highlighted improvements in design and interface work. Anthropic says testers found Sonnet 5.5 better at adding polish to user interfaces and following visual templates. In one internal test, Anthropic provided the model with earnings materials, transcripts and an existing presentation template, then asked it to create a ten-slide operating review. Two experts judged its first draft ready to send without further editing.

That matters because the boundary between AI coding and AI design is disappearing. Modern frontend development requires more than technically valid React components. A model needs to understand spacing, hierarchy, typography, responsive behaviour, interaction states and the overall visual system surrounding the component.

For creative development teams, improvements in this area may ultimately be more valuable than another few percentage points on a programming benchmark.

Claude Sonnet 5.5 vs Opus 5.5: which should you use?

The release makes this question considerably more interesting.

Sonnet 5.5 costs half as much per input and output token as Opus 5.5. At lower and medium effort settings, Anthropic says it can also achieve substantially lower cost per task.

For everyday software development, debugging, feature implementation, content operations and well-defined agent workflows, Sonnet 5.5 therefore becomes the logical starting point.

Opus 5.5 still has an important advantage when the task becomes ambiguous. Anthropic says its flagship remains stronger at complex, open-ended work requiring sustained judgment. That could include architecture decisions, difficult research, large strategic problems or situations where determining what should be done is harder than executing it.

That suggests an increasingly useful two-model workflow, Opus determines the direction, Sonnet executes the work. For example, Opus might analyse a large application, establish the architecture and determine how a complicated migration should work. Sonnet could then implement that plan across the codebase. One early tester described essentially this exact workflow for game development.

Effort levels make the decision more flexible

Sonnet 5.5 also supports adjustable effort.

Claude's apps and Claude Code use Medium effort by default, while the Claude Platform defaults to High. Lower settings prioritise speed and reduced token consumption. Higher settings allow Claude to reason longer and check its work more thoroughly.

This creates another change in how we should think about AI models. Choosing an AI system is no longer simply choosing between a fast model and a smart model. The same model can increasingly move along that spectrum depending on the task. A simple UI adjustment may need very little reasoning. Debugging an unusual production problem might need considerably more.

Instead of paying for maximum intelligence every time, developers can allocate more compute only when the problem requires it.

Safety is getting stronger too

Sonnet 5.5 also introduces stronger safeguards as its technical capabilities increase.

Anthropic evaluated the model across roughly 1,850 behavioural scenarios and reports that it improves on or matches Sonnet 5 across most measures involving alignment, misuse resistance and honesty. The company says it found no evidence in those evaluations that Sonnet 5.5 systematically pursued objectives conflicting with user intentions, while acknowledging that evaluations cannot detect every possible failure mode.

Cybersecurity capabilities have increased enough that Sonnet 5.5 becomes the first Sonnet model to launch with safeguards similar to those used for Anthropic's highest-capability models. Routine software development remains available, while certain higher-risk cybersecurity requests can fall back to Sonnet 5.

It is another indication of how quickly mid-tier AI models are approaching capabilities that were recently reserved for flagship systems.

What this means for creative development

For agencies and digital product teams, Sonnet 5.5 may represent a more meaningful upgrade than a dramatically more expensive frontier model.

The reason is simple. Most production work does not require the smartest available AI model every second. A typical website project contains hundreds of smaller tasks, building components, checking responsiveness, fixing bugs, refactoring code, analysing accessibility, generating structured content, testing interactions and connecting APIs.

What matters is having a model capable enough to perform those tasks reliably, but efficient enough to use continuously. Sonnet 5.5 moves closer to that balance.

For teams building custom interactive websites, AI does not remove the need for designers, developers or creative direction. It changes where their time is spent. Less time can go into repetitive implementation. More can go into deciding what the experience should actually feel like.

The middle model is becoming difficult to call "mid-tier"

That may ultimately be the most interesting part of Claude Sonnet 5.5.

AI companies traditionally create clear product hierarchies. The smaller model is fast. The middle model balances everything. The flagship is where the serious intelligence lives. Those boundaries are becoming less obvious.

Sonnet 5.5 reaches close to Opus 5.5 across several of Anthropic's evaluations while costing half as much per standard input and output token. It is faster than its predecessor, consumes fewer resources per task and appears considerably more capable of sustained agentic work.

Opus still has a reason to exist. But Sonnet increasingly looks capable of handling the work that most developers actually need to do every day.

And that may be more important than winning the title of the world's most powerful AI model.

The AI model that changes how people work may not be the smartest one available. It may be the one that becomes capable, fast and affordable enough to use for almost everything.

Claude Sonnet 5.5 official announcement

Frequently asked questions

Claude Sonnet 5.5 is Anthropic's second Claude 5.5 family release, launched September 28, 2026, following Opus 5.5. It's designed to balance capability, speed and cost for everyday coding, agentic and business workflows, keeping Sonnet 5's pricing while running more than 30% faster and up to 30% cheaper per task.
No, API token pricing is unchanged at $2 per million input tokens and $10 per million output tokens, with cache reads at $0.20 per million tokens. Anthropic says efficiency gains, fewer tokens and tool calls per task, can still cut typical costs by up to 30%.
Sonnet 5.5 costs half as much per token as Opus 5.5 and nearly matches it on several benchmarks, even outscoring it on Terminal-Bench 4.0 (70.6% versus 66.4%). Opus 5.5 remains stronger for complex, open-ended work requiring sustained judgment, while Sonnet 5.5 is positioned as the everyday workhorse.
Yes. Anthropic reports significant coding gains, including 55.5% on CursorBench 4.0 versus 34.1% for Sonnet 5, and notes the model batches tool calls more efficiently, completing tasks in fewer steps. Base44 reported Sonnet 5.5 averaged 3.6 iterations per app build versus 7.7 for Opus 5 across 118 real builds.
Anthropic evaluated Sonnet 5.5 across roughly 1,850 behavioural scenarios, reporting it matches or improves on Sonnet 5 on alignment, misuse resistance and honesty measures. It's also the first Sonnet model to launch with cybersecurity safeguards similar to Anthropic's highest-capability models, with certain high-risk requests falling back to Sonnet 5.