CompaniesInvestorsPeople
Home
Loading

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

Get in Touch

  • Contact

  • Request a Demo

  • Request Data Updates

  • Add a Company

Research

  • Companies

  • Investors

  • People

aVenture

  • Download App

  • Pricing

Download the aVenture Research beta for iOS and iPadOSDownload aVenture Research on the Mac App Store

Resources

  • Documentation

  • CLI

  • MCP

  • Feature Requests

  • Sitemap

Member

Backed by

© aVenture Investment Company, 2026. All rights reserved.

San Francisco, CA, USA

Privacy Policy · Terms of Service

aVenture Investment Company ("aVenture") is an independent research platform providing detailed analysis and data on startups, venture capital investments, and key industry individuals. It is not a registered investment adviser, broker-dealer, or investment advisor and does not provide investment advice or recommendations. The data provided by aVenture does not constitute recommendations or advice, whether by methodology, analysis, AI-generated content, or a statement written by a staff member of aVenture.

aVenture is not affiliated with any of the people, companies, organizations, government agencies, regulatory bodies, or investment funds we provide coverage for on this site unless explicitly stated otherwise. Users assume full responsibility for decisions made based on information obtained from this platform. Links to external websites do not imply endorsement or affiliation with aVenture. Any links that provide the ability to invest in a primary or secondary transaction in a company are for convenience only and do not constitute solicitations or offers to buy or sell an investment. Investors should exercise heightened precaution and due diligence when investing in private companies, especially those not independently audited.

While we strive to provide valuable insights with objectivity and professional diligence, we cannot guarantee the accuracy of the information provided on our platform. Before making any investment decisions, you should verify the accuracy of all pertinent details for your decision. To the fullest extent permitted by law, aVenture shall not be liable for any direct, indirect, incidental, consequential, or financial damages arising from use of this site, whether by consumers of its contents directly or by persons or organizations covered by our research, even if we are advised of the possibility. Our best-efforts processes and correction request forms do not create a warranty or duty of care.

Profiles on this platform may include content generated in part by large language models (LLMs, artificial intelligence) that aggregate publicly available sources (e.g., SEC EDGAR, public filings, press releases). Source attribution is provided where known; always verify statements and claims here against original sources before relying on any data. Content on our site may contain inaccuracies, omissions, or what are commonly called 'hallucinations' if generated in part or in full by AI / LLMs. The risk can also exist even when content is written by a human, as internal and third-party sources may also have inaccuracies for the same or different reasons. While we randomly audit a proportion of content, this is not exhaustive.

We recommend that an independent auditor be hired to verify the accuracy of the information before relying on it for any sensitive decisions. By accessing this platform, you agree not to rely solely on any information generated by AI, aggregated, or sourced or written otherwise on this site, for investment, financial, or other decisions. aVenture assumes no responsibility for inaccuracies, omissions, or hallucinations. You must independently verify all data from primary sources. Use of this platform constitutes your waiver of claims for reliance-based damages, including negligent misrepresentation. To report an error, request a correction, or dispute information about a company or individual, contact us via our request data updates form.

Loading
Loading
Home
News
Introducing Claude Sonnet 5.5

From Anthropic

By

September 28, 2026

Introducing Claude Sonnet 5.5

Introducing Claude Sonnet 5.5

Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.

Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5. Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It’s also got a sharp eye for design. Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.

Sonnet 5.5 improves over Sonnet 5 on:

Performance. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, compared to Sonnet 5’s 10.3%. It scores two points below Opus 5.5 on GDPval-AA, a test of real-world work across a variety of occupations. And it’s strong on long-horizon work and image understanding—it’s the first Sonnet model to beat Pokémon Red working only from screenshots.

Collaboration. Like Opus 5.5, Sonnet 5.5 writes more clearly than our previous generation of models; early testers described it as a better partner for collaboration than Sonnet 5. Its speed also makes it well suited to fast iteration on less complex tasks.

Cost. Sonnet 5.5 is priced the same as Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads, but it typically needs far fewer tokens to do the same work. In our testing, it costs up to 30% less per task than its predecessor.

Speed. Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.

Alignment and safety. On our automated behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment. Because its cybersecurity capabilities are comparable to Opus 5’s, it’s the first Sonnet model to launch with cyber safeguards and fallbacks like those we’ve developed for our most capable models. Its biology safeguards are the same as Sonnet 5’s. Both safeguards target a narrow set of high-risk requests; routine software development and most life sciences work are unaffected.

Performance

Sonnet 5.5 improves on Sonnet 5 across domains—in some cases dramatically. On several evaluations, Sonnet 5.5 at Max effort even performs comparably to Opus 5.5. However, benchmark scores capture only one facet of a model’s capabilities; in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.

Sonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Agentic codingTerminal-Bench 4.0
Agentic codingTerminal-Bench 4.070.6%10.3%66.4%¹—
Agentic codingFrontierCode 1.1 (Main)
Agentic codingFrontierCode 1.1 (Main)46.2%Max²42.4%54.4%49.3%
52.1%Xhigh
Agentic codingCursorBench 4.0
Agentic codingCursorBench 4.055.5%34.1%57.8%—
Knowledge workGDPval-AA v2.1³
Knowledge workGDPval-AA v2.1³1844144918461487⁴
Knowledge workAA-Briefcase v1.1³
Knowledge workAA-Briefcase v1.1³1811135918221483⁴
Multidisciplinary reasoningHumanity’s Last Exam
Multidisciplinary reasoningHumanity’s Last Exam64.5%with tools54.9%with tools67.7%with tools—
Computer useOSWorld 2.1
Computer useOSWorld 2.180.1%partial57.0%partial81.8%partial—
Visual chart recognitionChartography
Visual chart recognitionChartography61.6%no tools15.6%no tools64.4%no tools53.6%⁴no tools

The charts below plot each model’s score against its cost per task at every effort level. As effort goes up, models typically work for longer, leading to a higher cost per task but generally also a higher score. The closer a point is to the top left of the chart, the more capability it delivers per dollar.

On several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5’s best score for about a tenth of the cost per task. It complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost.

Agentic terminal codingAgentic coding: FrontierCodeAgentic coding: CursorBenchKnowledge work: AA-Briefcase

Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command-line interface. At Medium effort, the default in the Claude apps, Sonnet 5.5 far exceeds Sonnet 5’s best score for less than a tenth of the cost per task.

Terminal-Bench and OpenAI did not report GPT-6 Sol performance publicly, so we report GPT-5.6 Sol here.

FrontierCode measures whether an agent’s code changes would be merged. At High effort, the default on the Claude Platform, Sonnet 5.5 matches GPT-6 Sol’s best score for about a fifth of the cost per task.²

CursorBench evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions. Sonnet 5.5 at Low effort exceeds Sonnet 5’s best score for less than a tenth of the cost per task.

CursorBench 4.0 does not report GPT-6 Sol performance publicly, so we report GPT-5.6 Sol here.

On AA-Briefcase, a new benchmark of long-horizon knowledge work, Sonnet 5.5 at Medium effort bests Sonnet 5’s best score for about one ninth of the cost per task.³

Coding

Sonnet 5.5’s jump in performance is particularly noticeable in coding. At High effort on FrontierCode, it scores 10 points higher than Sonnet 5 at the same setting, at about one fifteenth of the cost per task. On CursorBench, which tests models on tasks from real Cursor coding sessions, its best score is within about two points of Opus 5.5.

Early testers appreciated how quickly Sonnet 5.5 can understand a codebase. They were also struck by its efficiency: in head-to-head runs, it batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs.

Epic GamesEveryCodeRabbitSpaceXAIBase44UnityCreator
Quote

“In Epic’s early testing, Claude Sonnet 5.5 cleared the same quality bar you’d expect from a higher-tier model, holding up on a system design audit and a data flow review. The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting.”

CompanyEpic Games
AuthorDaniel Vogel, Chief Operating Officer
Quote

“Claude Sonnet 5.5 cooks. Fast at coding and can be steered quickly in iterative workflows. But it can still work long if it needs to. It’s got some of Opus 5.5’s natural writing upgrades, which makes it more fun to work with.”

CompanyEvery
AuthorTyler Nishida, Designer
Quote

“Claude Sonnet 5.5 shows better judgment than Sonnet 5 across different levels of complexity, while spending significantly fewer output tokens. Sonnet 5’s tendency to reach for web search too often and its high token use are both gone in this new model. We plan to move simple and moderate reviews over now, and more in the coming weeks.”

CompanyCodeRabbit
AuthorDavid Loker, VP of AI
Quote

“Claude Sonnet 5.5 delivers frontier-level performance on CursorBench 4.0 at 55.5%, second only to Opus 5.5. We think it will be a hit with developers looking to balance performance with cost.”

CompanySpaceXAI
AuthorSualeh Asif, Director of ML
Quote

“Across 118 real app builds, Claude Sonnet 5.5 produced apps that scored level with Opus 5. It got there in 3.6 iterations per build on average, where Opus 5 took 7.7. It had the fewest failed tool calls of any model we compared. It also rarely stopped mid-build to ask the user a question, so fewer builds stall waiting on someone to answer.”

CompanyBase44
AuthorGabriel Grinberg, AI Engineering Lead
Quote

“At Unity, we have a high bar for task completion. Projects are reopened and results are checked at runtime, so a task only counts when the change works, not when the model says it’s done. The majority of Claude Sonnet 5.5’s work passed that check. It also completed 90% of tasks in our multi-step Unity Editor and coding benchmark, beating similar models.”

CompanyUnity
AuthorSam Zhang, Creative Technologist
Quote

“When Claude Opus 5.5 sets the architecture and general framework for a game, I would feel confident in letting Sonnet 5.5 implement it. I’m impressed with Sonnet 5.5’s handling of long-running, complex tasks.”

CompanyCreator
AuthorKevin Ngo, Creative Coder

Knowledge work

Sonnet 5.5 shows gains in multiple areas of knowledge work. On GDPval-AA, which tests models on real-world tasks across 44 occupations and nine major industries, Sonnet 5.5 scores nearly level with Opus 5.5 and about 400 points above Sonnet 5. It’s close to Opus 5.5 in computer use and chart recognition, and clearly outperforms Sonnet 5 and GPT-6 Sol on long-horizon knowledge work.

Early testers highlighted less quantifiable improvements. They found it to be a more natural conversational partner and remarked on its knack for design, noting that it adds polish to user interfaces and can follow slide templates to create decks that require minimal editing. In one internal test, we gave it a public company’s quarterly earnings materials and call transcripts, along with a slide template, and asked for a 10-slide operating review. Two experts judged its first draft to be ready to send as is.

SlackZendeskBalyasny Asset ManagementBoxLovableAtlassian
Quote

“Without changing any of our prompts, Claude Sonnet 5.5 did better than Sonnet 5 on almost all of our offline Slackbot evals, in fewer steps and with about 14% fewer output tokens. When someone gives Slackbot a task, quality and speed are what matter most, and Sonnet 5.5 allows Slackbot to deliver better outcomes for users, faster.”

CompanySlack
AuthorCurtis Allen, Principal Engineer
Quote

“We fed Claude Sonnet 5.5 hundreds of real support use cases across replies and escalation requests. It made fewer wrong decisions and resolved tickets faster than the Claude models we use in production today. Tickets were processed 20% faster, getting our customers the help they need without the wait.”

CompanyZendesk
AuthorAbhinay Kathuria, Director of AI
Quote

“On our private suite of 2,441 finance tasks covering Q&A, extraction, analysis, and forecasting, Claude Sonnet 5.5 scored ahead of Sonnet 5 and used about 121k tokens per answer, where Sonnet 5 used 497k. On our analyst search and retrieval work, it was better than Sonnet 5 in almost every way. For high-volume workflows, it had the best quality-to-cost tradeoff of the seven models we ran.”

CompanyBalyasny Asset Management
AuthorJoe Poirier, Senior AI Engineer
Quote

“Claude Sonnet 5.5 will give our customers in financial services and healthcare the confidence to use it for their most sensitive work. Sonnet 5.5 rechecks data in source documents, catching errors that Sonnet 5 failed to spot. Compared to the last model, Sonnet 5.5 was more accurate, 2.4x faster, and used 12% fewer total tokens.”

CompanyBox
AuthorYashodha Bhavnani, VP of AI Products
Quote

“Claude Sonnet 5.5 thinks in fewer, more robust steps, so builders wait less to see progress. Our coding evals showed a third fewer tool calls and roughly half the shell runs to finish a task. For everyday coding and higher-effort conversations, that means faster iteration and a smoother build loop.”

CompanyLovable
AuthorFabian Hedin, Co-founder and CTO
Quote

“With millions of Rovo-assisted actions powering our customers’ workflows each month, execution speed is critical. Claude Sonnet 5.5 will allow teams to run their Rovo Agents up to 30% faster than they could with Sonnet 5. I am excited to offer customers this choice.”

CompanyAtlassian
AuthorJamil Valliani, Head of Product, AI

Cost and speed

Pricing

Price per 1M tokensClaude Sonnet 5.5Claude Opus 5.5
Cache reads$0.20$0.20
Cache writes$2.50$5
Input tokens$2$4
Output tokens$10$20

Sonnet 5.5 requires fewer tokens per task than Sonnet 5, so it’s less expensive to run. It also generates output 30%+ faster, and its efficiency is immediately noticeable:

Prompt:

A murmuration of 400 starlings in one HTML file

Claude Sonnet 5
Claude Sonnet 5.5
Prompt:

Wind shaping sand dunes in one HTML file

Claude Sonnet 5
Claude Sonnet 5.5
Prompt:

A clock made of 24 small clocks in one HTML file

Claude Sonnet 5
Claude Sonnet 5.5

Adjusting the effort level lets you balance cost and speed against overall quality. In Claude Code and our apps, the default effort is set to Medium, while the Claude Platform defaults to High. At lower settings, Claude answers faster and uses fewer tokens, which suits routine work. At higher settings, Claude reasons for longer and checks its work more thoroughly.

Safety

Alignment

Sonnet 5.5 doesn’t advance the frontier of our models’ capabilities, so our alignment assessment focused on a targeted set of risks that apply to models of any capability level, including acting against users’ interests, misleading users, and cooperating with high-stakes misuse.

On our automated behavioral audit, which tests Claude across roughly 1,850 scenarios, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse, and honesty. On our newer containment evaluations, Sonnet 5.5 comes close to Opus 5.5, the best model we tested, in how rarely it tries to escape its sandbox, and it’s the least likely of any of our models to probe the limits of its containers. Across the full audit, Opus 5.5 still performs slightly better overall, but we found no evidence that Sonnet 5.5 pursues goals that conflict with the user’s intention.

As we described in our recent alignment assessment, no set of evaluations reliably catches every failure, and Sonnet 5.5 may have tendencies we haven’t found, which is why we pair our own alignment work with the safeguards described below.

Safeguards

Cybersecurity. Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5. Soon, cyberdefenders will be able to apply to our expanded Cyber Verification Program for tiered access to more advanced capabilities on Sonnet 5.5, Opus 5.5, and Claude Mythos models.

Biology. Sonnet 5.5 uses the same set of biology safeguards as Sonnet 5. These target harmful requests; most research, education, and clinical work is unaffected, though some microbiology and virology requests may be flagged in error. Organizations can apply to our Life Sciences Verification Program for access to safeguards designed for the full breadth of biology-related work.

Distillation. Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, allow bad actors to create highly capable models without the safeguards we build into Claude. Because Sonnet 5.5 is far more capable than its predecessor, it’s the first Sonnet model to launch with safety classifiers that prevent reasoning extraction. Sonnet 5.5 also expands preserved thinking, so Claude’s thinking cannot be decoupled from the account that created it. Most developers won’t notice a change. If you move conversations between accounts, including switching accounts mid-session in Claude Code, our docs article explains the change.

Getting started

As with Opus 5.5 and Sonnet 5, Claude Sonnet 5.5 is available with zero data retention.

Claude Sonnet 5.5 is now available on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Developers can get started on the Claude Platform with claude-sonnet-5-5. If you run Sonnet with thinking off, you’ll need to switch to the new between_tools setting, which keeps up-front thinking off, before moving to Sonnet 5.5. See our migration guide for details.

Footnotes

1 Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at Xhigh effort which represents the model’s highest score.

2 Sonnet 5.5 scores lower at Max effort than at Xhigh. FrontierCode evaluates whether a code change could be merged without human edits. It penalizes out-of-scope changes, even if they are high-quality or helpful. At Max effort, Sonnet 5.5 more often ran Claude Code’s code-review skill, which splits the review across many subagents, and in two cases Cognition examined, this led to a timeout or to extra edits beyond the task’s scope, and therefore to a lower score.

3 Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment of Sonnet 5.5 on the Claude Platform, which we found to have a bug that could degrade responses to requests that use structured outputs. We expect the effect on Sonnet 5.5’s scores, if any, to be small and to understate its performance. That bug has since been fixed.

4 OpenAI recently fixed a bug that degraded image understanding in GPT-6 Sol. Official AA-Briefcase v1.1 and GDPval-AA v2.1 scores from Artificial Analysis, and Chartography scores from Surge AI, may not have been updated yet to reflect the latest version of the model. Artificial Analysis does not expect major impacts to AA-Briefcase v1.1 and GDPval-AA v2.1. Internal testing of Chartography suggests its score was not impacted.

View original article on anthropic.com

Most Recent

PaleBlueDot AI Seeks $600 Million Private Credit to Buy Chips

US-based PaleBlueDot AI is in talks with potential lenders including Brookfield Asset Management for $600 million of private credit …

Sep 29, 2026

South Korea's Kospi was the world's worst-performing stock market in Q3, down 18.8% due to Samsung and SK Hynix sell-offs in early Q3, but remains up 60% YTD

Kospi tumbled almost a fifth in three months to September after July sell-off but is still up 60% this year

Sep 29, 2026

DeepSeek says it has partnered with Huawei to develop programming tools for Huawei's Ascend chips, including TileLang, an open-source CUDA alternative

Chinese AI firm DeepSeek said on Wednesday it has partnered with Huawei Technologies to develop programming tools for Huawei's Ascend chips …

Sep 29, 2026

Apple Pay launches in India: iPhone XR and later supported, Apple Pay VP Jennifer Bailey on why now

Apple Pay has finally launched in India, more than a decade after its US debut. The service is initially available on iPhone, iPad and Apple Watch for eligible Axis Bank Visa and Mastercard credit cards, allowing users to make contactless payments and pay on apps and websites. Apple says it views Ap

Sep 29, 2026

Similar Posts

Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for Fable cache reads

It's only the first day of September 2026, but the month and fall season are already off to the races in AI land, as Anthropic has just released its latest and most powerful large language models yet — Claude Fable 5.1 and Claude Mythos 5.1. The two names refer to the same underlying model. Fable 5.

Sep 1, 2026

Claude Opus 5.5 matches Fable 5.1 performance at lower cost and promises less "Claudish" writing

Anthropic is launching Claude Opus 5.5, the first model in a new generation. The company says it matches Claude Fable 5.1 on most tasks while costing about 40 percent less to run than Opus 5. Anthropic's benchmarks also put it ahead of OpenAI's GPT-6 Astra on most tasks, despite being significantly

Sep 22, 2026

How we made claude.ai 3x faster in two weeks / claude.dev

Inside our performance sprint: the benchmarks Claude built, the loop each Slack thread ran, and the guardrails that let us ship 3,000 changes safely.

Sep 23, 2026

Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index

Anthropic's new Sonnet model scores 56, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we have measured

Sep 27, 2026