CompaniesInvestorsPeople
Home
Loading

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

Get in Touch

  • Contact

  • Request a Demo

  • Request Data Updates

  • Add a Company

Research

  • Companies

  • Investors

  • People

aVenture

  • Download App

  • Pricing

Download the aVenture Research beta for iOS and iPadOSDownload aVenture Research on the Mac App Store

Resources

  • Documentation

  • CLI

  • MCP

  • Feature Requests

  • Sitemap

Member

Backed by

© aVenture Investment Company, 2026. All rights reserved.

San Francisco, CA, USA

Privacy Policy · Terms of Service

aVenture Investment Company ("aVenture") is an independent research platform providing detailed analysis and data on startups, venture capital investments, and key industry individuals. It is not a registered investment adviser, broker-dealer, or investment advisor and does not provide investment advice or recommendations. The data provided by aVenture does not constitute recommendations or advice, whether by methodology, analysis, AI-generated content, or a statement written by a staff member of aVenture.

aVenture is not affiliated with any of the people, companies, organizations, government agencies, regulatory bodies, or investment funds we provide coverage for on this site unless explicitly stated otherwise. Users assume full responsibility for decisions made based on information obtained from this platform. Links to external websites do not imply endorsement or affiliation with aVenture. Any links that provide the ability to invest in a primary or secondary transaction in a company are for convenience only and do not constitute solicitations or offers to buy or sell an investment. Investors should exercise heightened precaution and due diligence when investing in private companies, especially those not independently audited.

While we strive to provide valuable insights with objectivity and professional diligence, we cannot guarantee the accuracy of the information provided on our platform. Before making any investment decisions, you should verify the accuracy of all pertinent details for your decision. To the fullest extent permitted by law, aVenture shall not be liable for any direct, indirect, incidental, consequential, or financial damages arising from use of this site, whether by consumers of its contents directly or by persons or organizations covered by our research, even if we are advised of the possibility. Our best-efforts processes and correction request forms do not create a warranty or duty of care.

Profiles on this platform may include content generated in part by large language models (LLMs, artificial intelligence) that aggregate publicly available sources (e.g., SEC EDGAR, public filings, press releases). Source attribution is provided where known; always verify statements and claims here against original sources before relying on any data. Content on our site may contain inaccuracies, omissions, or what are commonly called 'hallucinations' if generated in part or in full by AI / LLMs. The risk can also exist even when content is written by a human, as internal and third-party sources may also have inaccuracies for the same or different reasons. While we randomly audit a proportion of content, this is not exhaustive.

We recommend that an independent auditor be hired to verify the accuracy of the information before relying on it for any sensitive decisions. By accessing this platform, you agree not to rely solely on any information generated by AI, aggregated, or sourced or written otherwise on this site, for investment, financial, or other decisions. aVenture assumes no responsibility for inaccuracies, omissions, or hallucinations. You must independently verify all data from primary sources. Use of this platform constitutes your waiver of claims for reliance-based damages, including negligent misrepresentation. To report an error, request a correction, or dispute information about a company or individual, contact us via our request data updates form.

Loading
Loading
Home
News
Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index

From Artificial Analysis

By

September 27, 2026

Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index

Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index
Artificial Analysis

All articles

September 28, 2026

Anthropic has launched Claude Sonnet 5.5: it scores 56 on the Artificial Analysis Intelligence Index, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we've seen

See model page

With max effort, Sonnet 5.5 gains 18 points over Sonnet 5 and moves to #2 on the Intelligence Index, behind only Opus 5.5 (max). Anthropic has priced Sonnet 5.5 identically to Sonnet 5 at $0.2/$2/$10 per 1M cache input/input/output tokens. However, it outputs a higher number of Output Tokens per Task and costs $7.60 per task (~50% higher than Sonnet 5's Cost per Task).

Key takeaways:

➤ Meets leading models on agentic terminal use and knowledge work: In Terminal-Bench 4.0, Claude Sonnet 5.5 reaches 64% against 60% for Opus 5.5 and GPT-6 Astra. On AA-Briefcase (1811 vs 1822 Elo), GDPval-AA (1844 vs 1846 Elo), and AutomationBench-AA (71% vs 70% headline score), Sonnet 5.5 reaches parity with Opus 5.5, albeit with significantly higher token usage to achieve it

➤ Heaviest token use we have measured: At max effort, where it reaches performance nearing that of Opus 5.5, Claude Sonnet 5.5 used ~193k Output Tokens per Intelligence Index Task. This is the highest token use we have measured, around 60% higher than Opus 5.5 (max) or Sonnet 5 (max) and ~7x GPT-6 Astra (max)

➤ Pricing remains at $2/$10 per million tokens of input/output, matching GPT-6 Sol: At this pricing Claude Sonnet 5.5 sits off the Intelligence vs. Cost per Task Pareto Frontier. At high effort levels it sits behind Opus 5.5, while lower efforts have GPT-6 Astra or Sol configurations delivering equivalent performance for lower cost. The high effort setting is the most competitive on this basis, sitting very narrowly behind GPT-6 Sol on Intelligence at effectively the same Cost per Task

➤ Behind Opus 5.5 on factual knowledge and scientific reasoning: As a smaller class model, Sonnet 5.5 still lags on factual knowledge in AA-Omniscience compared to Opus 5.5. It scores 54% against 66% for factual accuracy, though with a lower hallucination rate (47% against 59%). It also sits ~6 points lower on Humanity's Last Exam and SciCode compared to Opus

These evaluations were conducted on a pre-release deployment of Claude Sonnet 5.5, which Anthropic found to have a bug that can degrade responses to requests that use structured outputs. This is fixed for the public release and Anthropic expects minimal change or slightly understated performance, but we will be re-running relevant evaluations soon.

Other model details:

➤ Context window: 1 million tokens with image and text input, unchanged from Sonnet 5

➤ Pricing: Unchanged from Sonnet 5's latest $2/$10 per 1M input/output tokens; cache writes at $2.5, cache reads $0.2

➤ Effort settings: Five (low, medium, high, xhigh, max). Intelligence Index evaluations were run at all five with Anthropic's default fallback enabled. We see Sonnet 5.5 fall back in ~0.1% of tasks across the Intelligence Index, primarily in Terminal-Bench 4.0, falling back to Sonnet 5 in all cases

Claude Sonnet 5.5 (max) makes large strides on Terminal-Bench, sitting among the top models for both Terminal-Bench 4.0 and Terminal-Bench-Science. In Terminal-Bench 4.0 it scores 64%, a 50 point increase over Claude Sonnet 5 (max), and slightly above 60% for Opus 5.5 and GPT-6 Astra (xhigh).

On our leaderboard for Terminal-Bench-Science, a benchmark of agentic terminal use to complete realistic scientific research workflows across domains, it scores 53% and sits behind only GPT-6 Astra and Opus 5.5. Terminal-Bench-Science is not currently included in the Artificial Analysis Intelligence Index.

To achieve its outsized performance, Claude Sonnet 5.5 (max) uses ~193k output tokens per Intelligence Index task, the most we have measured and ~7x that of GPT-6 Astra (max).

However, its lower effort settings span a broader area of Intelligence versus Output Tokens per Task tradeoffs, with low, medium, and high effort settings sitting behind GPT-6 Sol high, xhigh, and max efforts. In this area Sol provides higher performance with fewer output tokens.

Full breakdown of the individual evaluations in the Artificial Analysis Intelligence Index for Claude Sonnet 5.5 across all reasoning efforts:

Compare Claude Sonnet 5.5 with other leading models at: https://artificialanalysis.ai/models/releases/claude-sonnet-5-5

Read the latest

Announcing the Artificial Analysis Cyber Index Alliance

The Artificial Analysis Cyber Index Alliance brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities.

September 28, 2026

GPT-6 Sol and Luna push the cost efficiency frontier

GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others

September 22, 2026

Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index

Anthropic's new Opus scores 58 and arrives with a 20% price cut and a larger cache hit discount

September 22, 2026

View original article on artificialanalysis.ai

Most Recent

PaleBlueDot AI Seeks $600 Million Private Credit to Buy Chips

US-based PaleBlueDot AI is in talks with potential lenders including Brookfield Asset Management for $600 million of private credit …

Sep 29, 2026

South Korea's Kospi was the world's worst-performing stock market in Q3, down 18.8% due to Samsung and SK Hynix sell-offs in early Q3, but remains up 60% YTD

Kospi tumbled almost a fifth in three months to September after July sell-off but is still up 60% this year

Sep 29, 2026

DeepSeek says it has partnered with Huawei to develop programming tools for Huawei's Ascend chips, including TileLang, an open-source CUDA alternative

Chinese AI firm DeepSeek said on Wednesday it has partnered with Huawei Technologies to develop programming tools for Huawei's Ascend chips …

Sep 29, 2026

Apple Pay launches in India: iPhone XR and later supported, Apple Pay VP Jennifer Bailey on why now

Apple Pay has finally launched in India, more than a decade after its US debut. The service is initially available on iPhone, iPad and Apple Watch for eligible Axis Bank Visa and Mastercard credit cards, allowing users to make contactless payments and pay on apps and websites. Apple says it views Ap

Sep 29, 2026

Similar Posts

Claude Opus 5.5 matches Fable 5.1 performance at lower cost and promises less "Claudish" writing

Anthropic is launching Claude Opus 5.5, the first model in a new generation. The company says it matches Claude Fable 5.1 on most tasks while costing about 40 percent less to run than Opus 5. Anthropic's benchmarks also put it ahead of OpenAI's GPT-6 Astra on most tasks, despite being significantly

Sep 22, 2026

Claude Opus 5.5 delivers Fable 5.1 performance - and costs 40% less

Elyse Betters Picaro/ZDNET ZDNET’s key takeaways Claude Opus 5.5 could be a big win for power users. Developers may see faster coding with fewer steps. Anthropic says the upgrade is safer, cheaper, and less wordy. Less than two months after the release of Claude Opus 5, Anthropic is back with Opus 5

Sep 22, 2026

OpenAI's GPT-6 Sol doubles its accuracy rate - for half the cost

GPT-6 Sol and Luna arrive less than three months after GPT-5.6, with OpenAI claiming major accuracy gains, dramatic cost reductions, and oddly revealing Claude comparisons across key benchmarks.

Sep 22, 2026

Anthropic releases Claude Opus 5.5 and OpenAI counters with two cheaper GPT-6 models

Despite rampant worries about runaway artificial intelligence, the two big AI model makers aren’t yet slowing down: Anthropic PBC released Claude Opus 5.5 today and cut its price 20%, and minutes later OpenAI Group PBC put out two new GPT-6 models, Sol and Luna, at half what their predecessors cost.

Sep 22, 2026