Skip to main content
Back to Blog
AI/MLData Analysis
17 September 20264 min readUpdated 29 September 2026

Claude Sonnet 5.5 Scores 56 on the Artificial Analysis Intelligence Index, Nearing Opus 5.5 at Much Higher Token Use

Claude Sonnet 5.5 Scores 56 on the Artificial Analysis Intelligence Index, Nearing Opus 5.5 at Much Higher Token Use Anthropic has launched Claude Sonnet 5.5. At maximum effort,...

By AI Engineering Team

Claude Sonnet 5.5 Scores 56 on the Artificial Analysis Intelligence Index, Nearing Opus 5.5 at Much Higher Token Use

Anthropic has launched Claude Sonnet 5.5. At maximum effort, the model scores 56 on the Artificial Analysis Intelligence Index, an 18-point improvement over Sonnet 5 and the second-highest result behind Opus 5.5, which scores 58.

Anthropic has kept Sonnet 5.5's pricing unchanged from Sonnet 5: $0.2 per 1 million cache input tokens, $2 per 1 million input tokens, and $10 per 1 million output tokens. Despite the unchanged rates, the model uses substantially more output tokens per task. Its measured cost per task is $7.60, approximately 50% higher than Sonnet 5's.

Key findings

  • Performance close to leading models on agentic and knowledge-work tasks: In Terminal-Bench 4.0, Claude Sonnet 5.5 scores 64%, compared with 60% for Opus 5.5 and GPT-6 Astra. It also reaches near-parity with Opus 5.5 on AA-Briefcase, with scores of 1811 versus 1822 Elo; GDPval-AA, with 1844 versus 1846 Elo; and AutomationBench-AA, with headline scores of 71% versus 70%. Sonnet 5.5 uses significantly more tokens to achieve these results.
  • The highest measured token use: At maximum effort, Claude Sonnet 5.5 generates approximately 193,000 output tokens per Intelligence Index task. This is about 60% more than Opus 5.5 or Sonnet 5 at maximum effort, and roughly seven times the output of GPT-6 Astra at maximum effort.
  • Pricing and cost efficiency: Sonnet 5.5 is priced at $2 per million input tokens and $10 per million output tokens, matching GPT-6 Sol. At these rates, it falls outside the Intelligence versus Cost per Task Pareto frontier. At high effort, it trails Opus 5.5, while lower-effort configurations of GPT-6 Astra and GPT-6 Sol provide similar performance at lower cost. The maximum-effort configuration is the most competitive for Sonnet 5.5, narrowly behind GPT-6 Sol in intelligence at nearly the same cost per task.
  • Lower factual and scientific performance than Opus 5.5: In AA-Omniscience, Sonnet 5.5 achieves 54% factual accuracy, compared with 66% for Opus 5.5. Its hallucination rate is lower, at 47% versus 59%. Sonnet 5.5 also scores approximately six points below Opus 5.5 on Humanity's Last Exam and SciCode.

The evaluations used a pre-release deployment of Claude Sonnet 5.5. Anthropic identified a bug that could degrade responses to requests using structured outputs. The issue was fixed before the public release. Anthropic expects minimal changes or slightly understated performance, and the relevant evaluations will be repeated.

Model details

  • Context window: 1 million tokens with image and text input, unchanged from Sonnet 5.
  • Pricing: $2 per 1 million input tokens and $10 per 1 million output tokens. Cache writes cost $2.5 per 1 million tokens, while cache reads cost $0.2 per 1 million tokens.
  • Effort settings: Low, medium, high, xhigh, and max. Intelligence Index evaluations were run at all five settings with Anthropic's default fallback enabled. Sonnet 5.5 fell back on approximately 0.1% of Intelligence Index tasks, primarily in Terminal-Bench 4.0. In every fallback case, it used Sonnet 5.

Terminal benchmark results

Claude Sonnet 5.5 at maximum effort makes a substantial improvement on Terminal-Bench. It scores 64% on Terminal-Bench 4.0, a 50-point increase over Claude Sonnet 5 at maximum effort. The result is slightly above the 60% scores recorded by Opus 5.5 and GPT-6 Astra at xhigh effort.

On Terminal-Bench-Science, which evaluates agentic terminal use for realistic scientific research workflows across multiple domains, Sonnet 5.5 scores 53%. It ranks behind only GPT-6 Astra and Opus 5.5. Terminal-Bench-Science is not currently part of the Artificial Analysis Intelligence Index.

Token usage across effort settings

Claude Sonnet 5.5 at maximum effort uses approximately 193,000 output tokens per Intelligence Index task, the highest level measured in the evaluation. This is roughly seven times the output token use of GPT-6 Astra at maximum effort.

The lower effort settings cover a wider range of intelligence and output-token tradeoffs. The low, medium, and high settings rank behind GPT-6 Sol's high, xhigh, and max configurations in this comparison. GPT-6 Sol delivers higher performance with fewer output tokens in this range.

The Intelligence Index results for Claude Sonnet 5.5 vary across all five reasoning-effort settings, with the maximum setting producing the model's highest score of 56.