Claude Opus 5.5 leads the Artificial Analysis Intelligence Index with lower pricing
Claude Opus 5.5 leads the Artificial Analysis Intelligence Index with lower pricing September 22, 2026 Claude Opus 5.5 matches GPT 6 Astra on evaluations including Terminal Benc...
By Software Development Team
Claude Opus 5.5 leads the Artificial Analysis Intelligence Index with lower pricing
September 22, 2026
Claude Opus 5.5 matches GPT-6 Astra on evaluations including Terminal-Bench 4.0 and AutomationBench-AA, while extending Anthropic's lead in agentic knowledge work.
At maximum effort, Opus 5.5 scores 58 on the Artificial Analysis Intelligence Index, the highest score measured by several points. Anthropic has also reduced pricing to $4 per 1 million input tokens and $20 per 1 million output tokens, compared with $5 and $25 for Opus 5. Cache reads now cost $0.20 per 1 million tokens, down from $0.50.
Key results
- Strong performance across the index: Opus 5.5 records the leading score on six of the ten Intelligence Index evaluations. Results include 61.4% on Humanity's Last Exam, compared with the previous best of 59.1% from Claude Fable 5.1, and 66.9% on SciCode, compared with 63.1% from Fable 5.1. It also leads on GDPval-AA v2.1, AA-Briefcase v1.1, AA-Omniscience, and AutomationBench-AA. On Terminal-Bench 4.0, it scores 59.6%, matching GPT-6 Astra at the xhigh setting and finishing 11 points ahead of Opus 5. It remains behind on CritPt, AA-LCR, and GDP.pdf.
- Leadership in agentic knowledge work: On AA-Briefcase, a private frontier knowledge-work evaluation, Opus 5.5 reaches an Elo score of 1,822. That is 143 points above Fable 5.1, with higher scores for both analytical quality and presentation. It is the first time Anthropic has achieved presentation quality above GPT-5.6 Sol in this evaluation. AA-Briefcase measures whether models can produce accurate, well-presented professional outputs using the open-source Stirrup reference-agent harness.
- Comparable cost per task: Opus 5.5 at maximum effort uses approximately 119,000 output tokens per Intelligence Index task. Opus 5 uses about 73,000, Fable 5.1 about 78,000, and GPT-6 Astra about 27,000. Despite using 1.6 times as many output tokens as Opus 5, Opus 5.5 reaches a similar cost per task.
- Multiple effort settings on the cost-performance frontier: The max, xhigh, high, and medium settings all sit on the Pareto frontier for Intelligence score versus cost per task. They either cost less or outperform other models scoring above 50, including GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.
Model details
- Context window: 1 million tokens, with image and text input support, unchanged from Opus 5.
- Pricing: $4 per 1 million input tokens and $20 per 1 million output tokens, a 20% reduction from Opus 5's $5 and $25 rates.
- Cache pricing: Cache writes cost $5 per 1 million tokens for the five-minute time-to-live period, down from $6.25. Cache reads cost $0.20 per 1 million tokens, a 60% reduction from Opus 5's $0.50 rate. This represents a 95% discount compared with uncached input pricing, up from 90% for previous Opus models.
- Effort settings: The model offers low, medium, high, xhigh, and max settings. The Intelligence Index evaluations were run at all five settings with Anthropic's default fallback enabled.
Agentic knowledge-work results
At maximum effort, Claude Opus 5.5 leads AA-Briefcase v1.1 with an Elo score of 1,822, 143 points above Fable 5.1. It also leads GDPval-AA v2.1 with an Elo score of 1,846, which is 111 points above Claude Fable 5.1 and 138 points above Claude Opus 5.
On AA-Briefcase, Opus 5.5 leads on analytical quality and presentation sub-scores, while remaining just behind Fable 5.1 on rubric-based scoring.