Anthropic Releases Claude Haiku 5.5 with a Score of 43 on the Artificial Analysis Intelligence Index
Anthropic Releases Claude Haiku 5.5 with a Score of 43 on the Artificial Analysis Intelligence Index Anthropic has released Claude Haiku 5.5, which scores 43 on the Artificial A...
By AI Engineering Team
Anthropic Releases Claude Haiku 5.5 with a Score of 43 on the Artificial Analysis Intelligence Index
Anthropic has released Claude Haiku 5.5, which scores 43 on the Artificial Analysis Intelligence Index. The result is 26 points higher than the previous Haiku release, introduced one year earlier.
Haiku 5.5 is the first Haiku model to include Anthropic's effort settings and adaptive thinking. Anthropic has also introduced tiered pricing for the model.
Pricing and token usage
For prompts of up to 100,000 tokens, Haiku 5.5 costs $0.10 per 1 million input tokens and $0.50 per 1 million output tokens. This is the same price as GPT-6 Luna and represents 10% of the previous Haiku model's price.
For prompts above 100,000 tokens, the price increases fivefold to $0.50 per 1 million input tokens and $2.50 per 1 million output tokens. Current provisional cost figures do not include this higher pricing tier.
Cache reads cost $0.01 per 1 million tokens, increasing to $0.05 above 100,000 tokens. Five-minute cache writes cost $0.125 per 1 million tokens, or $0.625 above 100,000 tokens.
Benchmark results
Artificial Analysis Intelligence Index
At maximum effort, Haiku 5.5 scores slightly higher than several other small-class models:
- GLM-5.3 Flash: 42
- Gemini 3.8 Flash: 41
- GPT-6 Luna: 38
Its score is close to Kimi K3, an open-weights model with 2.8 trillion parameters, which scores 44. Haiku 5.5 trails Claude Sonnet 5.5 at maximum effort, which scores 56, by 13 points.
Haiku 5.5 uses substantially more output tokens than GPT-6 Luna for comparable Intelligence Index results. At maximum effort, it uses approximately 162,000 output tokens per task, about three times GPT-6 Luna's approximately 50,000 tokens.
Increasing Haiku 5.5's setting from xhigh to max adds two points while requiring approximately 1.8 times as many tokens. At high effort, Haiku 5.5 scores 38 and uses approximately 55,000 tokens per task. GPT-6 Luna reaches the same score at maximum effort with approximately 50,000 tokens. The difference is larger at lower effort settings.
AA-Briefcase
On AA-Briefcase, a private evaluation of realistic knowledge-work tasks, Haiku 5.5 at maximum effort reaches 1578 Elo. It scores ahead of models including Kimi K3 and GLM-5.3 and is comparable to Muse Spark 1.3 at maximum effort.
Terminal-Bench 4.0
Haiku 5.5 scores 33% on Terminal-Bench 4.0, compared with 0% for Haiku 4.5. This result is level with GLM-5.3 Flash and ahead of Gemini 3.8 Flash at 20% and GPT-6 Luna at 13%.
AA-Omniscience
Haiku 5.5 has lower factual-knowledge accuracy than some larger sibling models. Its AA-Omniscience accuracy is 36%, compared with 55% for Gemini 3.8 Flash and 44% for GPT-6 Luna.
The lower accuracy is partly associated with a greater willingness to acknowledge uncertainty. Haiku 5.5's hallucination rate is 40%, compared with 55% for Gemini 3.8 Flash and 77% for GPT-6 Luna.
AutomationBench-AA
Haiku 5.5 scores 35% on AutomationBench-AA, compared with 53% to 60% for GPT-6 Luna, Gemini 3.8 Flash, and GLM-5.3 Flash.
During pre-release testing, a safety-refusal issue caused Haiku 5.5 to refuse too many tasks. The reported score may therefore be understated, and the evaluation is expected to be repeated after the issue is resolved.
Other model details
- Context window: 1 million tokens, up from 200,000 tokens for Claude 4.5 Haiku
- Input and output: Text and image input, with text output
- Pricing up to 100,000 tokens: $0.10 per 1 million input tokens and $0.50 per 1 million output tokens
- Pricing above 100,000 tokens: $0.50 per 1 million input tokens and $2.50 per 1 million output tokens
- Cache reads: $0.01 per 1 million tokens, increasing to $0.05 above 100,000 tokens
- Five-minute cache writes: $0.125 per 1 million tokens, increasing to $0.625 above 100,000 tokens
Claude Haiku 5.5 at maximum effort uses approximately 162,000 output tokens per Intelligence Index task. This is more than Opus 5.5 at maximum effort and about three times the usage of GPT-6 Luna at maximum effort. Across effort settings, Haiku 5.5 uses more output tokens than GPT-6 Luna for similar Intelligence Index scores.
On AA-Briefcase, Haiku 5.5 at maximum effort reaches 1578 Elo. The result falls within the confidence intervals of GPT-6 Astra at maximum effort and Claude Fable 5.1 at high effort.