Claude Fable 5.1 Leads the Artificial Analysis Intelligence Index at a Higher Cost per Task
Claude Fable 5.1 Leads the Artificial Analysis Intelligence Index at a Higher Cost per Task September 1, 2026 Claude Fable 5.1 achieved a score of 66 at maximum effort on the Ar...
By AI Engineering Team
Claude Fable 5.1 Leads the Artificial Analysis Intelligence Index at a Higher Cost per Task
September 1, 2026
Claude Fable 5.1 achieved a score of 66 at maximum effort on the Artificial Analysis Intelligence Index, the highest score measured by the evaluation. It was followed by Claude Opus 5 at 63, Claude Fable 5 at 62, GPT-5.6 Sol at 61, and Grok 4.6 at 61.
The evaluation used Anthropic's default server-side fallback configuration. Safety-flagged requests were routed to Claude Opus 4.8 or Claude Opus 5, and fallback responses accounted for approximately 4% of output tokens across the Intelligence Index.
Benchmark performance
Claude Fable 5.1 improved by four points over Claude Fable 5 on the Intelligence Index. It also produced higher results on several individual evaluations:
- On HLE, it scored 59.1%, compared with the previous best of 55.5% from Claude Fable 5.
- It recorded the highest measured scores on Terminal-Bench v2.1, at 91.4%, and SciCode, at 62.0%.
- On 𝜏³-Banking, it gained nine points over Claude Fable 5.
- On GDPval-AA v2, it reached 1,853 Elo, 130 points above Claude Fable 5.
- On AA-Briefcase, it reached 1,694 Elo, 122 points above Claude Fable 5.
GDPval-AA v2 and AA-Briefcase evaluate agentic knowledge work. Against Claude Opus 5, Claude Fable 5.1's lead on GDPval-AA v2 is within the confidence interval. The models are also effectively tied on AA-Briefcase, where Claude Fable 5.1 scored 1,694 and Claude Opus 5 scored 1,685. Fable 5.1 performed better on analytical quality and rubric correctness, while Opus 5 performed better on presentation.
Both evaluations use Stirrup, an open source reference agent harness, to assess whether models can produce accurate and well-presented professional outputs.
Pricing and token usage
Anthropic reduced the cache read price for Claude Fable 5.1 from $1 to $0.25 per 1 million cached input tokens, a 75% reduction. Standard input and output prices remain $10 and $50 per 1 million tokens, while cache writes cost $12.50 per 1 million tokens.
Despite the cache price reduction, Claude Fable 5.1 at maximum effort costs more per Intelligence Index task than Claude Fable 5. Fable 5.1 costs $3.76 per task, compared with $3.14 for Fable 5 and $2.34 for Claude Opus 5. Its higher cost is primarily associated with using approximately 1.7 times as many output tokens as Fable 5.
The lower cache read price saves approximately $1.40 per task, with most of the savings concentrated in agentic evaluations where a large share of input tokens are cache reads. Without the price reduction, the estimated cost would be approximately $5.16 per task.
At the xhigh effort setting, Claude Fable 5.1 scores 65 and costs $2.72 per task, $1.04 less than at maximum effort. This remains above the $2.34 cost of Claude Opus 5 at maximum effort, which scores 63.
Intelligence and output-token efficiency
Claude Fable 5.1 occupies the upper end of the Intelligence versus Output Tokens per Task Pareto frontier. Every model variant scoring higher than GPT-5.6 Sol at medium effort is matched or exceeded by at least one Fable 5.1 effort setting in both intelligence and token usage.
The five Fable 5.1 effort settings span an 11-fold range in output token usage, from 13.1 million tokens at low effort to 143.7 million at maximum effort. Scores range from 58 to 66 on the Artificial Analysis Intelligence Index.
GPT-5.6 Sol at medium effort uses marginally fewer output tokens than Fable 5.1 at low effort, 12 million compared with 13.1 million. However, Fable 5.1's low-effort score is part of the same broader frontier across its available effort settings.
Additional evaluation results
On AA-Omniscience, Claude Fable 5.1 at maximum effort attempted 93.4% of questions, compared with 87.8% for Claude Opus 5. It recorded an accuracy of 67.2%, the highest measured result and above Claude Fable 5's 65.4%.
The higher attempt rate also led to more incorrect responses being attempted. Among questions it did not answer correctly, Fable 5.1 attempted a response 72.6% of the time, compared with 63.6% for Claude Fable 5. These effects offset one another, leaving Claude Fable 5.1 level with Claude Fable 5 on the AA-Omniscience Index.
Model details
- Context window: 1 million tokens
- Inputs: Image and text inputs are supported, as with Anthropic's other recent launches.
- Input pricing: $10 per 1 million tokens
- Output pricing: $50 per 1 million tokens
- Cache write pricing: $12.50 per 1 million tokens
- Cache read pricing: $0.25 per 1 million tokens
Claude Fable 5.1 therefore combines the highest measured Artificial Analysis Intelligence Index score with increased output-token use and a higher per-task cost than Claude Fable 5. The cache read price reduction lowers costs for workloads that reuse large amounts of context, particularly agentic tasks.