Artificial Analysis Releases Intelligence Index v4.2
Artificial Analysis Releases Intelligence Index v4.2 Artificial Analysis Intelligence Index v4.2 introduces more complex and realistic evaluations, adds private test sets, and u...
By AI Engineering Team
Artificial Analysis Releases Intelligence Index v4.2
Artificial Analysis Intelligence Index v4.2 introduces more complex and realistic evaluations, adds private test sets, and updates grading infrastructure while work continues on the planned v5 release.
Released on September 4, 2026, version 4.2 arrives eight months after Intelligence Index v4, which launched in January. The update is intended to keep the index relevant as new models continue to advance rapidly.
Changes in Intelligence Index v4.2
- AA-Briefcase added: An agentic knowledge-work evaluation using a private held-out test set.
- Surge’s GDP.pdf added: A long-context document-reasoning benchmark covering 4,592 PDF pages.
- GPQA Diamond removed: The scientific reasoning benchmark has become saturated.
- Private test-set weighting increased: Held-out evaluations now account for 40% of the index, twice the proportion used in v4.1.
- Grading infrastructure upgraded: Several benchmarks received changes intended to improve scoring accuracy, stability, and robustness.
AA-Briefcase
AA-Briefcase evaluates models on realistic agentic knowledge-work tasks within complex projects created by industry experts. The projects span multiple weeks, contain many linked tasks, and use thousands of source files.
The evaluation combines rubric-based and pairwise grading to measure verifiable task completion, analytical quality, and presentation quality. Its private held-out test set is designed to provide a broader view of agentic capability in knowledge work while reducing opportunities to optimize specifically for the benchmark.
GDP.pdf
Created by Surge AI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must combine evidence from 4,592 pages containing text, tables, charts, footnotes, and exclusions.
Responses are assessed against 1,275 expert-authored atomic criteria. The headline All-pass Rate counts a task only when every criterion has been satisfied.
Greater use of held-out data
Private held-out test sets now represent 40% of the Intelligence Index weighting, up from 20% in v4.1. The held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt.
Artificial Analysis says the change reduces the ability of model developers to game evaluations. The share of held-out data is expected to increase further in Index v5.
Grading infrastructure updates
For AA-LCR v1.1, the update adds a grading system prompt and corrects errors and ambiguities in answer keys, improving scoring accuracy.
GDPval-AA v2 and AA-Briefcase received sampling improvements and a re-anchored Elo scale to make ratings more stable as new models are added. SciCode received more robust grading sandboxes so that slow but correct code is not treated as a failure.
Key results
Anthropic and OpenAI lead the index
Anthropic’s Claude Fable 5.1 leads the index, followed by OpenAI’s GPT-6 Astra. GPT-6 Astra records a four-point gain over GPT-5.6 Sol.
Meta ranks third among labs on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z.AI, and Google.
Four labs share the Cost per Task frontier
Anthropic, OpenAI, Meta, and Z.AI occupy the updated Cost per Task Pareto frontier.
GPT-6 Astra leads the output-token frontier
GPT-6 Astra is more token-efficient than almost every other model near the intelligence frontier. Claude Fable 5.1, Grok 4.5, and Gemini 3.5 Flash-Lite appear at different ends of the curve. The comparison excludes models scoring below 25 on the Index.
AA-Briefcase results
Anthropic’s Claude Fable 5.1 and Opus 5 lead AA-Briefcase, followed by GPT-6 Astra and Muse Spark 1.3.
GPT-6 Astra records a substantial improvement over GPT-5.6 Sol, gaining approximately 85 Elo points. AA-Briefcase measures performance on multi-week agentic knowledge-work projects with linked tasks and thousands of input files, combining rubric and pairwise grading across task success, analysis, and presentation.
GDP.pdf results
OpenAI leads GDP.pdf, with GPT-6 Astra achieving an All-pass Rate of 33.2% and GPT-5.6 Sol reaching 28.2%. Claude Fable 5.1 follows at 26.2%.
GDP.pdf requires models to synthesize information across 100 professional documents and 4,592 pages. Its results are based on 1,275 expert-authored criteria, with credit awarded only when a task meets every criterion.
Index v5 development
Artificial Analysis says work on Intelligence Index v5 has been underway for several months. Some elements planned for that release are being introduced earlier through interim updates, while further incremental releases are planned before v5 becomes available.