Artificial Analysis Capability Indices v1.1 Updates Domain-Specific Evaluations
Artificial Analysis Capability Indices v1.1 Updates Domain Specific Evaluations September 14, 2026 Artificial Analysis has updated its Capability Indices to version 1.1. The rev...
By AI Engineering Team
Artificial Analysis Capability Indices v1.1 Updates Domain-Specific Evaluations
September 14, 2026
Artificial Analysis has updated its Capability Indices to version 1.1. The revision applies stronger domain tuning by combining domain-specific slices of Intelligence Index v4.3 evaluations with specialized evaluations.
The Capability Indices map tasks from O*NET occupations to benchmarks that represent those tasks. Each benchmark is weighted according to how frequently its capability appears across the relevant tasks. Version 1.1 includes updates from Intelligence Index v4.2 and v4.3, with domain-specific slicing where available.
Changelog
| Index | Added | Deleted |
|---|---|---|
| Finance & Accounting | Agentic Tool Use (AutomationBench-AA, Finance); Agentic Knowledge Work (AA-Briefcase); Long-Context (GDP.pdf) | Agentic Customer Interaction (𝜏³-Banking) |
| Strategy & Ops | Agentic Tool Use (AutomationBench-AA, Operations); Agentic Knowledge Work (AA-Briefcase); Long-Context (GDP.pdf) | Agentic Customer Interaction (𝜏³-Banking) |
| Legal | Agentic Tool Use (AutomationBench-AA, Operations and Support); Agentic Knowledge Work (AA-Briefcase); Long-Context (GDP.pdf) | Agentic Customer Interaction (𝜏³-Banking) |
| Healthcare & Medical | Agentic Tool Use (AutomationBench-AA, Operations and Support); Long-Context Reasoning (MLCR-AA) | Agentic Knowledge Work (AA-Briefcase); Agentic Customer Interaction (𝜏³-Banking) |
| Engineering | Agentic Terminal Use (Terminal-Bench v4.0) | Agentic Terminal Use (Terminal-Bench v2.1); Reasoning (GPQA Diamond) |
| Economics | Agentic Knowledge Work (AA-Briefcase) | None |
Changes by Index
Finance & Accounting
Agentic Tool Use was added using the Finance slice of AutomationBench-AA. Agentic Customer Interaction was removed. AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work, while GDP.pdf was added alongside LCR under Long-Context.
| Capability | Evaluations | v1.0 | v1.1 |
|---|---|---|---|
| Business Knowledge | AA-Omniscience | 30% | 30% |
| Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 30% | 30% |
| Reasoning | HLE | 20% | 20% |
| Agentic Tool Use | AutomationBench-AA | 0% | 10% |
| Long-Context | LCR, GDP.pdf | 5% | 5% |
| Non-Hallucination | AA-Omniscience | 5% | 5% |
Strategy & Ops
Agentic Tool Use was added using the Operations slice of AutomationBench-AA to represent tool use in day-to-day operational tasks. Agentic Customer Interaction was removed. AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work, and GDP.pdf was added alongside LCR under Long-Context.
| Capability | Evaluations | v1.0 | v1.1 |
|---|---|---|---|
| Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 35% | 35% |
| Business Knowledge | AA-Omniscience | 30% | 30% |
| Agentic Tool Use | AutomationBench-AA | 0% | 30% |
| Long-Context | LCR, GDP.pdf | 5% | 5% |
Legal
Agentic Tool Use was added using the Operations and Support slices of AutomationBench-AA to represent tool use in legal workflows. Agentic Customer Interaction was removed. AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work, and GDP.pdf was added alongside LCR under Long-Context.
| Capability | Evaluations | v1.0 | v1.1 |
|---|---|---|---|
| Legal Knowledge | AA-Omniscience | 35% | 35% |
| Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 25% | 25% |
| Reasoning | HLE | 15% | 15% |
| Long-Context | LCR, GDP.pdf | 10% | 10% |
| Non-Hallucination | AA-Omniscience | 10% | 10% |
| Agentic Tool Use | AutomationBench-AA | 0% | 5% |
Healthcare & Medical
Long-Context Reasoning was added using MLCR-AA, which evaluates reasoning across lengthy clinical records. Agentic Tool Use was also added using the Operations and Support slices of AutomationBench-AA to cover operational healthcare workflows. Agentic Customer Interaction was removed, and AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work.
| Capability | Evaluations | v1.0 | v1.1 |
|---|---|---|---|
| Medical & Health Knowledge | AA-Omniscience | 35% | 30% |
| Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 25% | 25% |
| Long-Context Reasoning | MLCR-AA | 0% | 15% |
| Non-Hallucination | AA-Omniscience | 15% | 10% |
| Reasoning | HLE | 15% | 10% |
| Agentic Tool Use | AutomationBench-AA | 0% | 10% |
Engineering
Agentic Terminal Use was updated to Terminal-Bench v4.0. GPQA Diamond was removed from Reasoning, and AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work.
| Capability | Evaluations | v1.0 | v1.1 |
|---|---|---|---|
| Engineering Knowledge | AA-Omniscience | 35% | 35% |
| Reasoning | HLE, CritPt | 35% | 30% |
| Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 25% | 20% |
| Agentic Terminal Use | Terminal-Bench 4.0 | 5% | 15% |
Economics
AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work.
| Capability | Evaluations | v1.0 | v1.1 |
|---|---|---|---|
| Economics Knowledge | AA-Omniscience | 35% | 35% |
| Reasoning | HLE | 35% | 35% |
| Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | 15% | 25% |
| Long-Context Reasoning | LCR | 15% | 5% |