Skip to main content
Back to Blog
AI/MLData Analysis
3 September 20265 min readUpdated 17 September 2026

Artificial Analysis Capability Indices v1.1 Updates Domain-Specific Evaluations

Artificial Analysis Capability Indices v1.1 Updates Domain Specific Evaluations September 14, 2026 Artificial Analysis has updated its Capability Indices to version 1.1. The rev...

By AI Engineering Team

Artificial Analysis Capability Indices v1.1 Updates Domain-Specific Evaluations

September 14, 2026

Artificial Analysis has updated its Capability Indices to version 1.1. The revision applies stronger domain tuning by combining domain-specific slices of Intelligence Index v4.3 evaluations with specialized evaluations.

The Capability Indices map tasks from O*NET occupations to benchmarks that represent those tasks. Each benchmark is weighted according to how frequently its capability appears across the relevant tasks. Version 1.1 includes updates from Intelligence Index v4.2 and v4.3, with domain-specific slicing where available.

Changelog

IndexAddedDeleted
Finance & AccountingAgentic Tool Use (AutomationBench-AA, Finance); Agentic Knowledge Work (AA-Briefcase); Long-Context (GDP.pdf)Agentic Customer Interaction (𝜏³-Banking)
Strategy & OpsAgentic Tool Use (AutomationBench-AA, Operations); Agentic Knowledge Work (AA-Briefcase); Long-Context (GDP.pdf)Agentic Customer Interaction (𝜏³-Banking)
LegalAgentic Tool Use (AutomationBench-AA, Operations and Support); Agentic Knowledge Work (AA-Briefcase); Long-Context (GDP.pdf)Agentic Customer Interaction (𝜏³-Banking)
Healthcare & MedicalAgentic Tool Use (AutomationBench-AA, Operations and Support); Long-Context Reasoning (MLCR-AA)Agentic Knowledge Work (AA-Briefcase); Agentic Customer Interaction (𝜏³-Banking)
EngineeringAgentic Terminal Use (Terminal-Bench v4.0)Agentic Terminal Use (Terminal-Bench v2.1); Reasoning (GPQA Diamond)
EconomicsAgentic Knowledge Work (AA-Briefcase)None

Changes by Index

Finance & Accounting

Agentic Tool Use was added using the Finance slice of AutomationBench-AA. Agentic Customer Interaction was removed. AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work, while GDP.pdf was added alongside LCR under Long-Context.

CapabilityEvaluationsv1.0v1.1
Business KnowledgeAA-Omniscience30%30%
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase30%30%
ReasoningHLE20%20%
Agentic Tool UseAutomationBench-AA0%10%
Long-ContextLCR, GDP.pdf5%5%
Non-HallucinationAA-Omniscience5%5%

Strategy & Ops

Agentic Tool Use was added using the Operations slice of AutomationBench-AA to represent tool use in day-to-day operational tasks. Agentic Customer Interaction was removed. AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work, and GDP.pdf was added alongside LCR under Long-Context.

CapabilityEvaluationsv1.0v1.1
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase35%35%
Business KnowledgeAA-Omniscience30%30%
Agentic Tool UseAutomationBench-AA0%30%
Long-ContextLCR, GDP.pdf5%5%

Legal

Agentic Tool Use was added using the Operations and Support slices of AutomationBench-AA to represent tool use in legal workflows. Agentic Customer Interaction was removed. AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work, and GDP.pdf was added alongside LCR under Long-Context.

CapabilityEvaluationsv1.0v1.1
Legal KnowledgeAA-Omniscience35%35%
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase25%25%
ReasoningHLE15%15%
Long-ContextLCR, GDP.pdf10%10%
Non-HallucinationAA-Omniscience10%10%
Agentic Tool UseAutomationBench-AA0%5%

Healthcare & Medical

Long-Context Reasoning was added using MLCR-AA, which evaluates reasoning across lengthy clinical records. Agentic Tool Use was also added using the Operations and Support slices of AutomationBench-AA to cover operational healthcare workflows. Agentic Customer Interaction was removed, and AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work.

CapabilityEvaluationsv1.0v1.1
Medical & Health KnowledgeAA-Omniscience35%30%
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase25%25%
Long-Context ReasoningMLCR-AA0%15%
Non-HallucinationAA-Omniscience15%10%
ReasoningHLE15%10%
Agentic Tool UseAutomationBench-AA0%10%

Engineering

Agentic Terminal Use was updated to Terminal-Bench v4.0. GPQA Diamond was removed from Reasoning, and AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work.

CapabilityEvaluationsv1.0v1.1
Engineering KnowledgeAA-Omniscience35%35%
ReasoningHLE, CritPt35%30%
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase25%20%
Agentic Terminal UseTerminal-Bench 4.05%15%

Economics

AA-Briefcase was added alongside GDPval-AA v2 under Agentic Knowledge Work.

CapabilityEvaluationsv1.0v1.1
Economics KnowledgeAA-Omniscience35%35%
ReasoningHLE35%35%
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase15%25%
Long-Context ReasoningLCR15%5%