Artificial Analysis Intelligence Index v4.3 Adds AutomationBench-AA and Upgrades Terminal-Bench
Artificial Analysis Intelligence Index v4.3 Adds AutomationBench AA and Upgrades Terminal Bench September 7, 2026 Artificial Analysis Intelligence Index v4.3 upgrades Terminal B...
By AI Engineering Team
Artificial Analysis Intelligence Index v4.3 Adds AutomationBench-AA and Upgrades Terminal-Bench
September 7, 2026
Artificial Analysis Intelligence Index v4.3 upgrades Terminal-Bench to version 4.0 and adds AutomationBench-AA, an agentic workflow automation benchmark with a private test set. The release brings forward selected changes originally planned for Intelligence Index v5.
Index results
Intelligence Index v4.3 combines 10 evaluations:
- AA-Briefcase
- GDPval-AA v2
- AutomationBench-AA
- Terminal-Bench v4.0
- SciCode
- Humanity's Last Exam
- GDP.pdf
- CritPt
- AA-Omniscience
- AA-LCR v1.1
The index uses the following category weights: Agents, 30%; Coding, 20%; General, 30%; and Scientific Reasoning, 20%.
Key results include:
- Claude Fable 5.1 and GPT-6 Astra lead the index. Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) each score 53. They are followed by Claude Opus 5 (max) at 51, Claude Fable 5 (with fallback) at 50, Muse Spark 1.3 (max) at 48, and GPT-5.6 Sol (max) at 47.
- GLM-5.3 and Kimi K3 remain among the leading open-weights models. GLM-5.3 Flash scores 42, followed by Qwen3.8 2.4T A95B at 40 and DeepSeek V4 Pro 0813 (max) at 36.
- Four labs appear on the Intelligence versus Cost per Task Pareto frontier. OpenAI occupies most of the cost-efficiency frontier. All five reasoning efforts of GPT-6 Astra provide the lowest Cost per Task at their respective intelligence levels. MiMo-V2.5-Pro at 26, GLM-5.3-Flash at 42, and Claude Fable 5.1 (xhigh, max) at 53 make up the remainder of the frontier.
Changelog
Terminal-Bench upgraded from v2.1 to v4.0
The index now uses the latest version of Terminal-Bench. The update increases the difficulty of agentic coding tasks and improves task instructions, environments, verification, and compute and time allowances.
AutomationBench-AA replaces 𝜏³-Banking
AutomationBench-AA is Artificial Analysis's implementation of Zapier's business workflow automation benchmark. It replaces 𝜏³-Banking in the index and adds broader workflows across business applications.
The v4.2 and v4.3 changes are intended to improve real-world problem solving, introduce more private test sets to limit gaming, and reduce score saturation. Because AutomationBench uses a held-out test set in collaboration with Zapier, evaluations with private tasks or answers now account for 45% of the index, up from 40% in v4.2.
The category weights remain unchanged:
- Agents: 30%
- Coding: 20%
- General: 30%
- Scientific Reasoning: 20%
Evaluation updates
Terminal-Bench v4.0
Terminal-Bench v4.0 tests whether an agent can complete complex tasks through a terminal in areas including software, machine learning, science, operations, security, hardware, and media. Artificial Analysis runs all 66 tasks three times and reports the average pass@1 result.
GPT-6 Astra (max) scores 59.1%, compared with 52.0% for Claude Fable 5.1 (max with fallback) and 49.0% for Claude Opus 5 (max). Astra leads GPT-5.6 Sol (max), which scores 39.9%, by 19.2 percentage points.
AutomationBench-AA
AutomationBench-AA uses Zapier's benchmark and a held-out test set of 657 tasks from version 1.0.6. The tasks cover Finance, HR, Marketing, Operations, Sales, and Support. Agents work across simulated business applications and discover the relevant APIs needed to complete each workflow.
The benchmark reports two measures:
- Score: The share of objectives completed across all tasks. A task receives zero if the agent violates any guardrail.
- Tasks Completed: The share of workflows in which every objective is completed without a guardrail violation.
GPT-6 Astra (max) scores 68.5%, compared with 66.7% for Grok 4.6 (high) and 62.2% for GLM-5.3 (max). Astra completes every objective without a guardrail violation on 41.6% of workflows. The corresponding figures are 32.1% for Claude Fable 5.1 (max with fallback) and 28.3% for Claude Opus 5 (max).
These results show that completing part of a workflow is easier than completing every objective while respecting all guardrails.
Cost differences
Models with similar Intelligence Index scores can have substantially different task costs. GPT-6 Astra (max) and Claude Fable 5.1 (max with fallback) both score 53, but their average costs per Intelligence Index task are $3.26 and $7.63, respectively. Astra's cost is 57% lower.
GLM-5.3-Flash and GPT-5.6 Terra (max) both score 42. GLM-5.3-Flash costs 18% as much per task, at $0.25 compared with $1.40. GPT-5.6 Luna (max) scores 38 at $0.18 per task.
Each evaluation's cost is calculated from input, cache-hit, cache-write, reasoning, and answer-token prices. The result is divided by the task count and weighted according to the evaluation's Intelligence Index weight.
Evaluations and weights
The overall scores reflect different model strengths. Claude Fable 5.1 scores higher on AA-Briefcase and SciCode, while GPT-6 Astra scores higher on Terminal-Bench v4.0 and the AutomationBench-AA Score.
Terminal-Bench 4.0 retains the weighting used by Terminal-Bench 2.1. AutomationBench-AA replaces 𝜏³-Banking at a 5% weighting.
| Category | Evaluation | Private test set | Weight |
|---|---|---|---|
| Agents, 30% | AA-Briefcase | Yes | 15% |
| Agents, 30% | GDPval-AA v2 | No | 10% |
| Agents, 30% | AutomationBench-AA, new in the index | Yes | 5% |
| Coding, 20% | Terminal-Bench v4.0, new in the index | No | 10% |
| Coding, 20% | SciCode | No | 10% |
| General, 30% | AA-Omniscience, Accuracy | Yes | 10% |
| General, 30% | AA-Omniscience, Non-hallucination | Yes | 5% |
| General, 30% | GDP.pdf | No | 10% |
| General, 30% | AA-LCR v1.1 | No | 5% |
| Scientific Reasoning, 20% | HLE, Humanity's Last Exam | No | 10% |
| Scientific Reasoning, 20% | CritPt | Yes | 10% |
| Total | 100% |
Private questions or answers account for 45% of the Intelligence Index v4.3 weighting, compared with 40% in v4.2.