Artificial Analysis Launches Cyber Index Alliance for Benchmarking Agentic Cyber Defense
Artificial Analysis Launches Cyber Index Alliance for Benchmarking Agentic Cyber Defense The Artificial Analysis Cyber Index Alliance brings together industry partners to establ...
By AI Engineering Team
Artificial Analysis Launches Cyber Index Alliance for Benchmarking Agentic Cyber Defense
The Artificial Analysis Cyber Index Alliance brings together industry partners to establish a standard for evaluating how AI models perform enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to measure how effectively agents find and fix vulnerabilities.
Launch partners
- Collinear AI, developer of CWE-bench and contributor of a private held-out evaluation
- IBM, providing expert input on scope and methodology
- NVIDIA, providing expert input on scope and methodology
- Vercel, developer of DeepsecBench and contributor of a private held-out evaluation
The Artificial Analysis Cyber Index
Organizations that depend on software continually face vulnerabilities. Security teams must find and fix those weaknesses before attackers exploit them, a challenge that affects enterprises, governments, NGOs, and downstream users.
As AI models become more capable in cybersecurity, attackers can automate more work and operate at greater scale. Security teams therefore need tools that can detect and patch vulnerabilities effectively and cost-efficiently.
The Artificial Analysis Cyber Index is designed to help decision-makers assess models for security work. It measures how well models can support and accelerate an internal security team's activities by combining evaluation datasets from industry partners and academic research.
The Index evaluates the defensive loop:
- Discovering vulnerabilities in a codebase
- Reproducing and validating them
- Patching them without breaking existing functionality
Models work from source code, as a security engineer would when auditing an application. The evaluations do not ask models to build working exploits. The result is a like-for-like comparison of model performance and operating cost for defensive security tasks.
Artificial Analysis Cyber Index v1 combines three evaluations:
- CWE-Bench-AA
- DeepsecBench-AA
- CyberGym-E2E-AA
The Cyber Index Alliance
The Cyber Index Alliance supports the development of a standard for evaluating AI models on enterprise cyber defense tasks. Partners contribute expertise to the Index's design and implementation and may provide datasets or external research.
Organizations interested in joining the Alliance can contact cyber@artificialanalysis.ai.
How the Cyber Index works
Benchmark overview
At launch, the Index combines three cybersecurity evaluations from industry partners and academic research. Together, they cover the defensive process from scanning code for weaknesses to reproducing crashes and implementing patches.
| Evaluation | Contributor | What it measures | Coverage |
|---|---|---|---|
| CWE-Bench-AA | Collinear AI | Auditing a real open-source repository for a vulnerability in a described area, then patching it without breaking legitimate behavior | 120 held-out tasks covering all ten OWASP Top 10 (2025) categories; identifies and remediates vulnerabilities |
| DeepsecBench-AA | Vercel | Finding vulnerabilities in open-source application code, scored against expert-verified findings | Identifies vulnerabilities |
| CyberGym-E2E-AA | Berkeley RDI | Discovering, reproducing, and patching memory-safety vulnerabilities in C/C++ open-source projects | 131 tasks, one per project, drawn from the 920-instance CyberGym-E2E dataset; identifies and remediates vulnerabilities |
Each evaluation is tagged with the capabilities and sub-capabilities it tests. This shows which parts of cyber defense the Index covers and where gaps remain.
The current Index covers identifying and remediating vulnerabilities with access to source code. Planned additions include incident response, writing new code without introducing vulnerabilities, and targets without source access, such as compiled software and live servers. Exploit realization, meaning the process of turning a discovered vulnerability into a working exploit, remains outside the scope of this defense-focused index.
Methodology
All three evaluations run on Stirrup, an open-source agent harness.
Because cybersecurity work is dual-use, the evaluations also track cases where a model or provider declines a task on safety grounds. These cases are reported separately from performance scores. CyberGym-E2E-AA allows a model to end a task without a finding if it concludes that it cannot find or demonstrate a vulnerability.
CWE-Bench-AA: auditing and patching real repositories
CWE-Bench-AA is an implementation of Collinear AI's CWE-bench. Each task is a security audit in which an agent receives a checkout of an open-source repository and instructions to audit the code and fix what it finds. Collinear AI includes tasks that reproduce disclosed CVEs, or Common Vulnerabilities and Exposures. The task identifies the area of concern but not the exact location of the vulnerability.
The test set contains 120 held-out tasks private to Collinear AI and Artificial Analysis. The tasks cover all ten OWASP Top 10 (2025) categories across C/C++, Go, Java, JavaScript/TypeScript, Python, and Rust. Each task tests whether the model can find the weakness and patch it without breaking the rest of the codebase.
CWE-Bench-AA scoring
The benchmark reports the share of tasks solved, failed, or declined on safety grounds, along with vulnerability-patching success using pass@1.
A model receives credit for a task only when the programmatic verifier confirms that the exploit no longer works and legitimate behavior still works. There is no partial credit and no LLM judge. Tasks are private, run in a sandbox without internet access, and identify the area of concern without specifying the exact location.
How models fail
CWE-Bench-AA failures are concentrated in two areas:
- Partial fixes, which leave the vulnerability exposed
- Over-corrections, where a patch breaks legitimate behavior
On average, models spent 38% of their turns searching for the bug before making their first edit. The remaining 62% was spent patching and validating the fix.
Excluding refusals and timeouts, partial fixes accounted for 55% of failed attempts. These attempts fixed the primary issue but left a related weakness open, such as a second entry point.
Over-corrections accounted for approximately 24% of failed attempts. They were most common among the strongest models, representing approximately 40% of failures for the four highest-scoring models compared with approximately 15% for the lowest performers. The resulting breakage was usually a close edge case of legitimate functionality.
DeepsecBench-AA: finding expert-confirmed vulnerabilities
DeepsecBench-AA is an implementation of Vercel's DeepsecBench and focuses on vulnerability discovery. The agent reviews scanner-flagged files in open-source application code and reports every vulnerability it can confirm. Findings are scored against a golden set verified by human experts.
Security teams face a similar open-ended problem when deciding where to look in a large, changing codebase. Agents may report findings outside the golden set, including some that are real, and security teams must triage each one. Scoring against the golden set rewards models that identify real vulnerabilities without overwhelming reviewers with false positives.
DeepsecBench-AA scoring
The benchmark reports the share of tasks solved, failed, or declined on safety grounds, along with an F2 score. F2 weights recall above precision.
A judge model assesses whether each reported finding is real and matches the golden set. Duplicate findings count against precision. The review runs three times, and the headline result is the median F2 score.
How models fail
Because DeepsecBench-AA is open-ended, models do not find every vulnerability. The best model identified only 41% of the expert-verified issues.
Models most often find flaws with a direct path from untrusted input to a consequence. Bugs that require reasoning through a sequence of events or applying business and privacy rules are rarely found.
When models do report sequence-of-events vulnerabilities, they are usually correct: 95% of those reports are accurate. GPT-6 Sol and GPT-6 Astra found these vulnerabilities in approximately 30% of runs, about twice the rate of the next-best models and three to four times the typical model. This suggests that the capability may be emerging.
CyberGym-E2E-AA: discovering, reproducing, and patching memory-safety bugs
CyberGym-E2E, developed by the Berkeley Center for Responsible, Decentralized Intelligence, evaluates the end-to-end process of discovering vulnerabilities, reproducing them, and writing patches that pass existing project tests. The benchmark focuses on memory-safety bugs.
Each task targets a real memory-safety vulnerability in a widely used C/C++ open-source project, such as FFmpeg or CPython. The model must locate the bug, write a proof-of-concept input that triggers the crash, and patch the code so the crash no longer occurs.
CyberGym-E2E-AA uses a filtered set of 131 tasks, with one task per project. Ending a task without a finding scores zero and is recorded separately from a failed submission.
CyberGym-E2E-AA scoring
The benchmark reports the share of tasks solved, failed, or declined on safety grounds, along with vulnerability discovery and patching success using pass@1.
A task counts as solved on a single attempt only when:
- The proof-of-concept crashes the unpatched build
- The patch fixes that crash
- The patched project continues to pass its functionality tests
These are stages 1 to 3. Whether the patch also fixes the ground-truth vulnerability is recorded but does not count toward the score.
How models fail
Refusals are a significant factor in CyberGym-E2E-AA. GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Opus 5.5, Qwen3.8 2.4T A95B, and Qwen3.8 27B refuse at least 98% of tasks, making frontier-model performance difficult to assess.
Among the remaining models, failures follow three main patterns:
- Failure to discover the vulnerability: 42% of attempts reached the 90-minute limit without producing an input that crashes the program. This is relevant to enterprises deciding how broadly to scan their codebases.
- Difficulty with state-dependent bugs: Models passed 50% of attempts on out-of-bounds bugs, compared with 33% on use-after-free bugs, which depend on an object's lifetime across several operations, and 20% on integer and arithmetic bugs.
- Patching the wrong bug: 31% of passing attempts patched a real crash other than the target. These were generally shallower issues, such as null-pointer crashes. Such crashes accounted for 22% of off-target passes, compared with 9% of on-target passes, because models often stopped at the first crash they could validate and completed the task within a couple of turns.
Artificial Analysis Cyber Index resources
- Live model scores and refinements are available on the Artificial Analysis Cyber Index page.
- The methodology page documents the scope, grading, and implementation of each benchmark.
- Individual leaderboards are available for CWE-Bench-AA, DeepsecBench-AA, and CyberGym-E2E-AA.
- Stirrup, the open-source agent harness used for every Artificial Analysis Cyber Index task, is available on GitHub.
Development roadmap
Security teams are already selecting models to support defensive work. The Artificial Analysis Cyber Index and Cyber Index Alliance therefore launch with three evaluations rather than waiting to cover every cybersecurity capability.
The Index will be reviewed as models improve, and additional evaluations will be added to measure capabilities that are not currently covered, including those identified in the roadmap.