How UK AISI and EvalEval Are Making Benchmark Results Reproducible
How UK AISI and EvalEval Are Making Benchmark Results Reproducible Published September 22, 2026 The EvalEval Coalition and the UK AI Security Institute (AISI) are using EvalEval...
By AI Engineering Team
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
Published September 22, 2026
The EvalEval Coalition and the UK AI Security Institute (AISI) are using EvalEval's infrastructure to share evaluation results openly, with the goal of making evaluation research more reproducible and verifiable.
The organisations previously collaborated on research that began at a joint workshop held alongside NeurIPS 2025. Feedback from AISI also helped shape the Every Eval Ever (EEE) schema. This phase of the partnership applies that shared infrastructure to published evaluation results.
Why reproducible evaluation reporting matters
As AI deployment expands, evaluations are increasingly important sources of evidence about model and system performance. However, results are reported across many formats, platforms, and publications, often without enough information for others to reproduce them. Repeating an evaluation can also be prohibitively expensive.
EvalEval addresses this problem through Every Eval Ever, a shared reporting schema, and Evaluation Cards, an open platform that places evaluation results and the information needed to interpret them in a common structure.
The work complements AISI's efforts to improve evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas such as transcript analysis and capability elicitation. Together, AISI and EvalEval are examining gaps in evaluation reporting and developing shared infrastructure to address them.
What AISI is sharing
Transcript-level transparency supports reproducibility, analysis, and diagnosis. In this phase of the collaboration, AISI is making publicly reported evaluation methods and findings available through Evaluation Cards where appropriate.
The release includes verified results, context, and configuration information for five benchmarks used in the paper's main experiment:
- HealthBench
- FrontierMath
- Humanity's Last Exam
- SWE-Bench Pro
- Terminal-Bench 2.0
The results cover six frontier models:
- Claude Opus 4
- Claude Opus 4.5
- Claude Opus 4.6
- GPT-5
- GPT-5.2
- GPT-5.4
The release also contains results from two related cyber evaluations, Cyber CTFs and The Last Ones. These evaluations use a different, partially overlapping set of models.
The data accompany AISI's paper, How Inference Compute Shapes Frontier LLM Evaluation, which examines how benchmark performance depends on inference-time compute and evaluation protocol.
Performance on Humanity's Last Exam varies according to evaluation protocol and inference compute. The paper's curves show the cumulative share of attempted tasks solved within a given token count, using the earliest observed success for each task. When models received correctness feedback from an oracle after every attempt, they continued solving additional tasks as token use increased.
Open releases that include setup information allow researchers and practitioners to inspect individual studies more closely and compare findings across the broader evaluation ecosystem. When other reports omit comparable details, releases such as AISI's can provide verified reference points for interpreting results in context, including understanding how setup decisions may affect reported performance.
As more evaluators adopt EEE, these openly documented comparisons can support broader and more reliable meta-research.
AISI's Terminal-Bench 2.0 results can also be examined alongside other reported evaluations of the same models conducted under different setups. This helps show how evaluation conditions can affect apparently comparable results.
About the EvalEval Coalition
The EvalEval Coalition is a research community developing scientifically grounded research and deployment infrastructure for the evaluation ecosystem. Its goals include improving evaluation science, addressing the lack of consensus about documenting evaluation applicability and utility, and expanding coverage of impacts relevant to scientific research and policy analysis.
The coalition's main projects include Every Eval Ever, a shared schema and repository for evaluation results, and Evaluation Cards, which combines benchmark metadata, evaluation-run data, and model metadata into interpretable records. Together, these projects help distinguish scores produced under meaningfully different conditions, even when the evaluations appear similar.
About the UK AI Security Institute
The UK AI Security Institute is a research organisation within the UK government's Department for Science, Innovation and Technology. Its mission is to provide governments with a scientific understanding of the risks posed by advanced AI.
AISI conducts research and develops infrastructure to study advanced AI capabilities and impacts, develop and test mitigations, and inform policy.