Skip to main content
Back to Blog
AI/MLSecurityEnterprise
26 August 20263 min readUpdated 27 August 2026

Why Trustworthy AI Evaluation Requires Secrecy by Design

Organizations deploying AI need to determine whether a system performs reliably and safely with their specific use cases and data. That assessment requires more than connecting...

By AI Engineering Team

Organizations deploying AI need to determine whether a system performs reliably and safely with their specific use cases and data. That assessment requires more than connecting the relevant systems and running a test.

Deployers such as banks must protect sensitive information. AI providers, including frontier laboratories, must protect their model weights. Benchmark providers face an additional responsibility: they need technical and legal safeguards to ensure that tested models cannot learn from evaluation materials and undermine benchmark integrity.

Keeping an evaluation secret is not enough on its own. A robust benchmark stewardship program is also necessary to maintain a high-integrity benchmark over time.

A double-blind evaluation

To demonstrate this approach to benchmark integrity, MLCommons participated in a proof of concept with Google DeepMind, OpenMined, and AVERI. The project was the first double-blind evaluation of a closed-weight model. Its design protected every participating party through technical controls rather than relying only on legal contracts.

The method provided cryptographic guarantees that the evaluation components could be used without contaminating the benchmark or compromising its integrity.

AILuminate™ was selected for the proof of concept. MLCommons supplied a reserved subset of prompts from the AILuminate™ safety benchmark. Because this subset had been held back, no Google DeepMind model had previously been exposed to that specific evaluation set.

AVERI used OpenMined’s secure computation technology to run the prompts on a containerized instance of a Google DeepMind model. The process included cryptographically provable protections for both the model weights and the evaluation materials.

In this proof of concept, AILuminate measured the AI system’s reliability without revealing the test data to the developer. It also kept the developer’s proprietary model weights hidden from MLCommons, the benchmark instrument maker, and AVERI, the auditor.

Applying the approach to sensitive data

This type of intellectual-property-safe and saturation-resistant testing could support applied evaluations of AI models. Similar work has already been conducted with highly sensitive health informatics data through the MedPerf program.

MedPerf has been developing Trusted Execution Environments (TEEs) that allow multiple parties to bring encrypted data together, conduct high-integrity evaluations, and receive results without exposing one party’s sensitive information to another.

Toward purpose-specific AI evaluations

The goal is a near future in which AI adopters can quickly run trustworthy evaluations for purpose-specific deployments. An environment that structurally protects model weights, data, and tests, combined with independently supplied prompts that models have not previously encountered, can provide a strong basis for decisions about AI adoption, deployment, and reliable operation.

MLCommons is developing this approach through the AILuminate framework and its Benchmark Stewardship Program.