Skip to main content
Back to Blog
AI/MLData AnalysisInnovation
22 September 20263 min readUpdated 23 September 2026

MLCommons Joins EU-Funded AIRIS Project for Biomedical AI Development and Benchmarking

MLCommons joins the AIRIS research consortium MLCommons has joined AIRIS, a research project funded through the European Union’s Horizon Europe Program. AIRIS stands for Mechani...

By Software Development Team

MLCommons joins the AIRIS research consortium

MLCommons has joined AIRIS, a research project funded through the European Union’s Horizon Europe Program. AIRIS stands for Mechanism-Informed Multimodal Generative AI for Causal and Dynamical Modeling in Biomedical Research.

The consortium brings together 21 partners from Europe, Canada, and the United States. Over four years, it will receive €16.9 million from Horizon Europe to develop generative AI models that combine biological knowledge with clinical data.

Modeling disease mechanisms

AIRIS aims to create an AI collaborator that can help researchers understand disease progression and support advances in personalized medicine. Instead of relying only on statistical patterns, the planned platform will build and reason with mechanistic models of disease.

The project is intended to help scientists identify previously unknown disease pathways and develop new scientific hypotheses. Its broader objective is to support research into complex diseases by grounding AI reasoning in biological knowledge rather than producing only black-box predictions.

MLCommons’ benchmarking role

As an organization focused on AI benchmarking, MLCommons will lead the development of an evaluation framework and integrated benchmark suite for AIRIS. The framework will assess the platform across several dimensions:

  • Accuracy
  • Robustness
  • Fairness
  • Interpretability
  • Usability

The consortium will also create dedicated benchmarks for each of five disease areas. These evaluations will examine potential bias across patient subgroups, including differences associated with sex, ethnicity, and age.

“Rigorous, independent evaluation is what turns a promising AI system into one that researchers can actually trust,” said Alexandros Karargyris, Lead for the Medical working group at MLCommons. He said the open benchmarks would assess not only accuracy, but also fairness, robustness, and interpretability in real disease settings, providing a transparent standard for mechanism-informed biomedical AI.

MLCommons will coordinate AIRIS’ iterative evaluation rounds and track performance across successive versions of the platform. The benchmarks will also be made available to external researchers, allowing the wider scientific community to test AIRIS with their own data.

Through this process, the project aims to establish a reference benchmark for multimodal, mechanism-informed biomedical generative AI and provide a consistent way to evaluate future systems in the field.