MLCommons Joins EU-Funded AIRIS Project for Biomedical AI Development and Benchmarking
MLCommons joins the AIRIS research consortium MLCommons has joined AIRIS, a research project funded through the European Union’s Horizon Europe Program. AIRIS stands for Mechani...
By Software Development Team
MLCommons joins the AIRIS research consortium
MLCommons has joined AIRIS, a research project funded through the European Union’s Horizon Europe Program. AIRIS stands for Mechanism-Informed Multimodal Generative AI for Causal and Dynamical Modeling in Biomedical Research.
The consortium brings together 21 partners from Europe, Canada, and the United States. Over four years, it will receive €16.9 million from Horizon Europe to develop generative AI models that combine biological knowledge with clinical data.
Modeling disease mechanisms
AIRIS aims to create an AI collaborator that can help researchers understand disease progression and support advances in personalized medicine. Instead of relying only on statistical patterns, the planned platform will build and reason with mechanistic models of disease.
The project is intended to help scientists identify previously unknown disease pathways and develop new scientific hypotheses. Its broader objective is to support research into complex diseases by grounding AI reasoning in biological knowledge rather than producing only black-box predictions.
MLCommons’ benchmarking role
As an organization focused on AI benchmarking, MLCommons will lead the development of an evaluation framework and integrated benchmark suite for AIRIS. The framework will assess the platform across several dimensions:
- Accuracy
- Robustness
- Fairness
- Interpretability
- Usability
The consortium will also create dedicated benchmarks for each of five disease areas. These evaluations will examine potential bias across patient subgroups, including differences associated with sex, ethnicity, and age.
“Rigorous, independent evaluation is what turns a promising AI system into one that researchers can actually trust,” said Alexandros Karargyris, Lead for the Medical working group at MLCommons. He said the open benchmarks would assess not only accuracy, but also fairness, robustness, and interpretability in real disease settings, providing a transparent standard for mechanism-informed biomedical AI.
MLCommons will coordinate AIRIS’ iterative evaluation rounds and track performance across successive versions of the platform. The benchmarks will also be made available to external researchers, allowing the wider scientific community to test AIRIS with their own data.
Through this process, the project aims to establish a reference benchmark for multimodal, mechanism-informed biomedical generative AI and provide a consistent way to evaluate future systems in the field.