Skip to main content
Back to Blog
AI/MLEnterprise
24 August 20263 min readUpdated 24 August 2026

MLPerf Endpoints v0.7 Establishes a Foundation for Dynamic AI Inference Benchmarks

MLPerf expands its benchmarking focus Since its launch in 2018, MLPerf has measured the performance of AI systems. During that period, it recorded more than a 100X improvement i...

By AI Engineering Team

MLPerf expands its benchmarking focus

Since its launch in 2018, MLPerf has measured the performance of AI systems. During that period, it recorded more than a 100X improvement in inference performance per watt for large language models and more than a 50X improvement in training speed [1][2].

As AI services have become part of daily operations for enterprises and consumers worldwide, the requirements for evaluating AI infrastructure have also changed. MLPerf Endpoints is intended to address those requirements with a benchmarking system focused on AI inference procurement.

MLPerf was initially created to help a relatively small number of cloud providers and downstream system builders evaluate AI hardware purchases. Inference computing is now a significant purchasing decision for companies of all sizes. Buyers may need to evaluate neoclouds, cloud providers, and managed services at the same time, requiring independent and comparable performance data. The benchmarks must also be updated frequently to reflect an industry that introduces new models each week.

Four principles for MLPerf Endpoints

MLPerf Endpoints is being developed around four principles for enterprise-focused benchmarking:

  1. Current: Results should keep pace with the market. Buyers should not have to wait months for new hardware or models to appear in a benchmark round.
  2. Comprehensive: Benchmarks should cover the competing inference providers, systems, and workloads available for purchase.
  3. Comparable: Results should support apples-to-apples comparisons across vendors, including normalization for cost or power.
  4. Commentary: Visualizations, filtering, and analysis should provide additional context for interpreting benchmark results.

The v0.7 foundation release

MLPerf Endpoints v0.7 is a foundation release containing initial results from Coreweave, Google, Intel, KRAI, and Nvidia. The results span several orders of magnitude in performance across three benchmarks and provide the infrastructure for a more dynamic, comprehensive, and comparable inference benchmark suite for data centers.

The platform currently supports automated submission pipelines, continuous review tooling, and dynamic result visualizations. Its benchmarking rules are also being developed further to focus more directly on buyers' needs.

Planned v1.0 development

MLPerf Endpoints v1.0 is planned for later this year. It is expected to include more buyer-focused rules, normalization, and additional benchmarks, including agentic workloads. After that release, the rolling submission process will be opened to the broader MLPerf membership.

Rolling submissions are intended to keep results current and aligned with the pace of the AI infrastructure market.

More than 30 supporters have contributed to the development of MLPerf Endpoints, including AMD, Argonne National Laboratory, Broadcom, Core 42, Dell, HPE, Lambda, Oracle, and Red Hat. Their work has helped test the rules and processes and improve the supporting infrastructure.

References

[1] A. Tschand, A. T. R. Rajan, S. Idgunji, et al., “MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from µWatts to MWatts for Sustainable AI,” in 2025 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2025. arXiv:2410.12032.

[2] D. Kanter, M. Ahmad, H. Kassa, and S. Rishab, “MLPerf Training v4.1 Results - Press Briefing,” MLCommons, Nov. 13, 2024.