MLPerf Client v2.0 Adds Image Generation and Agentic AI Tests for AI PCs
MLCommons, an open engineering consortium focused on machine learning performance and transparency, has released MLPerf Client v2.0, the latest version of its benchmark for eval...
By Software Development Team
MLCommons, an open engineering consortium focused on machine learning performance and transparency, has released MLPerf Client v2.0, the latest version of its benchmark for evaluating AI performance on personal computers.
MLPerf Client measures how effectively laptops, desktops, and workstations run AI workloads locally. Version 2.0 expands the benchmark beyond its existing large language model testing to include image generation and agentic AI. It evaluates practical tasks such as summarization, content creation, and code analysis, while reporting both responsiveness and throughput.
New in MLPerf Client v2.0
Version 2.0 adds new workloads and updates existing tests to reflect changes in AI hardware and software.
Image Generation
A new Image Generation category evaluates generative visual capabilities. The category includes Flux.2 klein 4B as an experimental test.
Agentic AI
The new Agentic AI category measures performance in two scenarios:
- Software Engineering (SWE) Agent
- Data Analyst Agent
The results cover end-to-end performance and include separate measurements for large language model inference and tool execution.
Updated LLM Inference Tests
The required LLM workloads now use Phi 4 Mini Instruct instead of Phi 3.5 mini instruct. Qwen 3 8B has also been added as an experimental test.
The base tasks now include an Intermediate Summarization workload with an input prompt of approximately 4K tokens.
These changes expand MLPerf Client's coverage as a cross-platform benchmark for client AI computing.
Development and Availability
MLPerf Client was developed through collaboration among AMD, Intel, Microsoft, NVIDIA, Qualcomm Technologies, Inc., and leading PC OEMs. The participating organizations contributed resources and technical expertise to the benchmark.
MLPerf Client v2.0 is available for download, and its source code can be inspected and extended through the MLCommons GitHub repository. Releases are provided for Windows, macOS, and Linux.
MLCommons is an open engineering consortium that develops benchmarks, datasets, and best practices for machine learning applications, ranging from cloud training to resource-constrained edge devices. Its MLPerf benchmark suite covers multiple areas of AI performance evaluation.