Where the Industry Is Investing: A Look at MLPerf Inference v6.1

MLCommons ·

Where the Industry Is Investing: A Look at MLPerf Inference v6.1

In this analysis, MLPerf Inference Working Group chairs Miro Hodak and Frank Han break down a record field of 30 submitters and 120 systems, the arrival of agentic and end-to-end benchmarks, and what the results signal about where inference engineering is headed. The post Where the Industry Is Investing: A Look at MLPerf Inference v6.1 appeared first on MLCommons .

Every round of MLPerf ® Inference is a snapshot of where the industry is investing its engineering energy, and v6.1 stands out on two fronts: it is the broadest field of submitters we have ever seen , and it marks a clear inflection toward agentic and end-to-end benchmarking alongside a wave of newly submitted hardware .

The highlights of this round include a record number of submitters, new accelerators that significantly improve per-device performance, and a new multi-turn benchmark. Read on for our results analysis.

This round brought a record 30 submitters spanning silicon vendors, system builders (OEMs and ODMs), cloud and neocloud providers, and specialized inference-software companies. Several results this round were joint submissions , where two organizations collaborated: Dell_AMD, Dell_MangoBoost, RedHat_Intel, and RedHat_Supermicro.

The table below provides a breakdown of the systems and scores submitted by each organization:

Altogether, 120 systems across the Datacenter and Edge suites, in both the Closed and Open divisions, were submitted. Benchmarks MLPerf Inference benchmarks are divided into Datacenter and Edge categories. Each benchmark supports several scenarios, which may be one or more of: Offline, Server, Interactive, SingleStream, or MultiStream. Submissions can also be classified as either Closed or Open. Closed submissions require using a model that is mathematically equivalent to the reference implementation, essentially holding the model fixed to enable direct comparisons. For more details on the MLPerf Inference benchmark setup, visit https://mlcommons.org/benchmarks/inference-datacenter/ .

MLPerf Inference v6.1 features 10 Datacenter and 6 Edge benchmarks, two of which are new: End-to-End RAG and Agentic Edge Inference. Two benchmarks have been updated: For VLM (Visual Language Model based on Qwen3), an Interactive scenario has been newly defined, and in GPT-OSS-120B, the interactive scenario has been updated to allow speculative decoding.

Bar length is scaled within each suite, so Datacenter and Edge bars are not comparable to one another.

MLPerf Inference submitters can choose to submit to any of the benchmarks. In the Datacenter category, the most popular benchmark is gpt-oss-120b, followed by Llama2-70b. This is the first time gpt-oss-120b has topped the list, after Llama2-70b was the most popular for the last few rounds. This demonstrates that the MLPerf benchmarking community (which highly values stability and comparability) is now fully embracing MoE models.

To gauge performance improvements since the last round, we compared the best per-accelerator Offline and Server scores against those from v6.0.

Percentages are as reported in the results analysis. Bars are scaled per model, so lengths are comparable only within a row. rgat’s best Offline score sits marginally below v6.0.

The biggest gains are in VLM and DeepSeek R1 for both Offline and Server scenarios, significantly surpassing gains in other benchmarks. This is because those 2 benchmarks were submitted on a new Preview category system powered by Nvidia Vera Rubin. Other benchmarks’ top scores were achieved on the same hardware as in 6.0, so results reflect more gradual performance improvements from software stack and algorithmic optimizations.

Llama2-70b is the longest-running LLM in the MLPerf Inference suite, being introduced in early 2024, round v4.0. As such, it is a strong vehicle for showing LLM performance improvements over time. This is shown below in median per-accelerator performance for Server scenario submissions, which has improved 5.58x over 6 runs. Three main reasons contribute to the performance gains:

A similar trend appears in the DeepSeek R1 benchmark, despite being around for less time. The figure below shows that best per-accelerator performance has increased by 2.7x in Offline and 5.7x in Server scenarios, respectively, within a year. Similarly, the Interactive scenario, which had only 2 submission rounds, increased by 2.7x over that period.

Beyond performance improvements, results also contain additional important achievements. One continued trend is multi-node inference. In the last few rounds, the number of multi-node submissions has increased considerably, as shown in the figure below:

The upward trend started in round 4.1, 2 years ago, and this round reached an all-time record of 16 multi-node submissions. We also see the size of submitted systems increasing. Just last round, a record of 288 accelerators was set, only to be broken in this round by Cruose submissions 6.1-0026 and 6.1-0027, which used 512 accelerators. The latter of the submissions set a new record for the number of tokens/second generated in the MLPerf Inference benchmark by generating almost 5.8M tokens/second in GPT-OSS-120b offline test.

Another trend in MLPerf Inference is hybrid submissions that use different types of accelerators working together. These systems pose unique challenges because submitters must account for the different computational capabilities of the devices. This round contains two such submissions:

v6.1 introduced a diverse set of newly submitted accelerators and systems — from on-device parts to a next-generation rack-scale platform — spanning multiple accelerator vendors and different scales of deployment.

Many participating organizations submitted supplemental statements describing what they consider most significant about their v6.1 submission. Full statements are available in the results repository ; here are the highlights.

MLPerf Inference v6.1 is one of the most expansive rounds in the benchmark’s history: 30 organizations submitted 120 systems across the Datacenter and Edge suites and both the Closed and Open divisions.

Just as notable as the volume is the range — from single-accelerator edge devices to rack-scale platforms with hundreds of accelerators, all measured under one consistent methodology. A few clear signals emerge.

For the first time, GPT-OSS-120B is the most popular model, with DeepSeek-R1 and Qwen3-VL also drawing high participation.

A shift from single-shot benchmarking to agentic model evaluation.

From datacenter to edge devices, contributing to the best-ever per-accelerator scores seen this round.

The first-ever heterogeneous submission combining accelerators from different vendors, plus geographically distributed submissions spanning continents.

This round set a record for the largest-scale submission ever, at 512 GPUs — continued interest in scale-out inference deployments.

As always, these results are most valuable when read in context: the Closed division remains the foundation for apples-to-apples comparison, while the Open division surfaces the optimization techniques that often preview the mainstream. We encourage readers to consult both the results tables and the submitter statements.

A round like this reflects an enormous collective effort. Our thanks go to every submitting organization — established vendors and first-time participants alike — and to the many Inference Working Group members who build, review, and audit these benchmarks and keep MLPerf rigorous, relevant, and fair.

The post Where the Industry Is Investing: A Look at MLPerf Inference v6.1 appeared first on MLCommons .

Источник: MLCommons