Hardware/Analysis

MLPerf Inference v6.1 Results: Vera Rubin, MI350P and Arc Pro B70

A record 30 submitters, a 512-GPU system and first verified numbers for Nvidia’s Vera Rubin: what the latest AI inference benchmark round actually shows.

A long row of NVIDIA rack-scale AI server cabinets seen from the front against a black background
Nvidia’s rack-scale systems, including the Vera Rubin NVL72 preview, featured heavily in MLPerf Inference v6.1. Image: Nvidia.

MLCommons published the MLPerf Inference v6.1 results on 16 September 2026, and the round is the biggest yet: a record 30 submitting organisations and 120 systems. It also gives the first peer-reviewed numbers for Nvidia’s Vera Rubin NVL72 (in preview), AMD’s new Instinct MI350P PCIe card, the Ryzen AI Max+ 395 and Intel’s Arc Pro B70. MLCommons says the best per-accelerator DeepSeek-R1 result in the Server scenario is 5.7 times better than a year ago.

Key facts

  • Published: 16 September 2026 by MLCommons
  • Scale: 30 submitters, 120 systems, 10 datacenter and 6 edge benchmarks
  • New hardware: AMD Ryzen AI Max+ 395, AMD Instinct MI350P, Intel Arc Pro B70 (available); Nvidia Rubin and Vera Rubin NVL72 (preview)
  • New tests: End-to-End RAG (datacenter) and Edge Agentic Inference
  • Biggest system ever: 512 accelerators, submitted by Crusoe
  • Headline gains: up to 5.7x on DeepSeek-R1 in one year; 2.99x on the VLM test in six months

What is MLPerf Inference?

MLPerf Inference is the industry’s main peer-reviewed benchmark for how fast hardware can run trained AI models. It is run by MLCommons, an open engineering consortium that says it has more than 130 members and affiliates. Chipmakers, server vendors and cloud providers submit results on a fixed set of models; other submitters review them before publication.

Results are split into two suites. The Datacenter suite measures throughput and latency for servers and racks, and the Edge suite measures single devices such as workstations, embedded modules and PCs. Each benchmark can be run in one or more scenarios: Offline (maximum batch throughput), Server (queries arriving under a latency limit), Interactive (tighter per-user latency), and the single-stream and multi-stream scenarios used at the edge.

There are also two divisions. The Closed division requires a model mathematically equivalent to the reference, so results can be compared directly. The Open division lets submitters change the model or technique, which is where new optimisations tend to appear first. When a vendor claims a win, it is worth checking which division, scenario and system size it is talking about.

What’s new in MLPerf Inference v6.1?

Two new tests reflect how AI is now being deployed, according to MLCommons:

  • End-to-End Retrieval-Augmented Generation (RAG): a pipeline of several models rather than a single one. An embedding model turns a query into a vector, a retriever pulls passages from a vector database, a re-ranker refines them and one or more language models produce the answer. It measures both building the vector database from a document corpus and answering questions against it.
  • Edge Agentic Inference: a multi-turn, single-user coding workload for edge devices, where each query depends on the conversation so far. It borrows its method from the upcoming MLPerf Agentic datacenter benchmark and adds latency metrics and an accuracy gate.

The round also adds support for speculative decoding, a common production technique that predicts and verifies several tokens in one forward pass, in the Interactive scenario for two benchmarks including GPT-OSS-120B. The VLM test, based on Qwen3, gains a new Interactive scenario.

Participation shows where the industry’s attention is. In the analysis written by the MLPerf Inference working group chairs, Miro Hodak (AMD) and Frank Han (Dell), GPT-OSS-120B became the most popular datacenter benchmark for the first time, with 112 results from 54 systems, overtaking Llama 2 70B. The chairs describe this as the benchmarking community “fully embracing” mixture-of-experts models.

Which new chips appeared this round?

MLCommons lists five processors or accelerators making their first appearance:

ChipVendorTypeStatus in v6.1
Vera Rubin NVL72NvidiaRack-scale system (72 GPUs)Preview
RubinNvidiaData center GPUPreview
Instinct MI350PAMDPCIe accelerator cardAvailable
Ryzen AI Max+ 395AMDPC processor with integrated GPU and NPUAvailable
Arc Pro B70IntelWorkstation graphics cardAvailable

“Preview” in MLPerf terms means the system was not yet generally available when it was submitted. Preview results go through the same peer review, but they are early numbers on hardware and software that are still maturing, so buyers should expect them to move.

Nvidia: Vera Rubin NVL72 debuts

Nvidia submitted Vera Rubin NVL72 preview results on two benchmarks: DeepSeek-R1 and Qwen3-VL. In its results blog post, the company says Vera Rubin NVL72 delivered up to 3.7 times the throughput of its current GB300 NVL72 rack on Qwen3-VL, and up to 2.5 times on DeepSeek-R1. Cloud provider Nebius also submitted Vera Rubin NVL72 preview results.

Those two submissions explain the round’s biggest headline numbers. According to the working group chairs, the largest per-accelerator gains, on the VLM and DeepSeek-R1 tests, came from the new Vera Rubin preview system. On other benchmarks, top scores were set on the same hardware as in v6.0, so improvements there came from software and algorithms.

Nvidia also highlighted scaling. Its DeepSeek-R1 submission grew from one GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs) with what Nvidia says was 99% scaling efficiency in the Offline scenario. It says software changes improved GB300 NVL72 performance on Qwen3-VL by up to 1.6 times compared with v6.0. Nvidia also submitted Jetson AGX Thor results on the new Edge Agentic benchmark.

Nvidia cites further gains from post-deadline software work and from a third-party agentic benchmark, but notes itself that the post-submission results have not been verified by MLCommons. Only the numbers published in the official results tables carry MLPerf’s peer review.

AMD: broadest submission yet, and the MI350P arrives

AMD calls v6.1 its broadest MLPerf Inference submission to date, expanding from three model families in v6.0 to six across the Instinct MI355X, MI350X and the new MI350P. The main points AMD makes:

  • More from the same hardware: on the same eight MI355X GPUs, AMD says ROCm and open-source software work raised GPT-OSS-120B throughput by 28% in Offline and 38% in Server, and improved Wan-2.2 text-to-video SingleStream performance by 70%, within one MLPerf cycle.
  • Scaling: going from 8 to 72 MI355X GPUs on GPT-OSS-120B kept 95% scale efficiency in both Offline and Server, according to AMD.
  • Head-to-head claims: AMD says eight MI355X GPUs led selected Nvidia B200 and B300 GPT-OSS-120B results, and that its 72-GPU result led Nvidia’s GB200 result. It says the MI350P led selected RTX PRO 6000 Server Edition and H200 NVL results in its first round.
  • Partners: AMD says the average of comparable partner MI355X results landed within 4% of its own.

The MI350P is the notable newcomer. AMD describes it as a dual-slot PCIe 5.0 card on the CDNA 4 architecture with 128 compute units, 144GB of HBM3E memory, 4TB/s of memory bandwidth and up to 4.6 petaflops of peak theoretical MXFP4/MXFP6 matrix performance. That is half the compute units and memory of the MI355X, which AMD lists at 256 compute units, 288GB of HBM3E and 8TB/s. The point of the MI350P is to let companies add AMD accelerators to ordinary PCIe servers rather than buying dense eight-GPU platforms.

AMD graphic showing an Instinct MI350 Series GPU board and the Instinct MI350P PCIe card
AMD’s MLPerf Inference v6.1 graphic, showing an Instinct MI350 Series board and the new MI350P PCIe card. Image: AMD.

Intel: Xeon 6 gains from software, Arc Pro B70 debuts

Intel’s story this round is mostly about software. In its supplemental statement, quoted by MLCommons, Intel says that on the same Xeon 6980P silicon and socket count as v6.0, Llama 3.1-8B Server throughput rose 2.4 times (+142%) and Offline throughput rose 56%, from software alone.

The Arc Pro B70, a workstation graphics card, appears for the first time as an available product. Supermicro’s statement confirms it submitted results for the Arc Pro B70 alongside Intel Xeon 6 and 6+ CPUs, Nvidia B300 and AMD MI355X systems. Specific Arc Pro B70 scores are listed in the MLCommons results tables.

What it means for PCs and workstations

Most coverage of MLPerf focuses on racks, but v6.1 has more for desktop and workstation buyers than usual. The new Edge Agentic Inference benchmark models one user running a coding agent on a local machine, with a growing conversation history, fixed memory and a single stream of requests. According to the chairs’ tally, five systems submitted to it in its first round.

Two of the round’s new chips sit in this space. AMD’s Ryzen AI Max+ 395 is a PC processor rather than a data center part, and Intel’s Arc Pro B70 is a workstation graphics card. Having a PC processor and a workstation card measured under the same peer-reviewed rules as data center racks gives buyers of local AI hardware a rare independent reference point, although the number of edge submissions is still small compared with the datacenter suite.

Records: 512 GPUs, cross-vendor and cross-ocean systems

Beyond the new chips, v6.1 set several firsts, according to MLCommons and the submitters:

  • Largest system ever: Crusoe submitted two results using 512 accelerators, beating the previous record of 288. One of them generated almost 5.8 million tokens per second on GPT-OSS-120B Offline, a new MLPerf record. AMD says these runs used MI355X GPUs and reached 5.75 million Offline and 5.39 million Server tokens per second on GPT-OSS-120B, plus 2.90 million Offline tokens per second on DeepSeek-R1.
  • First cross-vendor system: Cisco combined eight Nvidia H200 GPUs and eight AMD Instinct MI350X GPUs into one inference pool over a Cisco G200 network.
  • Geographically distributed system: Dell and MangoBoost combined 16 MI300X GPUs in Korea with 16 MI355X GPUs in the United States as one GPT-OSS-120B endpoint, which MLCommons describes as spanning the Pacific Ocean.
  • Multi-node growth: the chairs count 16 multi-node submissions, an all-time high.

The chairs also track Llama 2 70B, the longest-running language model in the suite. Median per-accelerator Server performance on that test has improved 5.58 times over six rounds, which they attribute to lower-precision formats (some v6.1 entries use FP4 where earlier rounds used FP8), newer accelerators and steady software gains.

How to read vendor claims from this round

Nvidia and AMD both claimed leadership on 16 September, and both sets of claims can be true at once because they are scoped narrowly. A few rules help:

  1. Check the division. Closed-division results are the apples-to-apples ones. MLCommons itself says the Closed division remains the foundation for comparison.
  2. Check the scenario. A chip that leads in Offline throughput may not lead in Server or Interactive, where latency limits apply.
  3. Compare like system sizes. Per-accelerator figures and whole-system figures tell very different stories; a 72-GPU rack will always beat an 8-GPU server in total tokens.
  4. Look for “selected” in the claim. When a vendor says it led “selected” competitor results, it has chosen the comparison points.
  5. Separate preview from available. Vera Rubin and Rubin results are previews; the MI350P, Arc Pro B70 and Ryzen AI Max+ 395 are listed as available.

The full tables are on the MLCommons Datacenter results page and the equivalent Edge page, and MLCommons also offers an interactive results dashboard.

What happens next

MLCommons has signalled a bigger change coming. David Kanter, Head of MLPerf, said that “MLPerf Endpoints will replace Inference in our family of benchmarks for the datacenter.” More than half of v6.1 submitters already used the new API-based harness that underpins Endpoints, which sends queries to the system under test over industry-standard APIs, much as a real deployment would.

For buyers, the takeaways from v6.1 are straightforward. Nvidia’s next rack-scale platform has its first verified numbers and they are large gains over GB300, although they are previews on two benchmarks. AMD has broadened its coverage and now has a PCIe option in the MI350P that slots into existing servers. Intel is making a case that CPUs already in the data center can handle smaller models, helped by software gains. And the fastest-growing benchmarks are the ones built around mixture-of-experts models, agents and multi-step pipelines, which MLCommons says reflects how AI inference is actually being deployed.

Frequently asked questions

When were the MLPerf Inference v6.1 results released?

MLCommons published the MLPerf Inference v6.1 results on 16 September 2026, with 30 submitting organisations and 120 systems.

How much faster is Nvidia Vera Rubin NVL72 than GB300 NVL72 in MLPerf?

Nvidia says Vera Rubin NVL72 delivered up to 3.7 times the throughput of GB300 NVL72 on Qwen3-VL and up to 2.5 times on DeepSeek-R1. These were preview submissions on two benchmarks.

What is the AMD Instinct MI350P?

It is a dual-slot PCIe 5.0 accelerator card on AMD’s CDNA 4 architecture with 128 compute units, 144GB of HBM3E and 4TB/s of memory bandwidth. MLPerf Inference v6.1 is its first MLPerf round.

What new tests are in MLPerf Inference v6.1?

Two new tests were added: an End-to-End Retrieval-Augmented Generation (RAG) benchmark for the datacenter and an Edge Agentic Inference benchmark that models a multi-turn coding agent on a single device.

What is the difference between the Closed and Open divisions in MLPerf?

The Closed division requires a model mathematically equivalent to the reference, so results are directly comparable. The Open division allows changes to the model or technique.

What is the largest system ever submitted to MLPerf Inference?

Crusoe submitted results using 512 accelerators in v6.1, beating the previous record of 288. AMD says those runs used Instinct MI355X GPUs.

Sources

More on MLPerf →MLPerfNvidia Vera RubinAMD InstinctIntel Arc ProAI acceleratorsBenchmarks
Mara Lindqvist
Written byMara Lindqvist

Mara Lindqvist edits the hardware desk. She covers graphics cards, processors, memory and storage, the foundries and chip designers behind them, and what the numbers on a spec sheet mean for people choosing a PC. Specifications in her stories come from manufacturer spec pages and datasheets.

More from Hardware

See all