Skip to content
Daily Edition · AI industry recordEdition of Monday, September 21, 2026
Live desk ●
ResearchNews Report2 min readByAI Tools Daily

NVIDIA Vera Rubin NVL72 Tops Its Debut MLPerf Inference v6.1 Round, Reaching 3.7x GB300 NVL72 Throughput

In its first MLPerf Inference preview submission, published September 16 with v6.1, Vera Rubin NVL72 hits up to 3.7x GB300 NVL72 throughput on Qwen3-VL, while a 288-GPU GB300 run posts 99% scaling efficiency.

ShareXLinkedIn

MLPerf Inference v6.1 results were released on September 16, 2026, and NVIDIA's blog reports that Vera Rubin NVL72 debuted with leading performance in its first MLPerf Inference preview submission: up to 3.7x better throughput than GB300 NVL72 on Qwen3-VL. All figures are MLPerf Inference v6.1 Closed Division results, retrieved from www.mlcommons.org on September 16, 2026.

What the debut submission actually covers

NVIDIA submitted Vera Rubin NVL72 preview results on two of the suite's most demanding benchmarks, with different software stacks:

  • Qwen3-VL: using vLLM with the NVIDIA Dynamo open-source inference framework, up to 3.7x higher throughput than GB300 NVL72 across offline, server and interactive scenarios.
  • DeepSeek-R1: using the NVIDIA TensorRT-LLM library, throughput up to 2.5x higher than GB300 NVL72.
  • The relevant entries are 6.1-0106 and 6.1-0074; partner Nebius also submitted Vera Rubin NVL72 preview results.

Efficiency and scale, in numbers

  • A 288-GPU GB300 NVL72 submission across four racks achieved 99% scaling efficiency in the offline scenario on DeepSeek-R1 (entries 6.1-0073 and 6.1-0074).
  • On the WAN 2.2 text-to-video benchmark, GB300 NVL72 reached 0.65 720p videos per second at 5.7 seconds per video — 9x higher throughput and 7.5x lower latency than a single node.
  • Software gains continued: on Qwen3-VL, GB300 NVL72 improved up to 1.6x in v6.1 over its v6.0 results, from lower KV cache precision, additional kernel fusion, better kernels and disaggregated serving with vLLM and Dynamo.
  • Outside MLPerf, in SemiAnalysis AgentX preview testing, Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72; NVIDIA also points to the upcoming MLPerf Endpoints benchmark as the standardized measure for agentic inference.

Why the gaps appear

NVIDIA attributes the results to full-stack co-design: Vera Rubin's enhanced Tensor Cores and Transformer Engine accelerate both the prefill and decode stages, while NVFP4 precision shrinks the memory footprint of weights, attention and KV cache. Submissions leaned heavily on disaggregated serving plus large-scale expert parallelism for the MoE layers behind DeepSeek-R1 and Qwen3-VL. For interconnect, the NVL72 scale-up domain uses sixth-generation NVLink and NVLink Switch, which NVIDIA says deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet.

How to read these numbers

Three caveats are in NVIDIA's own text: the Vera Rubin figures are preview submissions and described as early results; optimizations run after the v6.1 deadline, shown on GPT-OSS-120B and DLRMv3, are not yet verified by MLCommons; and the 30x agentic figure comes from SemiAnalysis's test, not MLPerf. Beyond the rack-scale platform, NVIDIA also submitted Jetson AGX Thor results using TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark with Qwen3.6-27B, and says 19 partners participated in the round, eight of them on multi-node Blackwell NVL72 systems.

This article aggregates official announcements and public reporting; original sources are linked below.

Source:NVIDIA

AI Tools Daily is a bilingual newsroom covering AI tool launches, product updates and industry trends. Editorial standards · Report a correction

All stories in this section · Research