Today's AI/ML headlines are brought to you by ThreatPerspective

Digital Event Horizon

NVIDIA Vera Rubin NVL72 Pioneers AI Inference Performance with Groundbreaking Results in MLPerf Inference v6.1 Debut


NVIDIA has made a significant breakthrough in AI inference performance with the introduction of its Vera Rubin NVL72 platform, delivering leading performance in the MLPerf Inference v6.1 benchmark. This achievement has set a new standard for AI inference and highlights the company's commitment to advancing the state-of-the-art in AI infrastructure.

  • NVIDIA's Vera Rubin NVL72 platform delivers leading performance in the MLPerf Inference v6.1 benchmark.
  • The platform optimizes performance, scaling efficiency, and software velocity for AI infrastructure decisions.
  • Advanced architecture and innovations, such as Tensor Cores and NVFP4 precision, result in significant performance gains.
  • The platform demonstrates exceptional scaling efficiency, with 99% scaling efficiency across four GPU racks.
  • Continuous software development and optimization ensure the platform adapts and improves over time.
  • 19 partners demonstrate excellent performance on the platform, setting a new standard for AI inference.



  • NVIDIA has made a significant breakthrough in the field of AI inference performance with the introduction of its Vera Rubin NVL72 platform, which has delivered leading performance in the MLPerf Inference v6.1 benchmark. This achievement marks a major milestone in the company's commitment to advancing the state-of-the-art in AI infrastructure, and its impact will be felt across various industries and applications.

    The Vera Rubin NVL72 platform has been designed to optimize performance, scaling efficiency, and software velocity, which are essential considerations for organizations making AI infrastructure decisions. The platform's advanced architecture, including enhanced Tensor Cores and Transformer Engine, accelerates both the prefill and decode stages of inference, while NVFP4 precision reduces memory footprint across model weights, attention, and KV cache. These innovations have resulted in significant performance gains, with the platform delivering up to 3.7x better throughput than the NVIDIA GB300 NVL72 on the Qwen3-VL benchmark.

    Furthermore, the Vera Rubin NVL72 platform has demonstrated exceptional scaling efficiency, with a 288-GPU submission across four GB300 NVL72 racks achieving 99% scaling efficiency. This means that the platform can handle increased GPU counts while maintaining a proportional growth in throughput, reducing the infrastructure cost and improving the overall economics of AI inference.

    In addition to its technical prowess, the Vera Rubin NVL72 platform has also showcased its commitment to continuous software development and optimization. The platform's ongoing software development has delivered performance and feature improvements, with the company continuing to push the boundaries of what is possible in AI inference. The recent post-submission results, not yet verified by MLCommons, on GPT-OSS-120B and DLRMv3, demonstrate further performance gains, highlighting the platform's ability to adapt and improve over time.

    The Vera Rubin NVL72 platform's success has been mirrored by its partner ecosystem, with 19 partners demonstrating excellent performance on the platform. This includes major players such as ASUS, Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, HPE, Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro, and Wiwynn.

    The NVIDIA Vera Rubin NVL72 platform is a significant milestone in the company's efforts to advance the state-of-the-art in AI infrastructure. Its groundbreaking performance in the MLPerf Inference v6.1 benchmark has set a new standard for AI inference, and its commitment to continuous software development and optimization has ensured that the platform will continue to evolve and improve over time. As the AI landscape continues to shift and evolve, the Vera Rubin NVL72 platform will play a critical role in shaping the future of AI inference.



    Related Information:
  • https://www.digitaleventhorizon.com/articles/NVIDIA-Vera-Rubin-NVL72-Pioneers-AI-Inference-Performance-with-Groundbreaking-Results-in-MLPerf-Inference-v61-Debut-deh.shtml

  • https://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/


  • Published: Wed Sep 16 14:14:35 2026 by llama3.2 3B Q4_K_M











    © Digital Event Horizon . All rights reserved.

    Privacy | Terms of Use | Contact Us