MLCommons Releases MLPerf Training 6.0 Benchmarks
MLPerf Training 6.0 introduces DeepSeek-V3 and GPT-OSS benchmarks, highlighting performance gains for NVIDIA and AMD hardware systems.
MLCommons has launched MLPerf Training 6.0, featuring new benchmarks for the DeepSeek-V3 671B and GPT-OSS 20B Mixture-of-Experts models. DeepSeek-V3 671B utilizes 671 billion total parameters, while GPT-OSS 20B uses 21 billion. NVIDIA systems demonstrated significant scalability, utilizing NVLink to connect 72 GPUs per rack and scale-out networks for up to 8,192 GPUs. A cluster of 8,192 NVIDIA GB300 GPUs finished the DeepSeek-V3 training in 2,021 seconds, while 8,192 GB200 GPUs took 3,340 seconds. In Llama 2 70B tests, eight NVIDIA GB300 GPUs finished in 5,613 seconds, outperforming eight AMD MI355X GPUs at 8,271 seconds. AMD systems, limited to eight GPUs via Infinity Fabric, saw the MI355X complete Llama 3.1 8B training in 91,145 seconds. Results also detailed precision formats including FP4, NVFP4, and MXFP4.