Nvidia · Jensen Huang · AI Inference · Datacenter Dynamics
Nvidia’s ultra-low-latency AI inference LPX racks hit full production
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
Nvidia’s LPX racks for AI inference accelerators have entered full production, the company has confirmed.
Key facts
- Inside the LPX rack itself are BlueField-4 data processing units, Vera CPU racks, and STX storage servers all tied together with the recently debuted Spectrum-6 Ethernet networking tech
- During the Hot Chips event in Palo Alto this week, Nvidia cited industry benchmarking results that saw the LPX platform support 3,400 output tokens per second
- Among its early adopters is neocloud darling Nebius, which plans to bring the Groq 3 LPX platform into its Token Factory offering
- Nvidia CEO Jensen Huang said the LPX will “transform how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness
Summary
Unveiled at GTC back in March, the rack-scale platform came about following Nvidia’s acqui-hire of the eponymous startup. Inside the LPX rack itself are BlueField-4 data processing units, Vera CPU racks, and STX storage servers all tied together with the recently debuted Spectrum-6 Ethernet networking tech. The platform is not a replacement for Nvidia’s flagship NVL72 platform, but rather a complementary add-on for operators wanting to power ultra-low-latency AI inference workloads. During the Hot Chips event in Palo Alto this week, Nvidia cited industry benchmarking results that saw the LPX platform support 3,400 output tokens per second.