← Back to KHAO

Nvidia · Jensen Huang · AI Inference ·

Nvidia’s ultra-low-latency AI inference LPX racks hit full production

2 min read

Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.

◌ Single Source

The Agentic AI Supplement.

Nvidia’s LPX racks for AI inference accelerators have entered full production, the company has confirmed.

Key facts

Summary

Unveiled at GTC back in March, the rack-scale platform came about following Nvidia’s acqui-hire of the eponymous startup. Inside the LPX rack itself are BlueField-4 data processing units, Vera CPU racks, and STX storage servers all tied together with the recently debuted Spectrum-6 Ethernet networking tech. The platform is not a replacement for Nvidia’s flagship NVL72 platform, but rather a complementary add-on for operators wanting to power ultra-low-latency AI inference workloads. During the Hot Chips event in Palo Alto this week, Nvidia cited industry benchmarking results that saw the LPX platform support 3,400 output tokens per second.

Read full article at Datacenter Dynamics →

#Nvidia #Jensen Huang #AI Inference