Nvidia · San Francisco · NVIDIA Blog
From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Power sent a signal to an AI factory to adjust its power consumption.
Key facts
- Lambda, a GPU cloud provider serving more than 10,000 customers from AI-native startups to hyperscalers, ran the software on a five-rack, 19-node cluster
- Lambda’s results, released at the AI Infra Summit, are the first validation of DSX MaxLPS on NVIDIA HGX B20 0 GPU Servers
- What they found: by running 19 nodes within the same power budget as 16 nodes at full power, Lambda achieved 24% more cluster-wide token throughput, from roughly 4 million tokens per second to 5
- Based on NVIDIA’s projections, DSX MaxLPS can enable up to 40% more GPU capacity for next-generation Vera Rubin NVL72 AI factories within the same megawatt power budget in suitable deployment
Summary
Varun Sivaram was watching on Zoom with about forty others, his team at Emerald AI in their San Francisco conference room, engineers at the data center and people from the utility itself. Emerald AI’ s Conductor platform, a grid-orchestration platform from NVIDIA partner Emerald AI, and an early example of the kind of flexibility NVIDIA DSX Flex is built to deliver, receives signals about grid conditions and adjusts the data center’s flexible computing workloads. The goal is to reduce electricity demand when the grid is constrained without interrupting critical AI workloads— exactly what Silicon Valley Powe r needed,. When the reduction showed on screen, everyone cheered.