Agentic AI · Nvidia · NVIDIA Blog
AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency at the AI Infra Summit, the Santa Clara Convention Center event that has morphed into a Coachella of infrastructure tech.
Key facts
- On a 100K-context Qwen 3.8 27B workload, Groq 3 LPX hit 2,529 output tokens per second per user
- That’s where NVIDIA Groq 3 LPX comes in, adding deterministic ultralow-latency inference to Vera Rubin and complements DSX MaxLPS
- The results were significant: Lambda ran 19 nodes within the same power budget typically allocated to 16 full-power nodes, increasing cluster-wide token throughput by 24%, from about 4 million to 5
- Up to 30x higher throughput per megawatt means up to 30x more agentic work from the same energy footprint, and the AgentX results show up to 45x lower cost per million tokens
Summary
Before a packed audience, with more than 8,000 attendees this year, up from 3,500 last year, Buck discussed new collaborations across NVIDIA platforms and more. Amazon’s Annapurna Labs is working with NVIDIA on the NVHBM custom high-bandwidth memory technology. D-Matrix is integrating with NVLink Fusion to combine NVIDIA Vera CPUs with d-Matrix Raptor XPUs to deliver ultra low-latency inference at scale. Emerald AI and NVIDIA demonstrated a commercial AI factory flexible-load program, working with Silicon Valley Power.