This rolls out on AWS’s previously released support for NVLink Fusion
Compiled by KHAO Editorial — aggregated from 2 sources. See llms.txt for citation guidance.
✓ KHAO Verified
“NVHBM represents a new architectural approach to advancing high-bandwidth memory performance and efficiency,” said Nafea Bshara, vice president of Annapurna Labs at Amazon.
Key facts
- By integrating the memory controller into the 3D HBM stack instead of the XPU, NVHBM delivers up to 30% greater memory bandwidth and 15% lower HBM power consumption, and frees up to 25% more area
- Amazon’s Annapurna Labs will work with NVIDIA on NVHBM technology and the NVLink scale-up architecture to enhance performance and efficiency for AI workloads
- NVHBM represents a new architectural approach to advancing high-bandwidth memory performance and efficiency,” said Nafea Bshara, vice president of Annapurna Labs at Amazon
- Amazon’s Annapurna Labs will be the first to work on NVHBM as part of its broader collaboration with NVIDIA around NVLink Fusion
Summary
As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not only on compute, but on how compute, memory, storage, networking and software are designed together as a unified system. To help hyperscalers and AI innovators build the next generation of semi-custom AI infrastructure, NVIDIA today expanded NVIDIA NVLink Fusion with NVIDIA NVHBM, a next-generation high-bandwidth memory technology that brings higher memory performance and efficiency to XPUs. Traditional HBM architectures place the memory controller on the XPU die, consuming valuable silicon area that could otherwise be dedicated to compute. By integrating the memory controller into the 3D HBM stack instead of the XPU, NVHBM delivers up to 30% greater memory bandwidth and 15% lower HBM power consumption, and frees up to 25% more area on XPU compute die compared with standard HBM4E. NVIDIA is establishing a standard NVHBM implementation, available from multiple memory providers.