← Back to KHAO

Nvidia · AMD · Blackwell · Nvidia B200 ·

GLM5.2 on AMD MI355X at 2626 tok/s/node at over 2x lower cost than Blackwell

2 min read

Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.

◌ Single Source

Wafer preview image.

The demand for inference is skyrocketing and outpacing supply.

Key facts

Summary

At around 2.75x cheaper per GPU on average (MI355X vs B300) with comparable hardware specs, the solution to cheap inference is hiding in plain sight, a message they at Wafer have been preaching for months. Conversely, on the MI355X / ROCm stack SOTA performance rarely comes out of the box for these frontier models (sometimes it does!). But as agents improve at kernel and model optimization, this gap is closing in real time. And again, on a 20k in / 1k out, 60% cache hit rate workload, they hit an aggregate throughput of 2626 tok/s/node @ 2.4 rps with a defined knee of ≤5s TTFT, only 80% of the performance measured on a B200, despite being over 2x cheaper.

Read full article at wafer.ai →

#Blackwell #Nvidia #AMD #Nvidia B200