Blackwell · Nvidia · AI Inference · Nvidia B200 · Hugging Face
Bringing Nunchaku 4-bit Diffusion Inference to Diffusers
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
Most of these backends are weight-only.
Key facts
- This checkpoint pairs a Nunchaku NVFP4 transformer with a bitsandbytes NF4 text encoder, and generates a 1024x1024 image in about 1.7 seconds on an RTX 5090 with a peak memory usage of about 12 GB
- NVFP4 checkpoints require an NVIDIA Blackwell GPU (RTX 50 series, RTX PRO 6000, B200)
- As shown above, Nunchaku reduces peak VRAM by up to 50% while still improving latency by roughly 30%
- All numbers below were measured on an NVIDIA RTX PRO 6000 (Blackwell) at 1024x1024 using rootonchair/ERNIE-Image-Turbo-nunchaku-lite-int4-bnb4-text-encoder
Summary
SVDQuant, the quantization method behind the popular Nunchaku inference engine, takes a different approach. With current Diffusers, loading a Nunchaku checkpoint is as simple as calling from_pretrained, with no local CUDA compilation required thanks to the kernels package. First, install the requirements. No custom pipeline class or separate inference engine is needed, and there is nothing to compile locally.