Lead dispatchhuggingface ▸ tooling· filed 24 Jul 2026
Nunchaku Lite Finally Lets You Run Big Diffusion Models on Consumer GPUs—Without the Usual Headaches
Nunchaku Lite integrates SVDQuant quantization into Diffusers, enabling W4A4 inference on consumer GPUs with half the VRAM and up to 1.8x speedup. It supports pre-quantized checkpoints via standard `from_pretrained()` calls, with NVFP4 for Blackwell and INT4 for older GPUs. While slightly slower than the native engine, it offers model-agnostic support and a straightforward quantization toolkit.