Bringing Nunchaku 4-bit Diffusion Inference to Diffusers

Imported from official source

Research

AI Classified by Officially

Most of these backends are weight-only. This means that they store the weights in low precision and dequantize them back to high precision at compute time. This reduces memory usage significantly, but it usually does not make inference faster, and can even add a small latency overhead.

SVDQuant, the quantization method behind the popular Nunchaku inference engine, takes a different approach. It runs the main transformer layers with 4-bit weights and activations (W4A4), reducing memory while also speeding up the denoising loop. The details are covered below, but until now, using these checkpoints required a separate inference library.

With current Diffusers, loading a Nunchaku checkpoint is as simple as calling from_pretrained(), with no local CUDA compilation required thanks to the kernels package. In addition, the companion diffuse-compressor toolkit lets you quantize new architectures yourself and publish them as regular Diffusers repositories.

First, install the requirements. You need a recent version of Diffusers and the Hugging Face kernels package:

pip install -U diffusers transformers accelerate kernels bitsandbytes

Then load a pre-quantized pipeline like any other Diffusers model:

import torch
from diffusers import ErnieImagePipeline

pipe = ErnieImagePipeline.from_pretrained(
    "lite-infer/ERNIE-Image-Turbo-nunchaku-lite-nvfp4_r32-bnb4-text-encoder",
    torch_dtype=torch.bfloat16,
).to("cuda")

image = pipe(
    prompt="A cinematic portrait of a red fox in a misty forest at sunrise, "
           "detailed fur, volumetric light",
    height=1024,
    width=1024,
    num_inference_steps=8,
    guidance_scale=1.0,
    generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("output.png")

This is an extract. The publication continues at the source.

Read the original at the source: https://huggingface.co/blog/nunchaku-diffusers

Officially imported this from Hugging Face’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

Provenance

Organization
Hugging Face — imported from official source
Official source
https://huggingface.co/blog/feed.xml RSS
Imported
September 18, 2026 09:42
Versions
1 recorded
Identity
https://huggingface.co/blog/nunchaku-diffusers

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.