Transformers now runs llama.cpp quants

Imported from official source

Research

AI Classified by Officially

Running AI models on your laptop has become much easier, and llama.cpp has been a big part of that. Its inference engine powers local AI tools such as Ollama, LM Studio, and Jan. Alongside projects like MLX, it has helped make local inference a practical option for everyday use.

A recent example of what local AI can feel like:

This is where we are right now. And i’m not gonna lie it feels pretty magical 🧙‍♀️

Qwen3.6 27B running inside of Pi coding agent via Llama.cpp on the MacBook Pro

For non-trivial tasks on the @huggingface codebases, this feels very, very close to hitting the latest Opus in Claude… pic.twitter.com/lsIxLoUneU

GGUF, developed by the llama.cpp team, is a widely used format for local inference. The team also shares quantized checkpoints under ggml-org on the Hub. Publishers such as Unsloth, LM Studio Community, and bartowski also provide ready-to-use GGUF checkpoints in a range of quantizations, so users can pick the version that fits their machine. GGUF models have been downloaded millions of times.

We want to make it easier to run these models locally with transformers, too. Compatibility is only useful if the model is pleasant to run. To bring performance close to llama.cpp, we're reusing its underlying ggml kernels through the kernels library, and reducing overhead in generate. Our initial focus is local inference on Apple Silicon, starting with the Qwen3.5 architecture.

GGUF packages model weights and metadata, including tokenizer information and an optional chat template, in one file. It supports different quantization levels, letting you trade some precision for a smaller memory footprint. Variants such as Q4_K_M mix tensor precisions, using mostly 4-bit weights while keeping sensitive tensors at higher precision.

Here's how quantization changes the file size of Unsloth's Qwen3.5-4B:

This is an extract. The publication continues at the source.

Read the original at the source: https://huggingface.co/blog/transformers-llama-cpp-quants

Officially imported this from Hugging Face’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.

Provenance

Organization
Hugging Face — imported from official source
Official source
https://huggingface.co/blog/feed.xml RSS
Imported
September 22, 2026 11:30
Versions
1 recorded
Identity
https://huggingface.co/blog/transformers-llama-cpp-quants

Officially records where a publication came from, not whether it is true. Imported records are reproduced from an organization's own official source.