Niwaki

Qwen3.6-19B-A3B-Niwaki-2bit-GGUF

GGUF builds of Qwen3.6-19B-A3B-Niwaki-2bit-mlx — Qwen3.6-35B-A3B pruned to 19B total / 3.3B active parameters — for llama.cpp and everything built on it.

Niwaki (庭木): every routed expert individually width-pruned to the neurons its own routed tokens actually use, reconstructed to compensate, then briefly distilled from the full model, and stored at low precision. A paper with the full method is coming soon.

Files

file size wt2 ppl (llama.cpp, 512-ctx)
Qwen3.6-19B-A3B-Niwaki-2bit-UD-Q3K.gguf (recommended) 9.0 GB 13.15 ±0.10
Qwen3.6-19B-A3B-Niwaki-2bit-Q4_K_M.gguf 11.4 GB 13.26 ±0.10

Reference Qwen3.6-35B-A3B at Q8_0 measures 6.95 under the identical protocol (llama-perplexity, WikiText-2 test, 512-token windows). These llama.cpp numbers are not directly comparable to the MLX repo's 2048-window benchmarks; the relative standings match across both.

Generation battery (measured on the canonical MLX weights; reference scores 0.89 / 0.78): bigram-diversity avg/min = 0.89 / 0.75 across an 8-prompt code/reasoning/chat/creative battery.

The recommended UD-Q3K build is quantized structure-aware (importance matrices calibrated on the same mixed web/code/chat/reasoning corpus as the model itself), mirroring the artifact's native allocation: the always-active backbone (attention, shared experts, embeddings) is kept at high precision (Q6_K) while the pruned routed experts ride a compact carrier (q3_k, imatrix-guided). It matches or beats uniform Q4_K_M quality at ~20% fewer bytes on this model family.

Model dimensions

total / active parameters 19B / ~3.3B
layers / routed experts / top-k 40 / 256 / 8
expert intermediate size 256 (from 512)
context as base model
conversion note speculative-decoding (MTP) draft block not included

Usage

llama-cli -m Qwen3.6-19B-A3B-Niwaki-2bit-UD-Q3K.gguf -p "your prompt" -n 256
# or serve:
llama-server -m Qwen3.6-19B-A3B-Niwaki-2bit-UD-Q3K.gguf

Requires a recent llama.cpp with Qwen3.6 (hybrid linear-attention) support. Canonical benchmarks, method outline, and the MLX-native artifact: Qwen3.6-19B-A3B-Niwaki-2bit-mlx. Family: 27B-2bit · 11B-4bit.

Base model by the Qwen team (Apache 2.0); pruning, distillation, and GGUF builds by the Niwaki project, 2026-08.

Downloads last month
347
GGUF
Model size
19B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for neopolita/Qwen3.6-19B-A3B-Niwaki-2bit-GGUF

Collection including neopolita/Qwen3.6-19B-A3B-Niwaki-2bit-GGUF