Niwaki

Qwen3.6-19B-A3B-Niwaki-v2-2bit-GGUF

GGUF builds of Qwen3.6-19B-A3B-Niwaki-v2-2bit-mlx — Qwen3.6-35B-A3B pruned to 19B total / ~3.3B active parameters — for llama.cpp and everything built on it. No custom code: unlike the MLX repo, these files run on stock llama.cpp.

Niwaki (庭木) are Japan's garden trees, sculpted by meticulous pruning so that every branch serves the form of the whole. This model applies that spirit to a Mixture-of-Experts. A paper with the full method is coming soon.

Files

file size wt2 ppl (llama.cpp, 512-ctx)
Qwen3.6-19B-A3B-Niwaki-v2-2bit-UD-Q3K.gguf (recommended) 12.2 GB 11.44 ±0.08
Qwen3.6-19B-A3B-Niwaki-v2-2bit-Q4_K_M.gguf 14.8 GB 11.49 ±0.08

First-generation GGUF builds under the identical protocol:

model (UD-Q3K builds) size wt2 ppl
Qwen3.6-27B-A3B-Niwaki-2bit 13.5 GB 10.41 ±0.07
Qwen3.6-19B-A3B-Niwaki-2bit 9.0 GB 13.15 ±0.10
Qwen3.6-11B-A3B-Niwaki-4bit 5.9 GB 17.24 ±0.13

This build beats the same-size first-generation 19B by 13% (11.44 vs 13.15) at 3.2 GB more on disk — the format pad described below.

Reference Qwen3.6-35B-A3B at Q8_0 measures 6.95 under the identical protocol (llama-perplexity, WikiText-2 test, 512-token windows). These llama.cpp numbers are not directly comparable to the MLX repo's 2048-window benchmarks; the relative standings match across both.

Generation battery (measured on the canonical MLX weights; reference scores 0.63 / 0.51 under the identical battery): bigram-diversity avg/min = 0.84 / 0.54 across an 8-prompt code/reasoning/chat/creative battery.

The recommended UD-Q3K build is quantized structure-aware (importance matrices calibrated on a mixed corpus), mirroring the artifact's native allocation: the always-active backbone (attention, shared experts, embeddings) is kept at high precision (Q6_K) while the routed experts ride a compact carrier (q3_k, imatrix-guided). It matches or beats uniform Q4_K_M quality at ~18% fewer bytes on this model family.

Format note: GGUF requires a uniform expert count per model, so the shared-only layers carry zero-valued expert tensors stored at 1.6 bits/weight (3.1 GB of the file). They contribute nothing to outputs; this is why these files are larger than the MLX repo at equal quality.

Model dimensions

total / active parameters ~19B / ~3.3B
layers / routed experts / top-k 40 / 256 / 8 (layers 10–29 are shared-expert-only)
expert intermediate size 512 (unchanged)
context as base model
conversion note speculative-decoding (MTP) draft block not included

Usage

llama-cli -m Qwen3.6-19B-A3B-Niwaki-v2-2bit-UD-Q3K.gguf -p "your prompt" -n 256
# or serve:
llama-server -m Qwen3.6-19B-A3B-Niwaki-v2-2bit-UD-Q3K.gguf

Requires a recent llama.cpp with Qwen3.6 (hybrid linear-attention) support. Canonical benchmarks and the MLX-native artifact: Qwen3.6-19B-A3B-Niwaki-v2-2bit-mlx. Family: v2-4bit · first-generation 27B-2bit · 19B-2bit · 11B-4bit.

Base model by the Qwen team (Apache 2.0); pruning, distillation, and GGUF builds by the Niwaki project, 2026-08.

Downloads last month
130
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for neopolita/Qwen3.6-19B-A3B-Niwaki-v2-2bit-gguf

Collection including neopolita/Qwen3.6-19B-A3B-Niwaki-v2-2bit-gguf