Welp-35B-A3B (GGUF)

A 13.21 GB GGUF of Qwen/Qwen3.6-35B-A3B that runs on a 16 GB GPU with the full 262k context, and scores higher than Unsloth's UD-IQ3_XXS at exactly the same size.

It uses only standard llama.cpp formats, the same ones as Unsloth's UD-IQ3_XXS, so it should run wherever that file runs. Tested with llama.cpp (PrismML build b10754 and a patched build of the same).

No fine-tuning or new training data: the same model, quantized more carefully.

File

Welp-35B-A3B.gguf, 13.21 GB. SHA-256 in SHA256SUMS.

tensors type
routed experts gate/up IQ2_S
routed experts down IQ3_XXS (IQ4_XS in 3 layers)
attention, DeltaNet, shared expert, embeddings, output Q6_K

This is the same per-tensor type mix as Unsloth's UD-IQ3_XXS. Only the expert weights are quantized differently (see below); everything else is byte-identical to Unsloth's file.

Running

Full 262k context on a 16 GB GPU (14.8 GB used, including about 0.8 GB for the desktop):

llama-server -m Welp-35B-A3B.gguf -ngl 99 -fa on -c 262144 -ctk q4_0 -ctv q4_0 -ub 256 --jinja

For a more precise KV cache at half the context: -c 131072 -ctk q8_0 -ctv q8_0.

Results

Models that fit entirely on a 16 GB GPU, measured on one RTX 4080 Super 16 GB with the same harness, single runs.

model size HumanEval long-exact suite context on 16 GB decode tok/s prefill tok/s (4k)
Welp-35B-A3B 13.21 GB 95.7% 23/37 262k 164 5,111
Unsloth UD-IQ3_XXS 13.21 GB 93.9% 21/37 262k 165 5,165
Qwen3.8-27B (Unsloth UD-Q3_K_XL) 13.15 GB 96.3% 26/37 131k 47 2,155
Bonsai 2 27B 5.95 GB 91.5% 15/37 262k 85 2,278
  • HumanEval: 164 problems, thinking off, temperature 0, max 1024 tokens.
  • Long-exact suite: bonsai-ada-surgery suite/ default plan, 37 long tool-using tasks, thinking on. Bonsai 2 scores 15/37 here against 17/37 in that repo's report (RTX 4070).
  • Context: the largest that fits entirely on the card with a q4_0 KV cache.
  • Decode: llama-bench tg128. Prefill: llama-bench pp4096.
  • Perplexity is only comparable between Welp and UD-IQ3_XXS, which share the base model: wikitext-2 5.762 vs 5.873, code 1.880 vs 1.900 (context 2048, 40 chunks).
  • Qwen3.8-27B is the dense model Bonsai 2 is built from. Welp solves one fewer HumanEval problem (157 vs 158) and three fewer long-exact tasks (23 vs 26), at 3.5 times its decode speed and twice its context. The HumanEval and suite differences against UD-IQ3_XXS are small enough to be run-to-run noise on their own; the perplexity gain is consistent.

HumanEval pass@1

HumanEval pass@1: Qwen3.8-27B 96.3%, Welp-35B-A3B 95.7%, UD-IQ3_XXS 93.9%, Bonsai 2 27B 91.5%

Decode speed (tok/s, llama-bench tg128)

Decode speed: UD-IQ3_XXS 165, Welp-35B-A3B 164, Bonsai 2 27B 85, Qwen3.8-27B 47 tok/s

Prefill speed (tok/s, llama-bench pp4096)

Prefill speed: UD-IQ3_XXS 5,165, Welp-35B-A3B 5,111, Bonsai 2 27B 2,278, Qwen3.8-27B 2,155 tok/s

HumanEval against GGUF size (log scale). The dashed line joins the best Qwen3.6-35B-A3B quantization at each size; Welp moves it up at 13.21 GB and ties Q4_K_M (157 of 164) at 62% of its size. The two ternary points are research builds from the Welp log and are not released.

HumanEval pass@1 against GGUF size: Welp 95.7% at 13.21 GB, above UD-IQ3_XXS at the same size and level with Q4_K_M at 21.17 GB

At the full 262k context (q4_0 KV), after a 214k-token prompt: decode 107 tok/s, prefill 2,043 tok/s. A passcode hidden at 10%, 50% and 90% of that prompt was retrieved 3/3.

Speed against context

Decode tok/s (128 tokens) and prefill tok/s (next 2,048 tokens) after the given amount of context, llama-bench -d, q4_0 KV and -ub 256 for every model, RTX 4080 Super 16 GB, mean of 2 runs.

Decode speed against context

Decode speed against context: Welp 161 to 105 tok/s from empty to 252k, Bonsai 2 27B 84 to 45, Qwen3.8-27B 47 to 38 and out of memory beyond 128k

Prefill speed against context

Prefill speed against context: Welp 3,826 to 1,396 tok/s from empty to 252k, Bonsai 2 27B 2,335 to 548, Qwen3.8-27B 2,132 to 851 and out of memory beyond 128k

context Welp-35B-A3B decode prefill Bonsai 2 27B decode prefill Qwen3.8-27B decode prefill
0 161.5 3,826 83.9 2,335 46.6 2,132
4k 161.8 3,667 83.8 2,247 46.4 2,044
16k 155.1 3,301 80.4 1,952 45.3 1,791
32k 151.8 3,063 76.5 1,660 44.1 1,550
64k 141.1 2,637 69.4 1,285 41.6 1,220
128k 126.9 2,029 58.8 882 37.5 851
192k 115.2 1,647 50.9 672 does not fit
252k 105.0 1,396 45.1 548
  • Welp keeps 65% of its decode speed at 252k; at 252k it still decodes faster than either 27B model with an empty cache. Only 10 of its 40 layers use full attention (the rest are Gated DeltaNet linear attention), and 3B of its parameters are active per token.
  • -ub 256 is the batch size Welp needs for 262k on 16 GB, so prefill here is lower than the prefill column in the table above (pp4096, default batch). UD-IQ3_XXS has the same architecture and formats as Welp and runs at the same speed.

How it was made

Each expert matrix is quantized one 256-weight block at a time with llama.cpp's own quantizer, using that expert's activation statistics as the importance matrix. After each block, its rounding error is pushed into the columns not yet quantized (GPTQ, applied block-wise so any llama.cpp format works). Layers are processed in order, each fed the output of the already quantized layers. Calibration: 256 sequences of 2048 tokens, half code, half wikitext-2 train.

Recipe and scripts: github.com/quaedra/welp.

Credits

Base model by the Qwen team. The per-tensor type mix follows Unsloth's UD-IQ3_XXS, whose non-expert tensors this file reuses. Quantization formats from llama.cpp.

License

Apache 2.0, same as the base model.

Downloads last month
65
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for quaedra/Welp-35B-A3B-GGUF

Quantized
(871)
this model

Space using quaedra/Welp-35B-A3B-GGUF 1