Please keep the MTP/nextn tensors — and consider a ~3.25 bpw target

#1
by palmanza - opened

Thanks for this release. We benchmarked IQ3_XXS against Unsloth's UD-IQ3_S (3.44 bpw), which isn't in your comparison table — you only compare against UD-IQ2_S and UD-Q2_K_XL.

Paired runs (same llama.cpp build, same wikitext2 corpus, same seed), IQ3_XXS wins all four cells, every one outside the combined error:

KV / ctx UD-IQ3_S GSQ-RCO IQ3_XXS Δ
q4_0 / 4k 6.4776 ± .042 6.3743 ± .040 −0.103
q4_0 / 16k 6.4652 ± .044 6.1398 ± .039 −0.325
q4_0 / 32k 6.8376 ± .046 6.6517 ± .042 −0.186
f16 / 4k 6.4617 ± .042 6.3541 ± .040 −0.108

Smaller file, fewer bits, better perplexity.

Two requests:

1. Keep the MTP head. IQ3_XXS ships without the multi-token-prediction tensors — we count 848 tensors and not one blk.*.nextn.eh_proj / .enorm / .hnorm / .shared_head_norm, and there's no blk.64. The base checkpoint has them and Unsloth's quants preserve them. In llama.cpp they enable --spec-type draft-mtp: speculative decoding with no external draft model. On a 16 GB card that's worth more than half a bit — it removes a separate 1.1 GB draft file. Passing them through, even at Q8_0, costs about 1 GB. It's also currently undocumented that they're missing.

2. A ~3.25 bpw build. Since RCO takes a total size budget, this should be a one-parameter change. 4 bpw is unusable for 16 GB: the file alone lands at 12.5 GiB, and with a draft model plus a 98k q4_0 KV cache we measure that ~1.2 GB over the card. 3.25 bpw is the largest target that still leaves room for both.

Also useful if cheap: the per-tensor type assignment RCO selected (JSON or table), the imatrix calibration corpus, and any long-context evaluation — your benchmarks are all short-form, and we see perplexity degrade sharply at 32k on every quant we test, which is the regime that matters for agentic coding.

IST Austria Distributed Algorithms and Systems Lab org

Thanks for the detailed benchmarks and suggestions!

I’ve uploaded the imatrix and RCO allocations as well. MTP is on our radar too and we’d like to train a small dedicated MTP layer for our models. We’ll also do 3.25-bit and 3.5-bit models. Thanks for the suggestion!

Thanks for shipping the imatrix and the RCO allocations so quickly — both answered things we couldn't get at from the GGUF alone, and both fed into what follows.

A correction on our side first

We said 848 tensors. The real count is 851, and your allocation file says 851 too — it matches our copy of the GGUF tensor-for-tensor, zero differences in either direction. Our number came from a bad hand-count, apologies. The conclusion is unchanged, and now it comes from your file rather than our reading.

The MTP head is already gone before quantization

This is the one that may save you work. Your imatrix covers blocks 0..63 and contains zero nextn tensors. So the head wasn't dropped by the quantizer — it was already absent when you collected the statistics. That points at the conversion step, not the quantization config.

It matters because you mentioned wanting to train a small dedicated MTP layer. You may not need to. Unsloth's Qwen3.8-27B-UD-IQ3_S.gguf has 866 tensors to your 851, and the difference is exactly the 15 tensors of blk.64:

blk.64.attn_norm.weight            blk.64.attn_q.weight
blk.64.attn_q_norm.weight          blk.64.attn_k.weight
blk.64.attn_k_norm.weight          blk.64.attn_v.weight
blk.64.attn_output.weight          blk.64.post_attention_norm.weight
blk.64.ffn_gate.weight             blk.64.ffn_up.weight
blk.64.ffn_down.weight             blk.64.nextn.eh_proj.weight
blk.64.nextn.enorm.weight          blk.64.nextn.hnorm.weight
blk.64.nextn.shared_head_norm.weight

A full transformer block plus the four nextn-specific tensors, already present in the base checkpoint. If your HF-to-GGUF path preserved it, --spec-type draft-mtp works for the cost of passing tensors through — no training run. Worth checking before spending the compute.

A third column for the table we sent you

We've since run the same paired setup on Unsloth's UD-IQ4_XS (4.24 bpw), same base model, same binary, same corpus, same --seed 1234. Adding it to the numbers from our earlier post:

ctx (KV) UD-IQ3_S (3.44) GSQ-RCO IQ3_XXS (3.00) UD-IQ4_XS (4.24)
4096 (q4_0) 6.4776 ±.042 6.3743 ±.040 6.3882 ±.041
16384 (q4_0) 6.4652 ±.044 6.1398 ±.039 6.0625 ±.039
32768 (q4_0) 6.8376 ±.046 6.6517 ±.042 6.4783 ±.041
4096 (f16) 6.4617 ±.042 6.3541 ±.040 6.3642 ±.041

Against 1.24 bpw more, your IQ3_XXS ties at 4k in both KV precisions (0.2x the combined error, in both directions) and loses only at 16k and 32k (1.4x and 2.9x). So the gap that a full extra bit buys is specifically a long-context gap — everything else it buys is nothing.

The part we didn't expect: your imatrix is 1000 chunks × 4096 tokens, so the calibration never saw past 4k, and the model still holds level with a 4.24 bpw quant there and beats a 3.44 bpw one everywhere. Whatever RCO is doing generalises well past its calibration length.

A guess at the mechanism from your allocation file — hypothesis, not something we measured. The 16 full-attention blocks (3, 7, 11 … 63) average 3.18 bpw, against 2.96 for the 48 DeltaNet blocks and 2.88 for the FFN. The budget went where long-range information lives. Two exceptions stood out: blk.3.attn_k and blk.11.attn_v at IQ1_S. That's where we'd look first if it ever does degrade in long context.

Still interested in any long-context evaluation from your side — all three quants above get worse from 16k to 32k, and that's the regime we actually work in.

3.25 vs 3.5, now with a measurement instead of arithmetic

Good news that both are coming. Our earlier "4 bpw is unusable on 16 GB" was a calculation; we've now measured it on that same 4.24 bpw file, and it's worse than we said. On a 16376 MiB card with a 98k q4_0 KV cache:

window tok/s peak
UD-IQ3_S (3.44) + native MTP 98304 80.0 14404
UD-IQ4_XS (4.24) + native MTP 98304 — OOM
UD-IQ4_XS (4.24), no speculation 98304 38.2 15398

At 4.24 bpw the 98k cache only fits if you give up speculative decoding entirely — a 2.1x clock penalty for 0.8 bpw. Its ceiling with speculation is between 72k and 80k.

Scaling from your measured 14402 MiB peak (IQ3_XXS + 1.1 GB external draft + 98k q4_0 KV), 3.25 bpw lands near 15200 and 3.5 near 16000 of 16376. So 3.25 works comfortably and 3.5 is on the edge with an external draft.

Which loops back to the MTP point: if the head survives conversion, the external draft goes away and 3.5 becomes comfortable. On a 16 GB card that conversion fix is worth more than the extra half bit.

One small thing

The README still doesn't mention that the MTP tensors are absent, and doesn't document the imatrix or the allocation files you just added.

Hi guys this is truely a phenominal model - running on rtx 5080 16GB VRAM and I can do in llama.cpp:
GGML_CUDA_DISABLE_GRAPHS=1 ./llama.cpp/build/bin/llama-server -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ3_XXS --alias local-model --no-mmproj -np 1 -fa on -ctk q4_0 -ctv q4_0 -ngl 66 -fit off -b 2048 -ub 256 -t 14 -tb 28 -rea on --reasoning-budget 4096 --reasoning-budget-message "Reasoning budget reached. Finish your analysis and provide the complete final answer." --no-reasoning-preserve --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --frequency-penalty 0.0 --repeat-penalty 1.0 --cache-ram 8192 --host 0.0.0.0 --port 8080 -lv 4 -c 262144
at full context 262k, q4 kv cache and (prompt prefill) prompt eval is 487.37 tokens per second and (decode) eval time is 62.97 tokens per second

and also :

GGML_CUDA_DISABLE_GRAPHS=1 ./llama.cpp/build/bin/llama-server -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ3_XXS --alias local-model --no-mmproj -np 1 -fa on -ngl 66 -fit off -b 2048 -ub 128 -t 14 -tb 28 -rea on --reasoning-budget 4096 --reasoning-budget-message "Reasoning budget reached. Finish your analysis and provide the complete final answer." --no-reasoning-preserve --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --frequency-penalty 0.0 --repeat-penalty 1.0 --cache-ram 8192 --host 0.0.0.0 --port 8080 -lv 4 -ctk q8_0 -ctv q4_0 -c 196608
at 192k, q8 kv cache and (prompt prefill) prompt eval is 362.41 tokens per second and (decode) eval time is 63.01 tokens per second

I agree, this model deserves more attention... I've tested IQ3_XXS with MTP headers merged from the Unsloth UD_IQ3_XXS and the speed and quality was amazing with just a RTX 5070Ti

Sign up or log in to comment