Why q4 model is so small?

#1
by bb2103 - opened

This Q4 Model: 15G
Other Q4 Models: 18G->21G
Im curious about the difference between models and want to know how to reach this result.
Im thinking about to reimplement this(+ remove mtp, vision and change some layer to mxfp4 to accelerate) to fit 27B to my 4090.
If anyone knows how to do this, please tell me. Thanks

LibertAI org

Why it's 15 GB

It's not the FFN-NVFP4 win that does most of the work β€” NVFP4 is ~4.5 bits/elem packed (4-bit weight + FP8 block scale
per 16 elems), versus Q4_K_M at ~4.84 effective bits/elem. That's only ~0.3 bits/weight cheaper. The bigger savings
come from not boosting embed/lm_head/etc. to higher precision, which most "Q4" GGUFs (unsloth UD, etc.) routinely do.

The actual tensor inventory of Qwen3.6-27B-NVFP4-Q4_K_M.gguf:

  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚ # tensors β”‚ dtype β”‚                                   what                                   β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ 192       β”‚ NVFP4 β”‚ every ffn_gate / ffn_up / ffn_down weight (64 layers Γ— 3)                β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ 33        β”‚ Q6_K  β”‚ output.weight + the biggest projections                                  β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ 273       β”‚ Q4_K  β”‚ token_embd, attn_gate, attn q/k/v on the 16 full-attn layers, ssm_out, … β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ 737       β”‚ F32   β”‚ norms, biases, NVFP4 per-tensor input scales                             β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

MTP and vision are already gone. The base GGUF was built without the MTP head; the vision tower is in the separate
mmproj-Qwen3.6-27B-F16.gguf (889 MB) and you only need it for image input. So for your "text-only on a 4090" goal you
can ignore the mmproj β€” that's already what you want.

How to reproduce / shrink further

  1. Source: mmangkad/Qwen3.6-27B-NVFP4 β€” NVIDIA ModelOpt v0.42 activation-aware NVFP4 calibration, repacked. You cannot
    add NVFP4 to more tensors without re-running ModelOpt β€” NVFP4 needs activation calibration, llama.cpp will not
    synthesize it from BF16. (MXFP4 can be assigned at quantize-time, no calibration needed.)
  2. llama.cpp at master (β‰₯ #22196 for Blackwell NVFP4 MMA, plus #20505/#20506/#22611 for the Qwen3.5/3.6 convert path).
  3. python convert_hf_to_gguf.py β†’ BF16 GGUF with the 192 FFN tensors already typed as NVFP4 (the convert script
    picks up hf_quant_config.json).
  4. llama-quantize --tensor-type '=' input.gguf out.gguf Q4_K_M to retype anything you want. To push lower
    than Q4_K_M:
    - --tensor-type 'attn_q.weight=MXFP4' --tensor-type 'attn_k.weight=MXFP4' --tensor-type 'attn_v.weight=MXFP4'
    --tensor-type 'attn_output.weight=MXFP4'
    - Or flip the SSM in_proj (attn_qkv.weight in this arch) β€” that's the 48Γ— Q6_K block that dominates the non-FFN
    footprint.
  5. For 4090 specifically: NVFP4 and MXFP4 both fall back to the dp4a / MMQ paths on Ada (sm_89) β€” no hardware
    tensor-core speedup there. The size shrinks, but it won't go faster than well-tuned Q4_K_M. Your real constraint on a
    24 GB 4090 won't be weights, it'll be KV cache (64 layers Γ— 4 KV heads Γ— 256 dim is heavy at long context). Use -ctk
    q4_0 -ctv q4_0 and you'll fit way more context than weight quant tweaks will buy you.

Sign up or log in to comment