So what's the difference between the IQ4_XS and IQ4_NL quants?

#34
by xwildfyre - opened

Apologies if this is a dumb question, but I don't understand why the IQ4_XS and the IQ4_NL quants weigh the exact same. I was looking at the llama.cpp output, and the bits per weight are also the same (4.25 bpw for both files.) I would expect the bpw on the _NL to be higher, as well as the file size. They do have different checksums, so they seem to be different files, at least in theory, but the tensors are the same on both files. The IQ4_XS doesn't actually use _XS tensors.

IQ4_XS:

Screenshot 2026-04-18 at 10.09.40 PM

IQ4_NL:

Screenshot 2026-04-18 at 10.10.12 PM

I also downloaded your Qwen3.6-35B-A3B quant and that one actually uses _XS in the IQ4_XS quant.

This is not a knock on you guys by any means, I greatly appreciate your work and thank you for all your contributions to the community! Just genuinely curious about this one weird quirk.

xwildfyre changed discussion status to closed
xwildfyre changed discussion status to open

Sign up or log in to comment