Q5 or Q6 quants ?
Hello!
I'm on Strix Halo and I use Unsloth's Q6_K_XL (+160k context) quant with stock llama. It works and uses all 124Gb of VRAM.
I know Halogen engine is different, but Iβm just curious - is it possible to run Q5 or Q6 on Strix Halo using Halogen engine?
By run I mean run with any reasonable speed π
I asked my "people". Here is what they say:
Not today.
0.7.0 reads unsloth's UD-IQ4_XS and any IQ4_NL/IQ4_XS/IQ3_S/Q4_0 build, and refuses the K-quants by name (Q5_K, Q6_K), which need their own kernels and are next on the list.
Even then Q6 would not fit here. Llama.cpp lets the kernel page everything, while halogen pins the trunk and experts (68 GiB at our 4.6 bits) and pages only the 48 GiB lookup table. A Q6_K_XL pin is about 99 GiB and the machine is 125 GiB, so with scratch and a KV pool it is gone. Q5 is borderline.
Speed if it ran:
prefill unchanged (compute-bound), decode about 20 tok/s serial by the bytes against 25 for IQ4_XS and 35 for our file, more with the drafters. The useful part: on this model the experts barely respond to precision in anything we can measure; what moves quality is the dense trunk at 8 bits, which UD-IQ4_XS already has (its perplexity reads 0.7 to 2.1% better than our own checkpoint). Q5/Q6 experts on top cost bytes for no quality we can see. If you want to try halogen with GGUF, that file is the one.
Quick update. As of 0.11.6 the engine reads the K-quant blocks Q4_K, Q5_K and Q5_1, so the higher-precision path you asked about works now.
unsloth's UD-Q4_K_XL is the one to try. It turned out to be the most accurate GGUF we run, slightly better perplexity than UD-IQ4_XS, at the same prefill and about 3% slower decode. Point HALOGEN_CHECKPOINT at its first shard and it repacks at startup like the IQ4 files do. The README's "Bring your own GGUF" section has the full numbers.
Q6_K across the whole model is still the memory limit from before, so a full Q6 build does not fit on this box. UD-Q4_K_XL is the practical way to get the extra precision you were after.
Whenever I try to run any gguf version (IQ4XS or otherwise) I get the following error. That one block cant be quantized as far as I can tell..
gguf: blk.1.ple_conv1d.weight: F16 is not read by this engine (Q4_1, Q5_0, Q2_K, Q3_K and the IQ2/IQ1/F16 families are not repacked losslessly; no lossy fallback is offered). Use a build whose experts are IQ4_NL, IQ4_XS, IQ3_S, Q4_0, Q4_K, Q5_K or Q5_1 and whose trunk is Q8_0, such as unsloth's UD-IQ4_XS or UD-Q4_K_XL
AI generated, human reviewed.
@davenetdev that line comes from a GGUF made by a quantizer other than unsloth. Those files carry one small tensor, blk.1.ple_conv1d.weight, as F16 where unsloth keeps it at F32, and the engine refused the whole file for it. 0.11.10 reads that tensor exactly, with the same bf16 clean check as before, so those files load now. If you see the line on unsloth's own UD-IQ4_XS or UD-Q4_K_XL, please say which repo the shards came from. Still refused by name: Q4_1, Q5_0, Q2_K, Q3_K and the IQ2/IQ1 families, because reading them needs kernels for their block layouts, not a repack. The README's "Bring your own GGUF" section lists what loads.
Thanks for the reply. Happy to report that after the recent updates I have been able to take the safe-tensors from orcarouter/Qwen3.8-Flash-Next-Uncensored and quantize the resulting bf16 with Unsloth's imatrix. The model seems to be running great at IQ4XS on Halogen. Loaded very quickly and so far giving around 1k+ tps prefill and a steady 35 tps decode. Very nice!
Interesting that only 35913 of 126976 MiB of GTT memory are reporting in use at any given time. Llama.cpp would be using 100+ MiB for the same model.
Interesting that only 35913 of 126976 MiB of GTT memory are reporting in use at any given time. Llama.cpp would be using 100+ MiB for the same model.
Don't believe that number. ~68GB pinned in memory... use htop.