Praise to HarmonicMinus

#1
by Koshkasa - opened

In testing, HarmonicMinus appears to be the sweet(est) spot for a 16-24GB GPU, although it takes a fully vacant 16GB GPU, small context, or aggressive KV cache quantization. That said, with KV cache in RAM, it correctly utilizes tool calls and stays on track way past 60k - I suppose I don't have a sufficiently large NIHS set to test it against.

To claw back some VRAM space on sm_120 capable GPUs specifically, what I tried is an edit of HarmonicMinus. It involves capping ffn_down and ffn_up at Q6_K instead of Q8_0, and every sparse expert tensor that was either at Q4_K, or ffn_down_exps Q3_K falling back to Q4_0 during quantization, now switched to MXFP4. The result is a 4.48bpw quant, 13.1GB in size. I am currently testing it against HarmonicMinus, but so far, the results seem promising.

If file size isn't a major concern, the regained extra 0.3 bpw budget could be spent on quantizing global attn to q8_0 or slightly increasing quality layer allocation.

My main conclusion from trying to utilize MXFP4 in APEX is never to quant attention at MXFP4, under any circumstance - the coarse attn outlier precision hits Gemma 4 particularly hard, breaks tool calls, punctuation, and on rare occasions, grammar too. The lower precision ffn_down_exps that were quantized to q4_0 seem not to suffer hard from lower outlier precision in mxfp4

UPD: recipe and test quants for comparison uploaded here.

Cheers, great idea and I'd love to test your MXFP4 variant. Swamped with life at the moment, but not abandoning the project, just forced to take a little break.

Sign up or log in to comment