Amazing Quant

#1
by Neiko2002 - opened

I was testing different Laguna quants on benchlocal.com benchmark and my 4x 3090 setup. Your quant scored the highest and had some amazing decode speed. It seems a little more verbose and produced more tokens.
image

Thank you for the feedback !! Hope it will be useful to you. The models are from my personal collections and are released as I do not have sufficient disk space to keep them. As always, if it doesn't work, tell us. If it works well, tell others. 🙃

If you are into Qwen 3.6 , this is another of my favorite -

https://huggingface.co/JasonW2025/Qwen-3.6-35B-HybridQuant-4-1

Thank you I love FP8 KV cache adjusted quants. Its a much more realistic production scenario. Since I'm using 4x 3090 I should have also looked into your JasonW2025/Laguna-S-2.1-ModelOpt-NVFP4-W4A16-vllm quant, to see if there is a visible difference between the two activation types. Or if W4A4 naturally falls back to W4A16 on GPUs which do not support NVFP4.

All my optimizations are based on Spark. I usually prefer W4A16 instead of W4A4 but never did a full bench . If you benched the W4A16, I would love to see the numbers.

Could you give me access to the W4A16 model?

The KV calibrated version was 15% more token efficient than the normale W4W4. The W4A16 Version was even more token efficient. Unfortunately the old W4A4 was benched with a 230w power cap (the other two on 250w), which lead to the slight different in decoding speed. Overall accurary of all three models is the same.
image

Sign up or log in to comment