I tried to improve on UD-Q8_K_XL and succeeded

#103
by KissMyShinyArse - opened
Quant Size (GiB) PPL (Q) PPL Ratio ΔPPL Mean KLD RMS Δp (%) Same Top-p (%)
UD-Q8_K_XL-Q8_to_Q8_CR 29.30 6.9565 1.00089 0.0062 0.00037 0.587 99.263
Q8_CR 27.05 6.9585 1.00118 0.0082 0.00043 0.598 99.099
UD-Q8_K_XL 29.30 6.9537 1.00049 0.0034 0.00083 0.823 98.992
Q8_0 27.05 6.9560 1.00082 0.0057 0.00095 0.942 98.742

https://huggingface.co/KissMyShinyArse/Qwen3.8-27B-GGUF

I was wondering why this wasn't being done with text models. INT8 ConvRot is really big for image gen models.

shimmyshimmer changed discussion status to closed

Sign up or log in to comment