How to use from
Lemonade
Pull the model
# Download Lemonade from https://lemonade-server.ai/
lemonade pull AesSedai/Qwen3.8-Flash-Next-GGUF:
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-GGUF-
List all available models
lemonade list
Quick Links

Updates / Notes

  • 09/02/2027: Added the PLEQ4_0 for IQ2_S and IQ3_S
  • 09/01/2027: The original quants all have Q8_0 for the engram embedding tensors, and I've uploaded three variants with Q4_0 for the embedded tensors. The only difference is the Q8_0 vs Q4_0 for those tensors. I'm leaving the original quants up for those who want to use the Q8_0 PLE's.

This repo contains specialized MoE-quants for Qwen/Qwen3.8-Flash-Next. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.

Quant Size Mixture PPL 1-(Mean PPL(Q)/PPL(base)) KLD
Q5_K_M 147.17 GiB (7.14 BPW) Q8_0 / Q5_K / Q5_K / Q6_K 4.600285 ± 0.027992 +0.2179% 0.030432 ± 0.000210
Q5_K_M (Q4_0 PLE) 126.64 GiB (6.15 BPW) Q8_0 / Q5_K / Q5_K / Q6_K 4.616165 ± 0.028059 +0.5638% 0.030814 ± 0.000217
Q4_K_M 126.08 GiB (6.12 BPW) Q8_0 / Q4_K / Q4_K / Q5_K 4.618153 ± 0.028121 +0.6071% 0.041479 ± 0.000274
IQ4_XS 109.09 GiB (5.30 BPW) Q8_0 / IQ3_S / IQ3_S / IQ4_XS 4.733298 ± 0.029068 +3.1156% 0.079730 ± 0.000510
Q4_K_M (Q4_0 PLE) 105.55 GiB (5.12 BPW) Q8_0 / Q4_K / Q4_K / Q5_K 4.633212 ± 0.028203 +0.9352% 0.043324 ± 0.000287
IQ3_S 100.01 GiB (4.85 BPW) Q6_K / IQ2_S / IQ2_S / IQ3_S 4.870320 ± 0.030101 +6.1006% 0.162764 ± 0.000918
IQ2_S 97.66 GiB (4.74 BPW) Q6_K / IQ2_XS / IQ2_XS / IQ3_XXS 5.037301 ± 0.031450 +9.7383% 0.205950 ± 0.001125
IQ4_XS (Q4_0 PLE) 88.56 GiB (4.30 BPW) Q8_0 / IQ3_S / IQ3_S / IQ4_XS 4.750635 ± 0.029142 +3.4933% 0.084118 ± 0.000515
IQ3_S (Q4_0 PLE) 79.55 GiB (3.86 BPW) Q6_K / IQ2_S / IQ2_S / IQ3_S 4.887335 ± 0.030175 +6.4713% 0.167344 ± 0.000940
IQ2_S (Q4_0 PLE) 77.21 GiB (3.75 BPW) Q6_K / IQ2_XS / IQ2_XS / IQ3_XXS 5.076957 ± 0.031749 +10.6023% 0.213340 ± 0.001164

kld_graph ppl_graph

Downloads last month
6,783
GGUF
Model size
177B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AesSedai/Qwen3.8-Flash-Next-GGUF

Quantized
(179)
this model