Does its job but noticeably less impressive than the full IQ3_S in 0-shot prompt benchmarks

#16
by dandandelion - opened

In fact, it's quite far from the full IQ3_S tho it is still impressive that it is not broken even after half of the experts are pruned.

image
image

For comparison, this is what the IQ3_S does on the same prompts: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/discussions/32
image

Although it's quite a bit faster in strata, around 1.5x to be precise:

image

I think it's logical that it gives a worse pelican result, since the non-coder layers were removed, so there's less data to work with

Sign up or log in to comment