Will a similar IQ3_S version be produced later?

#1
by DataSoul - opened

After downloading and testing, I found that the current IQ3_XXS version enters an infinite loop when performing certain tasks, whereas the original IQ3_S version does not have this issue.

I would suspect this is not as much because it is not IQ3_S - but rather that this is not a "pure" GSQ-RCO optimization. It is a shortcut. Therefore, simply adding more bits very likely wouldn't solve the problem.

From model card info:
"The important honesty clause: I reproduced ISTA-DASLab's published per-tensor RCO allocation verbatim from their stock-base artifacts. I did not independently re-run the multi-GPU budget search โ€” same map, applied to the uncensored base."

So this means he didn't re-optimize the orcarouter (decensored) model, he assumed the uncensored model was MOSTLY similar in distribution to the original. And it probably is... MOSTLY. Trying to do the same on a deep finetune liek a DavidAU model would produce pure garbage as it would be large mismatch. In this case it's pretty similar, but not perfectly the same, so there is some "drift" at work here. The optimization only fits this model 99% (made up number) - and that matters when the optimiations are pushed so hard as this.

My own experience is: It's semi-stable, but I would not use it for agentic long context work. For creative writing in Sillytavern it's ok for now. But I have seen that eventually breaks down too when context grows long (unlike the original GSQ-RCO). It starts repeating itself and going in circles increasingly. Treat this as a stopgap model until a "real" GSQ-RCO version decensored version becomes available.

This is not a knock on the author. He disclosed this fair and honest. It takes significant compute to re-run this method - and not everyone can do that easily. But as a user this is something you need to look for very carefully. Not all model cards outright state this as clearly as RentedNoodle does (and that can leave you very disappointed in the result).
It is possible to use a simple tensor-analysis script to check if the distribution is identical to the original GSQ-RCO version. if it is - it is a "lazy" version. Useful to test when model card does not specify. I'd be happy to share a script if you wish - but im not allowed to attach .py files here. Or you can just ask your favorite SOTA LLM to make you one (it's not hard).

EDIT: actually - looks like they already have made proper versions. Here you go:
https://huggingface.co/huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF/tree/main

Speed Compare
RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored-v1.1.gguf
OVERALL: mean 34.0 median 31.6
RentedNoodle/Qwen3.8-27B-OrcaRouter-GSQ-RCO-IQ3_XXS-v2.0.gguf
OVERALL: mean 35.4 median 35.2
huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_XXS-mtp.gguf
OVERALL: mean 21.5 median 21.0
huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp.gguf
OVERALL: mean 34.7 median 32.0

Sign up or log in to comment