Miracle did not happen

#32
by ThaiCat - opened

Tested Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf in llama-server with following arguments --threads 8 -c 131000 -np 1 -fa on --repeat-penalty 1.1 --repeat-last-n 512 --presence-penalty 0.1 --temperature 0.3 --top-p 0.9 --min-p 0.02 -ctk q8_0 -ctv q8_0 -b 2048 -ub 1024 --main-gpu 1

Got thinking loop at ~17k context at first attempt. Stopped and asked to output resulting simple game (not well known) faster. Got non-working game. So far other qwen q4 quants behaved better for me, even MoE. 😞
Never tested in real world use case (150k line project)

How does it behave when using the recommended settings?

We recommend using the following sets of sampling parameters for generation:

    Thinking Mode: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
    Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Please note that the support for sampling parameters varies according to inference frameworks.

I've run that exact quant for millions of tokens in Hermes @148Kcontex - never saw it even hesitate - much less loop. Even the XXS is very stable.
My first suspicion is the temp 0.3 - that is very low for qwen. If you are coding that's much too low on this model and might result in odd behaviour. I think I have seen cases before of low-temp making qwen go a bit lobotomized.
Assuming that you are coding - you should also be very careful with using repeat penalty at all. I don't think that would be a cause of the looping - but coding in general can have lots of repeating text, and sampler params should not try to block this. It's better used in creative writing / conversation ect. DRY (if available) is often a better means to the same end in those cases though.

How does it behave when using the recommended settings?

Frankly I didn't test it with recommended settings. I used the same settings I had best results with MoE (at 0.6 and more MoE variants produces too many errors for me), especially temperature. Perhaps dense model reacts to them differently

I've run that exact quant for millions of tokens in Hermes @148Kcontex - never saw it even hesitate - much less loop. Even the XXS is very stable.

How is the behavior at long contexts? For coding and tool calls?

How does it behave when using the recommended settings?

Frankly I didn't test it with recommended settings. I used the same settings I had best results with MoE (at 0.6 and more MoE variants produces too many errors for me), especially temperature. Perhaps dense model reacts to them differently

You complain about the model without using the right settings?

How does it behave when using the recommended settings?

Frankly I didn't test it with recommended settings. I used the same settings I had best results with MoE (at 0.6 and more MoE variants produces too many errors for me), especially temperature. Perhaps dense model reacts to them differently

I've never had good luck tweaking the parameters myself. Recommended settings always seemed to work the best.

Try it and see if you get the same issues. I tried IQ3_S on dual 7900XTX, but it was slower than non iMatrix quants. This is an AMD problem, according to the internet. But, the IQ3_S quant worked really well for its size, no issues on tool calling/long context work.

I've run that exact quant for millions of tokens in Hermes @148Kcontex - never saw it even hesitate - much less loop. Even the XXS is very stable.

How is the behavior at long contexts? For coding and tool calls?

I've been using it a lot recently in Hermes - for coding jobs - and so far I haven't been able to detect anything unusual. 148K context is the longest I can run, so I can't speak for really extreme context lengths - but if models start to "drift", then it's usually visible already at 16-32K. So far I have seen no evidence of any drift or loss of focus. Agent has never failed to finish a task or forgotten something important.
For tool calls - failed toolcalls are rare. Actually I have only seen it happen in "raw" terminal commands it runs from memory (no template unlike toolcalls) - and it always detects and corrects the problem if there is a command error. It's very impressive. If you give good instructions it will do the job reliably.
I think this has a lot to do with very high presicion being retained at the gated deltanet parts of layers

Both XXS and S variants are very stable. Subjectively it feels like the slightly larger "S" quant is a little more efficient at finding the solution to a problem sooner - but stability-wise they can run for hours and hours on a big task (via delegation / subagents - each get the full 148K). I customized the delegate_task tool to run sequentially since my system doesn't have the spare capacity for parallelism. But the result in total is the same = a big job autonomously run from start to end without issues.

Tested Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf in llama-server with following arguments --threads 8 -c 131000 -np 1 -fa on --repeat-penalty 1.1 --repeat-last-n 512 --presence-penalty 0.1 --temperature 0.3 --top-p 0.9 --min-p 0.02 -ctk q8_0 -ctv q8_0 -b 2048 -ub 1024 --main-gpu 1

Got thinking loop at ~17k context at first attempt. Stopped and asked to output resulting simple game (not well known) faster. Got non-working game. So far other qwen q4 quants behaved better for me, even MoE. 😞
Never tested in real world use case (150k line project)

To explain what's probably going on: Your repeat-last-n is pretty high and your temperature is very low. Counter-intuitively, a high repeat penalty and low temperature can restrict the model's options so much that it does things like repeating the same sentence over and over. If your repeat-last-n is too high it can get into a situation where there are only a small number of viable tokens it can choose from, and by the time it gets to the end of N it has a very high chance of repeating the token the first token in N, creating a loop of whatever length your repeat-last-n is (in this case 512 tokens). It's often caused by the trained thinking interjections like "Wait, ...". If that triggers at the end of your N a loop is almost guaranteed. This is especially likely with a low temperature, because the model isn't allowed to pick less likely tokens that would break it out of the loop.

Lower temperature can help very small models (~4B size) that go off in the weeds at higher temperatures, because they often don't have great options to choose from to begin with. They get best results from just picking the most likely token in most cases. But for larger models it's the opposite - you want to give the model more room to explore other possible tokens, because it typically has a lot of good options to choose from. I wouldn't go below 0.7 for 27B, and only for coding tasks and such at that level. Qwen recommends 1.0 if you have reasoning enabled, and they enable reasoning by default.

Also, you can't directly transfer one model's settings to another architecture and expect success. You can sort of have rules of thumb, but how they were trained will determine how they will perform at different temperature settings. Even then, I think the 35B MOE recommended temperature settings from Qwen are still 0.7 for non-reasoning and 1.0 for reasoning.

Recommended settings made the miracle happen for me. On a mixed setup(rtx3090, 10G 3080, 3x12G 3060) this gave me 33 token/s while utilizing all cards.(with bf16 embeddings)

Sign up or log in to comment