Excessive CoT verbosity on simple, well-defined tasks

#5
by desugar - opened

Hi! Thanks for sharing this model.
I’m testing Ornith for agentic coding tasks on my project and wanted to check if the length of its reasoning process is expected or if my setup needs adjusting.

Setup:
Harness: OMP(Oh My Pi)
Engine: llama.cpp (latest CUDA build)
Quantization: Q4_K_M model weights, Q8 KV cache
Context: 131K context window
Parameters: temperature=0.6, top_p=0.95, top_k=20

Issue:
Even for relatively simple bug-fixing tasks where clear instructions and the direct solution are already provided in the prompt, the model spends 200k–300k tokens strictly on pure reasoning. It pauses briefly for file-reading tool calls, doesn't loop, and eventually succeeds, but generating that massive volume of CoT tokens takes a huge amount of time, even running at a relatively fast 40–60 tokens/sec on my machine.

Is this level of verbosity typical for this model, or is there a specific prompt structure or parameter tweak recommended to keep reasoning concise?
TIA 😄

Hi! Thanks for sharing this model.
I’m testing Ornith for agentic coding tasks on my project and wanted to check if the length of its reasoning process is expected or if my setup needs adjusting.

Setup:
Harness: OMP(Oh My Pi)
Engine: llama.cpp (latest CUDA build)
Quantization: Q4_K_M model weights, Q8 KV cache
Context: 131K context window
Parameters: temperature=0.6, top_p=0.95, top_k=20

Issue:
Even for relatively simple bug-fixing tasks where clear instructions and the direct solution are already provided in the prompt, the model spends 200k–300k tokens strictly on pure reasoning. It pauses briefly for file-reading tool calls, doesn't loop, and eventually succeeds, but generating that massive volume of CoT tokens takes a huge amount of time, even running at a relatively fast 40–60 tokens/sec on my machine.

Is this level of verbosity typical for this model, or is there a specific prompt structure or parameter tweak recommended to keep reasoning concise?
TIA 😄

It is native to Qwen and probably would be even more pronounced with these RL trained models. TBH I'd rather have a model be thorough like this than miss a bug and suffer for much later.

Try without OMP, too. The oh-my-something setups tend to produce more thinking tokens. See what the model does without it.

I have tested this model, It thinks a lot, not productive thinking, the ornith-1 was way much better. sometimes it thinks for 4 minutes and the agent stops it.
Is there any fix for this issue?
I'm using opencode

I have tested this model, It thinks a lot, not productive thinking, the ornith-1 was way much better. sometimes it thinks for 4 minutes and the agent stops it.
Is there any fix for this issue?
I'm using opencode

Could you try pi cli?

For any qwen-based model, I think passing a reasoning-budget is a must, otherwise it will think too much. I ran this ornith with --no-reasoning-preserve --reasoning-budget 4096.

I ran this ornith with --no-reasoning-preserve --reasoning-budget 4096.

I think for Qwen models, preserving thinking traces in chat history is critical. Even Ornith’s guide mentions it. Also, setting a hard limit on the reasoning budget may hurt output quality, since the model tends to reason gradually, moving from less important points to more important ones. Cutting it off prematurely could leave you with incomplete or incorrect reasoning traces.

I ran this ornith with --no-reasoning-preserve --reasoning-budget 4096.

I think for Qwen models, preserving thinking traces in chat history is critical. Even Ornith’s guide mentions it. Also, setting a hard limit on the reasoning budget may hurt output quality, since the model tends to reason gradually, moving from less important points to more important ones. Cutting it off prematurely could leave you with incomplete or incorrect reasoning traces.

Yes it will effect quality. This issue is not in base qwen 3.6 35b at all even in ornith-1

I ran this ornith with --no-reasoning-preserve --reasoning-budget 4096.

I think for Qwen models, preserving thinking traces in chat history is critical. Even Ornith’s guide mentions it. Also, setting a hard limit on the reasoning budget may hurt output quality, since the model tends to reason gradually, moving from less important points to more important ones. Cutting it off prematurely could leave you with incomplete or incorrect reasoning traces.

Yes it will effect quality. This issue is not in base qwen 3.6 35b at all even in ornith-1

It is in both. I daily drove both. You shouldn't ever limit nor skip its thinking. The thinking is why these small models outssmart way larger models so there's no point in limiting it unless you want it to act like a real 35B-A3B. Also thinking is super important in MoE models; it's one way to let that 3B active explore that 35b knowledge.

I have tested this model, It thinks a lot, not productive thinking, the ornith-1 was way much better. sometimes it thinks for 4 minutes and the agent stops it.
Is there any fix for this issue?
I'm using opencode

Try it without kv cache

Sign up or log in to comment