Great stuff but...
A great setup - fast prefill and decoding when three/four sessions work concurrently.
However, based on my 2-day user experience:
Higher possibility of looping than the base Ornith1.5 35B during reasoning. Sometimes failed in call tools.
Look forward to further improvements for agentic usage!
Thanks, I'm wondering if it's something about the xml parser causing problems. I noticed some weird behavior in pi harness and I don't think it's the model it looks like chat template or something else messing up the openai compatibility and then the broken outputs creating quality loss
I'll check soon, right now I am adapting this for apodex which I just found out is made exactly for this. It'll be maybe even faster but intended for multi agent use specifically.
Actually I guess that's also linked to Ornith 1.5 35B model also...
Even I when try the Q6 gguf of Ornith 1.5 35B, looping is a quite common issue.
However, for fine-tuned model such as Cyber-Tiel-Coder-35B-A3B-GGUF-MTP - I barely saw looping.
I fixed the quality, but it's still just a 35b.
I'll try tiel.