1000+ prefill optimizations - ciru v4?

#9
by namore - opened

https://pwilkin.github.io/strix-halo/ just dropped an open source optimizations fork that reaches 1000+ PP with 3.8 flash. Maybe itβ€˜s worth a look for ciru ul4 v4? Maybe Opus/Sol can graft yours and his branch together?

I will stick to your version for now anyways since the quality of your quant is really good and the memory usage rather low with the ssd streaming in plays. Really amazing model.

Working on it, hope to have these changes integrated soon.

its out!

Will this be able to be used with the orca variant as well?

Let me get that out now! I love the orca version.

My halos are currently working on seeing if we can apply this to the new 35b agent models. Imagine if I can get them over 3k pp 🀞🀞

Let me get that out now! I love the orca version.

My halos are currently working on seeing if we can apply this to the new 35b agent models. Imagine if I can get them over 3k pp 🀞🀞

Could you release a ROCMFPX version of this model ? VLLM is good for multi-user but less token rate for single user if we compare to llama.cpp

https://llm.ciru.ai/research/ornith-strix/

My models are already faster than the rocmfpx versions I tested against and higher quality.

Even from vllm. Did a lot of work to get it caught up.

Try apodex version also I think I like it better.

Sign up or log in to comment