MTP support for faster inference?

#1
by Exo87 - opened

Thanks for the big release, will you support multi token prediction so the inference can be faster without big loss in quality?
Keep improving!!

Older Ornith models supported it, so odds are we'll see mtp or dspark or both. (Either community fine tunes or from actual ornith team)

I am seeing a huge + difference against Qwen3.5-9B along with faster tg and similar pp speed. thanks for the model!

the model identifies as Claude by Anthropic...

Updated the model with MTP head. Thanks for your feedback.

Sign up or log in to comment