Any plans for reviving this model with MTP support?

#14
by phakio - opened

I've been playing with bartowski's quants, and on IQ1_S with MTP full gpu offload, I get nearly 70 t/s TG and 600 t/s PP!

The fun part comes when I tested his IQ4_XS quant on ik_llama.cpp... I get a good 33 t/s TG and 150t/s PP... that's with CPU-MOE offload, leaving lots of room for context on the VRAM. .This, in my opinion, is very usable. I feel as though this model is still relevant today, as artificial analysis places it at 4th place out of all open source models, with reasoning disabled... this shows that it has quite good knowledge density for its size while not relying on reasoning.

I think it's worth revisiting and giving a little life to this guy!

Heya @phakio

Yes, the whole Qwen3.5 series is pretty amazing. With them teasing Qwen3.7 however I'd probably just wait to see if they drop an updated version and put the effort on that.

There was another thread with folks talking about MTP, but are you able to run one of these models using the MTP tensors from another repo on ik_llama.cpp?

I'm hoping they do release another big model with the 3.7 release, but there's a lot of speculation that this model performs too similar to their paid model, and they might just stick to open sourcing the smaller ones from here on out. we'll see soon though!

I saw an interesting comparison chart here https://kaitchup.substack.com/p/lessons-from-gguf-evaluations-ternary , and I'm inclined to heavily agree, that this model is heavily resilient to quantization, as compared to baseline, it only really sees a less than 20% error increase at IQ1... From my testing, even at IQ1 quants the model can do light coding work, accurately reason and solve EVERY medical exam question I come across and feed it, and overall is a really solid model.

I'm not saying it's the best, but I truly do think that with 96GB VRAM, this model, even at IQ1, is a solid choice many may overlook.


Trying to load bartowski's external MTP results in:

llama_model_load: error loading model: check_tensor_dims: tensor 'blk.0.attn_norm.weight' not found
llama_model_load_from_file: failed to load model
phakio changed discussion status to closed

@phakio

I was running the 122B in 96GB VRAM before all the Qwen3.7-27B MTP craze hit the world haha... I agree the qwen models are pretty resiliant to quantization given the recurrant delta net layers (keeping ssm higher BPW doesn't cost much size) is efficient.

here is an interesting one to try out by a new ik_llama.cpp quantizer @cafkafk with MTP (no imatrix, yet): https://huggingface.co/cafkafk/gemma-4-31B-it-assistant-GGUF-noimatrix

cheers!

Sign up or log in to comment