mtpOff faster than mtpOn? - LM Studio 0.4.21 (Build 2) with latest stable runtime packs
#6
by DGanno - opened
Interestingly in my case when I turn off the mtp the model becomes way faster than with mtp.
Has anyone experienced a similar phenomenon?
mtpOn: 50 t/s
mtpOff: 65 t/s
Interesting! I haven't tested that. I'll check it out.
Actually, I've tested Qwen3.6-27B on RTX 5090 in LM Studio. MTP hasn't improved the speed. But when I changed LM Studio to llama.cpp, MTP worked!