mtpOff faster than mtpOn? - LM Studio 0.4.21 (Build 2) with latest stable runtime packs

#6
by DGanno - opened

Interestingly in my case when I turn off the mtp the model becomes way faster than with mtp.
Has anyone experienced a similar phenomenon?
mtpOn: 50 t/s
mtpOff: 65 t/s

Interesting! I haven't tested that. I'll check it out.

Actually, I've tested Qwen3.6-27B on RTX 5090 in LM Studio. MTP hasn't improved the speed. But when I changed LM Studio to llama.cpp, MTP worked!

Sign up or log in to comment