Text Generation
GGUF
nvidia
nemotron-3.5
imatrix
conversational

Possible to strip MTP layer?

#5
by jcmyang750 - opened

Hi,
Is it possible to strip the MTP layer to make the file size smaller? I found that using MTP actually slows the decode by about 10% on Apple M1 Max using llama cpp. Tried both Q6_K and Q8_0. Thanks.

I completely agree with you, the widespread use of MTP is a bad practice because the sending speed decreases from 1600 t/s without using MTP, with MTP the speed drops to 800 t/s, winning maybe 10 t/s (not always) on receiving and losing colossal speed on sending is just disgusting, besides, the MTP layer takes up space, for me these methods are quite enough and work well --spec-type ngram-map-k4v,ngram-mod --spec-ngram-map-k4v-size-n 4 --spec-ngram-map-k4v-size-m 3 --spec-ngram-map-k4v-min-hits 1 --spec-ngram-mod-n-match 16 --spec-ngram-mod-n-min 32 --spec-ngram-mod-n-max 128 --repeat-penalty 1.0

and draft-mtp in some scenarios it even slows down the

I completely agree with you, the widespread use of MTP is a bad practice because the sending speed decreases from 1600 t/s without using MTP, with MTP the speed drops to 800 t/s, winning maybe 10 t/s (not always) on receiving and losing colossal speed on sending is just disgusting, besides, the MTP layer takes up space,

If you load an MTP model specifically not using spec drafting or using a different spec draft, the MTP layers are not loaded. You are not incurring any VRAM waste.

Sign up or log in to comment