What is the actual difference between these two drafts? (MTP Q4_K_M | 2.79 GB) vs ( MTP Q4_K_M | 1.91 GB)

#73
by sarkaritamminen - opened

Which i should DL, when base model i'm using is UD-Q4_K_XL | 111 GB.

Thank You!

One is shared ( use some layer of the model ) , one is full ( have all is own layer ).

I am not sure but for what i understand :

I you have a setup with one GPU => use shared
If you have multiple GPU => it may be better, if you isolate the MTP draft on one GPU, to use the full.

==> my own understanding :)

The actual difference between these variants (draft models) boils down to whether the MTP (Multi-Token Prediction) model includes hard-copied shared layers from the base model, or assumes that the inference engine (e.g., llama.cpp) will fetch those layers directly from the main model already loaded in memory (shared weights). Your intuitions regarding the single-GPU vs. multi-GPU distinction are partially incorrect, and the optimal choice depends on your software's technical capabilities.

Which version should you download for UD-Q4_K_XL (111 GB)? You should download the full version (2.79 GB) unless you are using the latest software with full support for shared-weights speculative decoding and your hardware allows the entire context to be maintained within a single process. With such a massive base model (111 GB), the full version (2.79 GB) offers significantly greater flexibility and operational stability, and the additional 900 MB of memory usage is negligible given a budget of over 110 GB. Verifying your theory (Single GPU vs. Multiple GPUs): Your assumption based on the number of graphics cards is technically incorrect: Single-GPU configuration: The shared version (1.91 GB) saves VRAM but is not the only valid choice. The full version (2.79 GB) will also work seamlessly on a single GPU; it simply duplicates some weights in memory. Multi-GPU configuration: Isolating the draft process on a separate GPU while using the full version is a bad idea due to PCIe latency. To generate subsequent tokens, the MTP heads must work in close coordination with the base model's final hidden states. If the base model is split across 2 or 3 cards while the draft model is placed on a separate, fourth card, you will incur a massive data transfer cost over the PCIe bus for every single token, completely negating the performance gains from MTP. Deployment recommendation for maximum speed: Both the base model (111 GB) and the draft file should reside in the same memory space—specifically, on the same GPUs where the processing of the base model's layers concludes. If your software (e.g., a specific fork of llama.cpp or vLLM) reports layer incompatibility errors with the 1.91 GB version, do not hesitate to use the 2.79 GB version instead.

Sign up or log in to comment