brought over details from your upstream/original work for people who find this directly

Correction to the quick-start in this PR: --draft-mtp 3 isn't a llama.cpp option, and llama-mtp-cli isn't a llama.cpp binary. MTP speculative decoding is enabled with --spec-type draft-mtp and --spec-draft-n-max N (matching the "n=3" in the speed note, so --spec-draft-n-max 3). For the vision example, the mainline tool is llama-mtmd-cli. Possible command:
llama-server -m Qwen3.8-27B-GSQstyle-IQ3S-mtp.gguf --mmproj mmproj-Qwen3.8-27B-Q5_K-MIX.gguf --spec-type draft-mtp --spec-draft-n-max 3 -ngl 99 -c 32768 --port 8080

Thanks for the correction β€” all three points check out against upstream (common/arg.cpp registers --spec-draft-n-max / --spec-type; the mtmd tool builds as llama-mtmd-cli). Merged into main: quick-start now uses --spec-type draft-mtp --spec-draft-n-max 3 and llama-mtmd-cli for the vision example.

Cannot merge
This branch has merge conflicts in the following files:
  • README.md

Sign up or log in to comment