error while handling argument "--spec-type": unknown speculative type: mtp

#7
by Jacksao1970 - opened

./llama.cpp/llama-server --host 127.0.0.1 --port 5678 --flash-attn on --jinja -c 262144 -ngl all --fit on -np 1 --spec-type mtp --spec-draft-n-max 2 -m /mnt/d/llamacpp/Qwen3.6-35B-A3B-UD-Q6_K_XL.gguf
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 16383 MiB):
Device 0: NVIDIA GeForce RTX 3070, compute capability 8.6, VMM: yes, VRAM: 16383 MiB
error while handling argument "--spec-type": unknown speculative type: mtp

usage:
--spec-type none,draft-simple,draft-eagle3,draft-mtp,ngram-simple,ngram-map-k,ngram-map-k4v,ngram-mod,ngram-cache
comma-separated list of types of speculative decoding to use (default:
none)

                                    (env: LLAMA_ARG_SPEC_TYPE)

to show complete usage, run with -h

same issue "unknown speculative type: mtp" with Apple M1 Max :(

now it's draft-mtp

./llama.cpp/llama-server
-hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL
-ngl 99 -c 8192 -fa on -np 1
--spec-type draft-mtp --spec-draft-n-max 2

https://github.com/ggml-org/llama.cpp/compare/440c8e0b0e1e755257775b75beea5e4043ffe5e0..e7b4848151377395b1693d326d1cda3fcd61c2d9

Sign up or log in to comment