noctrex commited on
Commit
de693c8
·
verified ·
1 Parent(s): 1f99273

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +10 -2
README.md CHANGED
@@ -11,8 +11,9 @@ These are **MXFP4** quantizations of the model [Qwen3.6-35B-A3B](https://hugging
11
 
12
  This is the multi-token prediction (MTP) version.
13
 
14
- For the time being, the official mainline **llama.cpp** Does not support MTP yet, so you must compile one for yourself.
15
- This usually will entail pulling the PR 22673, and compiling it. You can find guides online, [like here](https://huggingface.co/havenoammo/Qwen3.6-35B-A3B-MTP-GGUF)
 
16
 
17
  ## Which version should I choose?
18
  All variants use **MXFP4** for the MoE (Mixture of Experts) weights to keep the model efficient. The difference lies in how the remaining tensors are handled:
@@ -27,3 +28,10 @@ All variants use **MXFP4** for the MoE (Mixture of Experts) weights to keep the
27
 
28
  Read the guide from unsloth in order to set up the model's recommended settings for MTP:
29
  [Qwen3.6 - MTP Guide](https://unsloth.ai/docs/models/qwen3.6#mtp-guide)
 
 
 
 
 
 
 
 
11
 
12
  This is the multi-token prediction (MTP) version.
13
 
14
+ ## Quick Start
15
+ 1. Download the latest release of **llama.cpp**.
16
+ 2. Download your preferred model variant from below.
17
 
18
  ## Which version should I choose?
19
  All variants use **MXFP4** for the MoE (Mixture of Experts) weights to keep the model efficient. The difference lies in how the remaining tensors are handled:
 
28
 
29
  Read the guide from unsloth in order to set up the model's recommended settings for MTP:
30
  [Qwen3.6 - MTP Guide](https://unsloth.ai/docs/models/qwen3.6#mtp-guide)
31
+
32
+ On my system it works very well with the commands:
33
+ ```
34
+ --spec-type draft-mtp
35
+ --spec-draft-p-min 0.75
36
+ --spec-draft-n-max 3
37
+ ```