AnonimousA commited on
Commit
4bc31ec
·
verified ·
1 Parent(s): 5fb294f

README: which llama.cpp build the MTP head needs (Unsloth tag, source build for ROCm, token_embd symptom, shared-Q4_K_M)

Browse files
Files changed (1) hide show
  1. README.md +18 -2
README.md CHANGED
@@ -101,7 +101,7 @@ HumanEval-164 94.5% with MTP vs 95.1% without — Unsloth's verification is exac
101
  one-problem gap is sampling noise.
102
 
103
  ```bash
104
- # Needs Unsloth's build, NOT mainline (see trap 4 below)
105
  llama-server -m Q2/Qwen3.8-Flash-Next-UD-Q2_K_XL-reap320-00001-of-00002.gguf \
106
  -md mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf \
107
  --spec-type draft-mtp --spec-draft-n-max 2 \
@@ -110,6 +110,22 @@ llama-server -m Q2/Qwen3.8-Flash-Next-UD-Q2_K_XL-reap320-00001-of-00002.gguf \
110
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
111
  ```
112
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
113
  The head costs **~3.2 GiB of VRAM, flat** — the same at every offload level we tried
114
  (`--n-cpu-moe` 14 / 12 / 10 / 8), so budget for it once. On a 32 GB card at 96k context,
115
  `--n-cpu-moe 10` is the balanced point (30578 MiB, 93.8%, 92.5 tok/s); `8` is the fast one
@@ -163,7 +179,7 @@ the Windows port, which keeps the MTP head working. Full numbers, including the
163
  1. `--parallel > 1` requires `--kv-unified`.
164
  2. Empty `content` with tight `max_tokens`: reasoning burns the budget. Detect and retry with a higher cap.
165
  3. First requests after a cold load pay disk; throughput climbs over ~3-4 requests to its warm ceiling. The disk cost is the PLE table (see the prefill section above); `--lazy-mode on-direct` removes it.
166
- 4. **`--spec-type draft-mtp` exists in llama.cpp mainline but does not work here.** Mainline has the flag and not the MTP graph for `qwen4exp`, nor the tensor borrowing the *shared* heads need; it fails with `model doesn't contain MTP layers`. Use Unsloth's build.
167
  5. `"reasoning_effort": "none"` returns **HTTP 500** — the chat template accepts only `xhigh` / `medium` / `low`.
168
 
169
  ## Credits
 
101
  one-problem gap is sampling noise.
102
 
103
  ```bash
104
+ # Needs a build with the Flash-Next MTP graph, NOT stock mainline (see "Which build" below and trap 4)
105
  llama-server -m Q2/Qwen3.8-Flash-Next-UD-Q2_K_XL-reap320-00001-of-00002.gguf \
106
  -md mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf \
107
  --spec-type draft-mtp --spec-draft-n-max 2 \
 
110
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
111
  ```
112
 
113
+ **Which build.** Stock llama.cpp mainline has the `--spec-type draft-mtp` flag but not the
114
+ Qwen3.8-Flash-Next MTP graph, and not the loader that lets the *shared* heads borrow
115
+ `token_embd` / `output` from the main model. Any of these three works (same list as Unsloth's
116
+ `MTP/README.md`):
117
+
118
+ 1. **Unsloth's fork, release tag `b10715-mix-86bd2d3` or newer** — https://github.com/unslothai/llama.cpp/releases.
119
+ The prebuilt binaries there are **CUDA and CPU only**. For **ROCm** (and for Windows) take the
120
+ *source tarball* of that tag and build it exactly like you build mainline (`-DGGML_HIP=ON` etc.).
121
+ This repository's numbers come from a `b10798-mix` source build on Windows/CUDA.
122
+ 2. Upstream PR https://github.com/ggml-org/llama.cpp/pull/28243 (the same code on its way to mainline; still a draft as of 2026-09-09).
123
+ 3. Unsloth fork PR #144.
124
+
125
+ On a correct build the shared head prints **one error line at startup about borrowing the
126
+ embeddings**. That line is normal; it works. The `shared-Q4_K_M` head (1.78 GiB) saves another
127
+ ~840 MiB of VRAM versus `shared-Q8_0` at the same acceptance (0.70 vs 0.72 measured), useful on 24 GB cards.
128
+
129
  The head costs **~3.2 GiB of VRAM, flat** — the same at every offload level we tried
130
  (`--n-cpu-moe` 14 / 12 / 10 / 8), so budget for it once. On a 32 GB card at 96k context,
131
  `--n-cpu-moe 10` is the balanced point (30578 MiB, 93.8%, 92.5 tok/s); `8` is the fast one
 
179
  1. `--parallel > 1` requires `--kv-unified`.
180
  2. Empty `content` with tight `max_tokens`: reasoning burns the budget. Detect and retry with a higher cap.
181
  3. First requests after a cold load pay disk; throughput climbs over ~3-4 requests to its warm ceiling. The disk cost is the PLE table (see the prefill section above); `--lazy-mode on-direct` removes it.
182
+ 4. **`--spec-type draft-mtp` exists in llama.cpp mainline but does not work here.** Mainline has the flag and not the MTP graph for `qwen4exp`, nor the tensor borrowing the *shared* heads need. Symptoms on mainline: with a `shared-*` head, `tensor 'token_embd.weight' not found` → `failed to load draft model` (the head omits the embeddings on purpose, it is not a corrupt download); with a fused model, `model doesn't contain MTP layers`. The non-shared head is not a workaround: it gets past the missing tensor and fails later for the same reason. Use one of the builds listed under "Which build" above (ROCm users: build the Unsloth tag from source).
183
  5. `"reasoning_effort": "none"` returns **HTTP 500** — the chat template accepts only `xhigh` / `medium` / `low`.
184
 
185
  ## Credits