In-file MTP speculation refused on Hadamard-folded weights (prism-b10683); ~1.34x decode after the 7-line embedding fix

#23
by zhaokeqi - opened

Hi β€” sharing a reproducible finding from running Bonsai 2 27B locally (RTX 4080 SUPER, Ada SM89, CUDA 13.3, WSL2). Short version: in-file MTP is a real ~1.34x decode win, but the official fork refuses the MTP graph on Hadamard-folded weights; a 7-line embedding-boundary fix (already in the community runtime) makes it work.

What works: MTP block packaged inside the main GGUF

ProCreations/Ternary-Bonsai-2-27B-MTP = the unchanged PQ2_0 target + blk.64.nextn.* (nextn_predict_layers=1) in one file. Launching with -m <bundle> --spec-type draft-mtp (no -md) builds the draft context against the target model (log: creating MTP draft context against the target model), so the draft shares the target's token_embd / output instead of carrying its own 248320x5120 copies.

That detail is the whole economics. A separate MTP sidecar carries a duplicate vocabulary β€” ~92% of its bytes β€” which puts the per-token draft cost ratio at rho = draft/target ~ 0.43 and makes even 40% acceptance a net loss (I measured 5.16 t/s vs 16.6 t/s no-spec on a 27B sidecar). Sharing the vocabulary drops rho to ~0.06 and the same head becomes profitable. Speedup ~ (1-a^(k+1))/(1-a) / (k*rho + 1).

What fails on the official fork

With prism-b10683-d8f26ee the same bundle refuses to start:

E llama_init_from_model: failed to initialize the context: Hadamard-latent table 'token_embd.weight' is read without the inverse transform
E common_speculative_init_result: failed to create MTP context

The guard is llama_verify_hadamard_graph; both messages are visible via strings libllama.so, and I could not find any flag to bypass it. It looks intentional: the MTP graph does ggml_get_rows(tok_embd_w, inp->tokens) on a table stored in the rotated basis.

The fix

Apply the inverse transform (rotation + explicit signs) to the MTP embedding lookup, i.e. 7 lines in src/models/qwen35.cpp::graph_mtp immediately after the ggml_get_rows. The community runtime ships exactly this as bonsai-mtp-embedding.patch, and its runtime/manifest.json says the bundled source (prism-dflash2-source.tar.gz, base d8f26ee) is already patched. I built it locally for SM89 (their prebuilt binary is SM120-only):

cmake -S llama -B llama/build -G Ninja -DGGML_CUDA=ON \
      -DCMAKE_CUDA_COMPILER=/usr/local/cuda/bin/nvcc \
      -DCMAKE_CUDA_ARCHITECTURES=89 -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF

(note: with nvcc not on PATH, CMake fails with CMAKE_CUDA_COMPILER-NOTFOUND unless the compiler is passed explicitly.)

Measured (vs --spec-type none, identical flags otherwise)

-ngl 99 -fa on -c 262144 -ctk q4_0 -ctv q4_0, temp 0, top_k 1, seed 7, n=128, 12 prompts, draft length 2:

category (n) baseline t/s MTP t/s speedup acceptance
reasoning prose (3) 66.4 84.9 1.25-1.31x 47.7-57.6%
code continuation (3) 67.4 101.3 1.32-1.70x 62.5-94.3%
math, step-by-step (2) 68.2 90.6 1.23-1.43x 54.5-72.1%
format / repetitive (2) 68.3 111.9 1.63-1.64x 90.0-92.1%
Chinese (1) 68.5 91.7 1.34x 64.5%
median 67.5 90.4 1.34x 68.1% (803/1180)

Same configuration, prefill of a ~12.5k-token prompt: 1726 t/s with PQ2_0 vs 1073 t/s with PTQ1_0 at 262144 + q4_0 KV; VRAM 15.7 / 16.0 GiB. (Consistent with the model card's "PQ2_0 wins prefill, PTQ1_0 wins decode on Ada".)

Suggestion

Please consider upstreaming the MTP embedding inverse-transform (and auditing any other latent table the draft graph reads) so that in-file MTP works on the official binaries. Today it only works with the community-patched runtime, which is SM120-only for the prebuilt archive and therefore forces non-Blackwell users to build from source. Happy to provide the patch, logs, or a repro script.

Adding the raw numbers behind the summary table above, plus a public reproduction bundle.

Public reproduction bundle: https://huggingface.co/datasets/zhaokeqi/bonsai2-27b-mtp-repro

It contains the raw per-prompt JSON for both arms (with the full timings objects), the A/B harness, the driver scripts and these notes, so the numbers above can be checked rather than trusted.

Measurement protocol (identical in both arms except --spec-type):

llama-server -m Ternary-Bonsai-2-27B-PQ2_0-MTP-Q8_0.gguf \
  -ngl 99 -fa on -c 262144 -ctk q4_0 -ctv q4_0 -np 1 -t 16 --temp 0 \
  --spec-type none                             # arm A
  --spec-type draft-mtp --spec-draft-n-max 2   # arm B

POST /completion, temperature=0, top_k=1, seed=7, cache_prompt=false, n_predict=128. Acceptance is read from the response timings.draft_n / draft_n_accepted β€” note that llama-cli does not print acceptance at all, only the server does (or the log line slot print_timing: draft acceptance = ...).

Per-prompt results (single run per prompt β€” no repeats, so read the medians, not the last digit):

# prompt baseline t/s draft-mtp t/s speedup acceptance
R1 reasoning prose 67.2 88.0 1.31x 57.6%
R2 reasoning prose 65.4 82.0 1.25x 47.7%
R3 reasoning prose 66.6 84.6 1.27x 57.6%
C1 Python continuation 67.1 99.8 1.49x 77.0%
C2 Python continuation 67.4 114.8 1.70x 94.3%
C3 async Python 67.7 89.2 1.32x 62.5%
M1 step-by-step math 68.1 97.2 1.43x 72.1%
M2 probability recursion 68.2 84.0 1.23x 54.5%
F1 JSON repetition 68.3 111.5 1.63x 90.0%
F2 list continuation 68.3 112.3 1.64x 92.1%
Z1 Chinese rewrite 68.5 91.7 1.34x 64.5%
median 67.5 90.4 1.338x 68.1% (803/1180)

Two practical notes for 16 GB cards

  1. At -c 262144 on a 16 GB Ada card, the KV cache type decides everything: q8_0 makes prefill collapse to 101 -> 35 t/s (one 10,240-token prefill took 293 s), while q4_0 gives 1726 t/s. The common_fit_params: failed to fit params to free device memory warning shows up in both states, so it is not the signal β€” KV bytes are. VRAM in the q4_0 state is 15.7/16.0 GiB.
  2. Prefix reuse works within a conversation: re-sending an identical 12,485-token prompt evaluates only 4 tokens (0.28 s). Switching conversations re-prefills, since -np 1 keeps one slot.

Of the 12 prompts, one (a second Chinese prompt) stopped after 1 token with stop_type=eos and empty content in both arms β€” a raw-/completion-without-chat-template artifact, not a model defect. It is kept in the raw JSON and excluded from the statistics.

Not measured: the official 14-benchmark suite, repeats per prompt, contexts beyond 262,144, CPU-only runs, and the DFlash2 head.

Same data as a chart (baseline vs MTP throughput per prompt, and the acceptance rate that explains it):

A/B results

The pattern to read off it: the speedup tracks acceptance, and acceptance tracks how predictable the continuation is β€” Python/list/JSON continuations land at 77-94% acceptance and 1.5-1.7x, free-form reasoning prose at 48-58% and 1.25-1.3x.

Code and method now live on GitHub β€” https://github.com/zhaoyilun/bonsai2-27b-mtp-repro β€” with the build script (SM89), the A/B harness, the launch units and the full write-up. This dataset stays the raw-data half of the same work, and the two link to each other.

Also added long-context numbers (chat path, natural text, depths exact via POST /tokenize, -c 262144 + KV q4_0):

depth prefill t/s decode t/s (MTP) acceptance decode t/s (no spec) speedup
8,020 1855 86.9 65.8% 63.0 1.38x
31,939 1674 62.4 52.2% 53.7 1.16x
64,084 1332 55.4 67.9% 43.4 1.28x
127,870 974 46.1 82.3% 33.3 1.38x
191,099 722 35.0 84.1% 26.5 1.32x

Decode falls hard with depth (KV traffic), but MTP's edge does not decay β€” acceptance rises from 66% to 84%, so speculation buys more exactly where the target is slowest. One measurement gotcha worth knowing: timings.prompt_n reports only the newly evaluated tokens (a 191k-token prompt can come back as prompt_n = 63745) because llama-server reuses cached context checkpoints β€” the log shows n_tokens = 191139, truncated = 0, so nothing is dropped. A short "cache buster" request does not evict them; use /tokenize for the real length.

Filed the fix upstream as a pull request: https://github.com/PrismML-Eng/llama.cpp/pull/205 β€” qwen35: apply the Hadamard inverse to the MTP token-embedding lookup (1 file, +14 lines; it mirrors what the trunk already does in llm_graph_context::build_inp_embd, keyed on the table actually looked up).

Two things checked while preparing it:

  • The bug is not version-specific. I downloaded the newest official release prism-b10709-9a9394a (2026-09-18, ~26 builds newer than the b10683 I started from) and it still refuses the MTP context β€” with an extra diagnostic line now (latent lookup 'mtp_tok_embd-64' consumed by op=RMS_NORM name='norm-64').
  • The fix is semantically exact. Built from the patched source for sm_89, draft acceptance on three probes is byte-for-byte the same as the community-patched build: 49/90, 59/71, 57/75 (54.4% / 83.1% / 76.0%).

Credit for finding it goes to runtime/bonsai-mtp-embedding.patch in ProCreations/Ternary-Bonsai-2-27B-MTP; the PR is the equivalent against current prism, so nothing here is claimed as new beyond rebasing it onto today's guard.

Also left supporting measurements on two related open issues from the same run: #203 (ngram-* silent no-op β€” independent CUDA corroboration) and #85 (q4_0 K-cache at 262144 on a 16 GB card without --kv-mean-center, plus the q8_0 failure mode).

Sign up or log in to comment