File size: 2,175 Bytes
944fda8
 
 
 
 
 
0a3dba9
944fda8
 
 
1ea961e
3d7a4fb
c462057
 
 
944fda8
0a3dba9
944fda8
 
 
 
 
 
 
 
 
 
3d7a4fb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
944fda8
 
 
c462057
944fda8
 
 
c462057
944fda8
 
 
 
3d7a4fb
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
# Gemma 4 26B-A4B MTP drafter

Multi-Token Prediction (MTP) drafter for `unsloth/gemma-4-26B-A4B-it-GGUF`. It runs as a speculative draft model that shares the target's KV cache and speeds up text generation.

Verified on a single B200 with the `gemma-4-26B-A4B-it-UD-Q4_K_M.gguf` target: 159 tok/s without MTP, 257 tok/s with MTP, 0.72 draft acceptance.

MTP was merged into llama.cpp on 2026-06-07 (PR ggml-org/llama.cpp#23398). You need a llama.cpp build from after that date. Older builds cannot load these (arch `gemma4-assistant`).

## Files

For `-hf` auto-discovery a Q8_0 drafter sits at the repo root as `mtp-gemma-4-26B-A4B-it.gguf`. The same three precisions live in `MTP/`:

- `mtp-gemma-4-26B-A4B-it-Q8_0.gguf` (smallest, recommended; mirrored at the repo root as `mtp-gemma-4-26B-A4B-it.gguf`)
- `mtp-gemma-4-26B-A4B-it-BF16.gguf`
- `mtp-gemma-4-26B-A4B-it-F16.gguf`

## Build llama.cpp

```bash
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp

# CUDA build. Set the arch for your GPU: 89 (RTX 4090), 90 (H100), 100 (B200).
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=90
cmake --build build --config Release -j --target llama-server
```

## Run, the easy way

A recent llama.cpp finds the drafter automatically from the root `mtp-` file, so `-hf` is all you need. No `--model-draft`.

```bash
./build/bin/llama-server \
  -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_M \
  --spec-type draft-mtp --spec-draft-n-max 4 \
  -ngl 999 -fa on
```

If your build is too old to auto-discover the sibling, use the explicit form below.

## Run with an explicit drafter

Use this to choose a precision or point at a local file.

```bash
hf download unsloth/gemma-4-26B-A4B-it-GGUF gemma-4-26B-A4B-it-UD-Q4_K_M.gguf --local-dir .
hf download unsloth/gemma-4-26B-A4B-it-GGUF MTP/mtp-gemma-4-26B-A4B-it-Q8_0.gguf --local-dir .

./build/bin/llama-server \
  -m gemma-4-26B-A4B-it-UD-Q4_K_M.gguf \
  --model-draft MTP/mtp-gemma-4-26B-A4B-it-Q8_0.gguf \
  --spec-type draft-mtp --spec-draft-n-max 4 \
  -ngl 999 -fa on
```

Multi GPU: add `--spec-draft-device CUDA0 -sm layer`. The drafter pairs with any quant of the 26B-A4B. Quantized KV cache works.