Document the MTP heads and vision support
Browse files
README.md
CHANGED
|
@@ -21,16 +21,16 @@ tags:
|
|
| 21 |
machine. Built and tested on an NVIDIA DGX Spark (GB10).
|
| 22 |
|
| 23 |
The repo also includes Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and
|
| 24 |
-
quantized to 4 bpw.
|
| 25 |
|
| 26 |
| | |
|
| 27 |
|---|---|
|
| 28 |
-
| Weights |
|
| 29 |
| Bitrate | 2.27 bpw (excluding head), head 6 bpw |
|
| 30 |
| Perplexity | 5.40, wikitext-2 test, 64 x 2048 tokens |
|
| 31 |
-
| Decode, batch 1 | ~31 tok/s without drafter;
|
| 32 |
-
| Context | 262,144 tokens with the drafter
|
| 33 |
-
| Modalities | Text
|
| 34 |
|
| 35 |
Requires exllamav3 v1.5.2 or later; see [How to run](#how-to-run).
|
| 36 |
|
|
@@ -39,7 +39,7 @@ Requires exllamav3 v1.5.2 or later; see [How to run](#how-to-run).
|
|
| 39 |
```
|
| 40 |
model-*.safetensors, model.safetensors.index.json EXL3 weights
|
| 41 |
quantization_config.json per-tensor storage record
|
| 42 |
-
config.json, tokenizer files, chat_template.jinja
|
| 43 |
dflash/ drafter, 4 bpw EXL3 (use this one)
|
| 44 |
dflash-bf16/ the same drafter, unquantized
|
| 45 |
eval/ benchmark outputs
|
|
@@ -56,7 +56,8 @@ Converted with `convert.py -b 2.25 -hq -cr 250 -cc 2048`. The per-module result:
|
|
| 56 |
| attention | 4.0 |
|
| 57 |
| dense MLP (layer 0) | 3.0 |
|
| 58 |
| `lm_head` | 6.0 |
|
| 59 |
-
|
|
|
|
|
| 60 |
|
| 61 |
**Layer 47:** some of its experts produce intermediate values past the fp16 limit.
|
| 62 |
This build scales `up_proj` down by 128 in that layer (`interm_div`) and restores the scale in
|
|
@@ -84,15 +85,19 @@ Per-task scores and run settings are in [`eval/bench/`](eval/bench/).
|
|
| 84 |
|
| 85 |
## Speed
|
| 86 |
|
| 87 |
-
DGX Spark, batch 1, greedy, 512 generated tokens, median of 3 runs
|
|
|
|
| 88 |
|
| 89 |
-
| prompt | no drafter |
|
| 90 |
|---|---|---|---|
|
| 91 |
-
| coding | 31.
|
| 92 |
-
| prose | 31.
|
| 93 |
-
| reasoning | 31.
|
|
|
|
| 94 |
|
| 95 |
-
|
|
|
|
|
|
|
| 96 |
|
| 97 |
Long context with the BF16 drafter, generating ~400 tokens of code against a large
|
| 98 |
repository prompt:
|
|
@@ -152,6 +157,17 @@ draft_model:
|
|
| 152 |
dynamic_draft: true
|
| 153 |
```
|
| 154 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 155 |
Thinking can be turned off per request with `"chat_template_kwargs": {"enable_thinking": false}`.
|
| 156 |
Tool calls use the `qwen3_coder` format, which TabbyAPI detects automatically.
|
| 157 |
|
|
|
|
| 21 |
machine. Built and tested on an NVIDIA DGX Spark (GB10).
|
| 22 |
|
| 23 |
The repo also includes Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and
|
| 24 |
+
quantized to 4 bpw, and the model's own MTP heads and vision tower.
|
| 25 |
|
| 26 |
| | |
|
| 27 |
|---|---|
|
| 28 |
+
| Weights | 85.28 GiB, 12 shards, including the MTP heads and the vision tower |
|
| 29 |
| Bitrate | 2.27 bpw (excluding head), head 6 bpw |
|
| 30 |
| Perplexity | 5.40, wikitext-2 test, 64 x 2048 tokens |
|
| 31 |
+
| Decode, batch 1 | ~31 tok/s without a drafter; 35–80 tok/s with the DFlash drafter (see [Speed](#speed)) |
|
| 32 |
+
| Context | 262,144 tokens on a DGX Spark, with the drafter and vision loaded |
|
| 33 |
+
| Modalities | Text and images (no audio) |
|
| 34 |
|
| 35 |
Requires exllamav3 v1.5.2 or later; see [How to run](#how-to-run).
|
| 36 |
|
|
|
|
| 39 |
```
|
| 40 |
model-*.safetensors, model.safetensors.index.json EXL3 weights
|
| 41 |
quantization_config.json per-tensor storage record
|
| 42 |
+
config.json, tokenizer files, chat_template.jinja, preprocessor_config.json
|
| 43 |
dflash/ drafter, 4 bpw EXL3 (use this one)
|
| 44 |
dflash-bf16/ the same drafter, unquantized
|
| 45 |
eval/ benchmark outputs
|
|
|
|
| 56 |
| attention | 4.0 |
|
| 57 |
| dense MLP (layer 0) | 3.0 |
|
| 58 |
| `lm_head` | 6.0 |
|
| 59 |
+
| MTP heads | 4.0 |
|
| 60 |
+
| embeddings, norms, router, vision tower | BF16 |
|
| 61 |
|
| 62 |
**Layer 47:** some of its experts produce intermediate values past the fp16 limit.
|
| 63 |
This build scales `up_proj` down by 128 in that layer (`interm_div`) and restores the scale in
|
|
|
|
| 85 |
|
| 86 |
## Speed
|
| 87 |
|
| 88 |
+
exllamav3 v1.5.2 on a DGX Spark, batch 1, greedy, 512 generated tokens, median of 3 runs, in
|
| 89 |
+
tok/s:
|
| 90 |
|
| 91 |
+
| prompt | no drafter | DFlash drafter | MTP heads |
|
| 92 |
|---|---|---|---|
|
| 93 |
+
| coding | 31.5 | 49.5 | 40.7 |
|
| 94 |
+
| prose | 31.3 | 35.4 | 31.8 |
|
| 95 |
+
| reasoning | 31.2 | 61.1 | 45.1 |
|
| 96 |
+
| code edit | 31.0 | 80.7 | 51.3 |
|
| 97 |
|
| 98 |
+
DFlash is the faster drafter on every prompt. The MTP heads draft 3 tokens per step; asking for
|
| 99 |
+
6 was at most 5% faster. Prose gains the least and varies the most from prompt to prompt. The target model verifies every drafted token, so neither drafter
|
| 100 |
+
affects output quality.
|
| 101 |
|
| 102 |
Long context with the BF16 drafter, generating ~400 tokens of code against a large
|
| 103 |
repository prompt:
|
|
|
|
| 157 |
dynamic_draft: true
|
| 158 |
```
|
| 159 |
|
| 160 |
+
To use the MTP heads instead of the DFlash drafter:
|
| 161 |
+
|
| 162 |
+
```yaml
|
| 163 |
+
draft_model:
|
| 164 |
+
draft_mode: mtp
|
| 165 |
+
dynamic_draft: true
|
| 166 |
+
```
|
| 167 |
+
|
| 168 |
+
For image input, add `vision: true` under `model:`. The vision tower adds about 1.3 GiB; with it
|
| 169 |
+
and the DFlash drafter loaded at 262,144 context, a 250K-token prompt still left 15 GiB free.
|
| 170 |
+
|
| 171 |
Thinking can be turned off per request with `"chat_template_kwargs": {"enable_thinking": false}`.
|
| 172 |
Tool calls use the `qwen3_coder` format, which TabbyAPI detects automatically.
|
| 173 |
|