benthecarman commited on
Commit
5ae87b8
·
verified ·
1 Parent(s): 0ba71f4

Document the MTP heads and vision support

Browse files
Files changed (1) hide show
  1. README.md +29 -13
README.md CHANGED
@@ -21,16 +21,16 @@ tags:
21
  machine. Built and tested on an NVIDIA DGX Spark (GB10).
22
 
23
  The repo also includes Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and
24
- quantized to 4 bpw.
25
 
26
  | | |
27
  |---|---|
28
- | Weights | 83.45 GiB, 12 shards |
29
  | Bitrate | 2.27 bpw (excluding head), head 6 bpw |
30
  | Perplexity | 5.40, wikitext-2 test, 64 x 2048 tokens |
31
- | Decode, batch 1 | ~31 tok/s without drafter; 40–60 tok/s with the 4 bpw drafter (see [Speed](#speed)) |
32
- | Context | 262,144 tokens with the drafter on a DGX Spark |
33
- | Modalities | Text only (no vision, audio or MTP heads) |
34
 
35
  Requires exllamav3 v1.5.2 or later; see [How to run](#how-to-run).
36
 
@@ -39,7 +39,7 @@ Requires exllamav3 v1.5.2 or later; see [How to run](#how-to-run).
39
  ```
40
  model-*.safetensors, model.safetensors.index.json EXL3 weights
41
  quantization_config.json per-tensor storage record
42
- config.json, tokenizer files, chat_template.jinja
43
  dflash/ drafter, 4 bpw EXL3 (use this one)
44
  dflash-bf16/ the same drafter, unquantized
45
  eval/ benchmark outputs
@@ -56,7 +56,8 @@ Converted with `convert.py -b 2.25 -hq -cr 250 -cc 2048`. The per-module result:
56
  | attention | 4.0 |
57
  | dense MLP (layer 0) | 3.0 |
58
  | `lm_head` | 6.0 |
59
- | embeddings, norms, router | BF16 |
 
60
 
61
  **Layer 47:** some of its experts produce intermediate values past the fp16 limit.
62
  This build scales `up_proj` down by 128 in that layer (`interm_div`) and restores the scale in
@@ -84,15 +85,19 @@ Per-task scores and run settings are in [`eval/bench/`](eval/bench/).
84
 
85
  ## Speed
86
 
87
- DGX Spark, batch 1, greedy, 512 generated tokens, median of 3 runs:
 
88
 
89
- | prompt | no drafter | 4 bpw drafter | speedup |
90
  |---|---|---|---|
91
- | coding | 31.4 tok/s | 48.5 tok/s | 1.54x |
92
- | prose | 31.4 tok/s | 40.2 tok/s | 1.28x |
93
- | reasoning | 31.4 tok/s | 59.5 tok/s | 1.89x |
 
94
 
95
- The target model verifies every drafted token, so the drafter does not affect output quality.
 
 
96
 
97
  Long context with the BF16 drafter, generating ~400 tokens of code against a large
98
  repository prompt:
@@ -152,6 +157,17 @@ draft_model:
152
  dynamic_draft: true
153
  ```
154
 
 
 
 
 
 
 
 
 
 
 
 
155
  Thinking can be turned off per request with `"chat_template_kwargs": {"enable_thinking": false}`.
156
  Tool calls use the `qwen3_coder` format, which TabbyAPI detects automatically.
157
 
 
21
  machine. Built and tested on an NVIDIA DGX Spark (GB10).
22
 
23
  The repo also includes Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and
24
+ quantized to 4 bpw, and the model's own MTP heads and vision tower.
25
 
26
  | | |
27
  |---|---|
28
+ | Weights | 85.28 GiB, 12 shards, including the MTP heads and the vision tower |
29
  | Bitrate | 2.27 bpw (excluding head), head 6 bpw |
30
  | Perplexity | 5.40, wikitext-2 test, 64 x 2048 tokens |
31
+ | Decode, batch 1 | ~31 tok/s without a drafter; 35–80 tok/s with the DFlash drafter (see [Speed](#speed)) |
32
+ | Context | 262,144 tokens on a DGX Spark, with the drafter and vision loaded |
33
+ | Modalities | Text and images (no audio) |
34
 
35
  Requires exllamav3 v1.5.2 or later; see [How to run](#how-to-run).
36
 
 
39
  ```
40
  model-*.safetensors, model.safetensors.index.json EXL3 weights
41
  quantization_config.json per-tensor storage record
42
+ config.json, tokenizer files, chat_template.jinja, preprocessor_config.json
43
  dflash/ drafter, 4 bpw EXL3 (use this one)
44
  dflash-bf16/ the same drafter, unquantized
45
  eval/ benchmark outputs
 
56
  | attention | 4.0 |
57
  | dense MLP (layer 0) | 3.0 |
58
  | `lm_head` | 6.0 |
59
+ | MTP heads | 4.0 |
60
+ | embeddings, norms, router, vision tower | BF16 |
61
 
62
  **Layer 47:** some of its experts produce intermediate values past the fp16 limit.
63
  This build scales `up_proj` down by 128 in that layer (`interm_div`) and restores the scale in
 
85
 
86
  ## Speed
87
 
88
+ exllamav3 v1.5.2 on a DGX Spark, batch 1, greedy, 512 generated tokens, median of 3 runs, in
89
+ tok/s:
90
 
91
+ | prompt | no drafter | DFlash drafter | MTP heads |
92
  |---|---|---|---|
93
+ | coding | 31.5 | 49.5 | 40.7 |
94
+ | prose | 31.3 | 35.4 | 31.8 |
95
+ | reasoning | 31.2 | 61.1 | 45.1 |
96
+ | code edit | 31.0 | 80.7 | 51.3 |
97
 
98
+ DFlash is the faster drafter on every prompt. The MTP heads draft 3 tokens per step; asking for
99
+ 6 was at most 5% faster. Prose gains the least and varies the most from prompt to prompt. The target model verifies every drafted token, so neither drafter
100
+ affects output quality.
101
 
102
  Long context with the BF16 drafter, generating ~400 tokens of code against a large
103
  repository prompt:
 
157
  dynamic_draft: true
158
  ```
159
 
160
+ To use the MTP heads instead of the DFlash drafter:
161
+
162
+ ```yaml
163
+ draft_model:
164
+ draft_mode: mtp
165
+ dynamic_draft: true
166
+ ```
167
+
168
+ For image input, add `vision: true` under `model:`. The vision tower adds about 1.3 GiB; with it
169
+ and the DFlash drafter loaded at 262,144 context, a 250K-token prompt still left 15 GiB free.
170
+
171
  Thinking can be turned off per request with `"chat_template_kwargs": {"enable_thinking": false}`.
172
  Tool calls use the `qwen3_coder` format, which TabbyAPI detects automatically.
173