hermitdave commited on
Commit
9d0cdcf
·
verified ·
1 Parent(s): 3695817

Add MTP drafter usage instructions

Browse files
Files changed (1) hide show
  1. README.md +29 -0
README.md CHANGED
@@ -42,6 +42,35 @@ mlx_lm.generate --model hermitdave/Agnes-3.0-Flash-MLX-6bit --prompt "Hello" --m
42
 
43
  Drop the folder under `~/.lmstudio/models/hermitdave/` in LM Studio and it appears as a `qwen3_5` model.
44
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
45
  ## Attribution
46
 
47
  This conversion was produced by [Hermes Agent](https://hermes-agent.nousresearch.com) (Nous Research) — the autonomous research and conversion pipeline that identified the correct quantization parameters, fixed one-centered norm conversion, and validated output quality. Verified against the reference verison/Agnes-3.0-Flash-MLX-4bit model.
 
42
 
43
  Drop the folder under `~/.lmstudio/models/hermitdave/` in LM Studio and it appears as a `qwen3_5` model.
44
 
45
+ ## Speculative decoding with MTP drafter
46
+
47
+ A companion MTP drafter is available for speculative decoding (up to 2× faster generation):
48
+
49
+ ```bash
50
+ pip install mlx-vlm
51
+ python -m mlx_vlm.server \
52
+ --model hermitdave/Agnes-3.0-Flash-MLX-6bit \
53
+ --draft-model hermitdave/Agnes-3.0-Flash-MTP-drafter
54
+ ```
55
+
56
+ Or with the Python API:
57
+
58
+ ```python
59
+ from mlx_lm import load
60
+ from mlx_vlm.speculative.drafters.qwen3_5_mtp.config import Qwen3_5MTPConfig
61
+ from mlx_vlm.speculative.drafters.qwen3_5_mtp.qwen3_5_mtp import Qwen3_5MTPDraftModel
62
+ import json, mlx.core as mx, safetensors.torch
63
+
64
+ base_model, tokenizer = load("hermitdave/Agnes-3.0-Flash-MLX-6bit")
65
+ config = Qwen3_5MTPConfig.from_dict(json.load(open("path/to/Agnes-3.0-Flash-MTP-drafter/config.json")))
66
+ mtp = Qwen3_5MTPDraftModel(config)
67
+ weights = safetensors.torch.load_file("path/to/Agnes-3.0-Flash-MTP-drafter/model.safetensors")
68
+ mtp.load_weights([(k, mx.array(v)) for k, v in weights.items()])
69
+ mtp.bind(base_model)
70
+ ```
71
+
72
+ The drafter is architecture-compatible with Qwen3.5's gated attention and was extracted from the original Agnes-3.0-Flash model. See the [MTP drafter repo](https://huggingface.co/hermitdave/Agnes-3.0-Flash-MTP-drafter) for details.
73
+
74
  ## Attribution
75
 
76
  This conversion was produced by [Hermes Agent](https://hermes-agent.nousresearch.com) (Nous Research) — the autonomous research and conversion pipeline that identified the correct quantization parameters, fixed one-centered norm conversion, and validated output quality. Verified against the reference verison/Agnes-3.0-Flash-MLX-4bit model.