Capicua25x commited on
Commit
4c16e9b
Β·
verified Β·
1 Parent(s): e7267fe

model card

Browse files
Files changed (1) hide show
  1. README.md +118 -0
README.md ADDED
@@ -0,0 +1,118 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ base_model: deepreinforce-ai/Ornith-1.0-9B
4
+ base_model_relation: quantized
5
+ pipeline_tag: image-text-to-text
6
+ library_name: vllm
7
+ language:
8
+ - en
9
+ tags:
10
+ - mxfp4
11
+ - compressed-tensors
12
+ - speculative-decoding
13
+ - mtp
14
+ - multi-token-prediction
15
+ - qwen3.5
16
+ - vllm
17
+ - rocm
18
+ - rdna4
19
+ - amd
20
+ - vision
21
+ ---
22
+
23
+ # Ornith-1.0-9B β€” MXFP4 + MTP (vision), for AMD RDNA4 / vLLM
24
+
25
+ [`deepreinforce-ai/Ornith-1.0-9B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-9B) quantized to
26
+ **MXFP4** with a **grafted MTP (Multi-Token-Prediction) draft head**, packaged to run **out of the box on
27
+ AMD Radeon RDNA4 (gfx1201) under vLLM** with lossless self-speculative decoding. Vision retained.
28
+
29
+ - **Trunk** β€” Ornith-1.0-9B quantized to **MXFP4** (`compressed-tensors`, group size 32, symmetric).
30
+ `lm_head`, `embed_tokens`, norms, and the vision tower are kept **BF16**.
31
+ - **Draft head** β€” the **KL-distilled MTP head** from
32
+ [`protoLabsAI/Ornith-1.0-9B-MTP`](https://huggingface.co/protoLabsAI/Ornith-1.0-9B-MTP) (BF16), grafted
33
+ into a separate `model-mtp.safetensors` shard and marked unquantized in the quant config so the
34
+ mixed-precision model loads cleanly.
35
+ - **Speculative decoding** β€” vLLM native `mtp` method, `num_speculative_tokens=3`. **Lossless**: the
36
+ target verifies every drafted token, so the output distribution is unchanged β€” the head only buys speed.
37
+
38
+ ## Measured on AMD RDNA4
39
+
40
+ 2Γ— Radeon AI PRO R9700 (gfx1201), tensor-parallel 2, vLLM `0.19.1` (image below).
41
+
42
+ - **MTP draft acceptance (n=3): β‰ˆ66% average** across a mixed code + 6k-context run β€” per-position
43
+ **0.82 / 0.65 / 0.53**, mean acceptance length **~3.0** (max 4) β€” rising to **~85%** (length ~3.6) on
44
+ cache-warm short context. Lossless throughout.
45
+ - **Tool-calling** (`qwen3_xml`) and **vision** confirmed working.
46
+ - **Throughput** (TP2, 256 output tokens, per-user / aggregate tok/s):
47
+
48
+ | concurrent | short prompt | ~6k prompt |
49
+ |---:|---|---|
50
+ | 1 | 84.6 / 85 | 103.8 / 104 |
51
+ | 16 | 51.8 / 803 | 50.1 / 767 |
52
+ | 32 | 40.8 / 1260 | 33.0 / 1010 |
53
+ | 64 | 30.7 / 1902 | 21.1 / 1268 |
54
+ | 96 | 22.5 / 2075 | 15.0 / 1293 |
55
+ | 128 | 20.2 / 1893 | 12.6 / 1284 |
56
+
57
+ Usable concurrency ceiling (per-user β‰₯ 20 tok/s): **~128** at short context, **~64** at 6k context.
58
+ Single-stream decode is bound by the dense 9B's active-parameter count; a single GPU (TP1) is also supported.
59
+
60
+ ## Notes from the publisher
61
+
62
+ On AMD RDNA4, **tool-calling and core agentic tool-use were a clear, reliable step up** in our testing β€”
63
+ that's what this build is good at. We did **not** independently evaluate complex long-horizon reasoning;
64
+ evaluate on your own tasks. Base-model capabilities and limitations are
65
+ [DeepReinforce](https://huggingface.co/deepreinforce-ai)'s to characterize.
66
+
67
+ ## Run it on AMD RDNA4
68
+
69
+ Uses the prebuilt RDNA4 vLLM image
70
+ [`capicua25x/vllm-rocm-rdna4`](https://hub.docker.com/r/capicua25x/vllm-rocm-rdna4) (tag `0.19.1`):
71
+
72
+ ```bash
73
+ docker run --rm --network=host \
74
+ --device=/dev/kfd --device=/dev/dri \
75
+ --group-add=video --group-add=render --ipc=host --ulimit memlock=-1 \
76
+ capicua25x/vllm-rocm-rdna4:0.19.1 \
77
+ --model Capicua25x/Ornith-1.0-9B-MXFP4-Vision-MTP \
78
+ --served-model-name ornith --trust-remote-code \
79
+ --tensor-parallel-size 2 \
80
+ --gpu-memory-utilization 0.90 \
81
+ --max-model-len 16384 \
82
+ --attention-backend TRITON_ATTN \
83
+ --enable-prefix-caching \
84
+ --enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3 \
85
+ --speculative-config '{"method":"mtp","num_speculative_tokens":3}'
86
+ ```
87
+
88
+ ### The settings that actually matter on RDNA4
89
+ - `--attention-backend TRITON_ATTN` β€” required on gfx1201.
90
+ - `--speculative-config '{"method":"mtp","num_speculative_tokens":3}'` β€” enables the grafted MTP head.
91
+ **n=3 maximizes throughput; n=1–2 maximize per-token acceptance.** Tune per workload.
92
+ - `--tool-call-parser qwen3_xml --reasoning-parser qwen3` β€” Qwen3.5-family tool-calling + reasoning split.
93
+ - `--trust-remote-code` β€” the `qwen3_5` vision architecture.
94
+ - **Single GPU works too**: `--tensor-parallel-size 1` and pass one render node
95
+ (e.g. `--device=/dev/dri/renderD128`).
96
+
97
+ ## How it was built (reproducible)
98
+ 1. **MXFP4 quantize** Ornith-1.0-9B with `compressed-tensors` (4-bit float, group 32, symmetric;
99
+ `lm_head` / `embed_tokens` / norms / vision tower left BF16).
100
+ 2. **Graft** the 15 `mtp.*` head tensors from `protoLabsAI/Ornith-1.0-9B-MTP` into a new
101
+ `model-mtp.safetensors` shard and patch `model.safetensors.index.json`.
102
+ 3. **Mark the head unquantized** β€” add its Linear modules (`mtp.fc`, `mtp.layers.0.self_attn.*`,
103
+ `mtp.layers.0.mlp.*`) to `quantization_config.ignore`, so vLLM's compressed-tensors loader keeps the
104
+ BF16 head as-is instead of expecting MXFP4 weight-scales. This is the one mixed-precision gotcha.
105
+
106
+ Steps 2–3 are scripted in [`recipe_graft_mxfp4.py`](./recipe_graft_mxfp4.py) (run against an MXFP4
107
+ `compressed-tensors` trunk + the protoLabs head). The head's distillation recipe lives upstream at
108
+ [`protoLabsAI/Ornith-1.0-9B-MTP`](https://huggingface.co/protoLabsAI/Ornith-1.0-9B-MTP).
109
+
110
+ ## Credits
111
+ - **DeepReinforce** β€” [`Ornith-1.0-9B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-9B), the base model (MIT).
112
+ - **protoLabs** β€” [`Ornith-1.0-9B-MTP`](https://huggingface.co/protoLabsAI/Ornith-1.0-9B-MTP), the KL-distilled MTP draft head and its recipe (MIT).
113
+ - **Qwen / Alibaba** β€” the Qwen3.5 architecture the MTP head derives from (the head was initialized from `Qwen/Qwen3.5-9B`'s `mtp.*` tensors).
114
+ - **vLLM** and **compressed-tensors** β€” serving stack and quantization format.
115
+ - RDNA4 vLLM image: [`capicua25x/vllm-rocm-rdna4`](https://hub.docker.com/r/capicua25x/vllm-rocm-rdna4).
116
+
117
+ ## License
118
+ **MIT.** This is a derivative of Ornith-1.0-9B (MIT); merging the MTP head (MIT) produces a derivative whose MIT terms carry.