credits: the MTP head is Qwen's own trained block, not nerkyor's — byte-verified, and name pahajokiconsulting as the donor it was copied from
Browse files
README.md
CHANGED
|
@@ -25,7 +25,7 @@ tags:
|
|
| 25 |
|
| 26 |
# Ornith-1.0-35B-A3B — MXFP4 + MTP (vision), for AMD RDNA4 / vLLM
|
| 27 |
|
| 28 |
-
[`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/
|
| 29 |
**Qwen MoE model (A3B, ~3 B active parameters per token)** — quantized to **MXFP4** with a **grafted MTP
|
| 30 |
(Multi-Token-Prediction) draft head**, packaged to run **out of the box on AMD Radeon RDNA4 (gfx1201)
|
| 31 |
under vLLM** with lossless self-speculative decoding. Vision retained.
|
|
@@ -36,7 +36,10 @@ under vLLM** with lossless self-speculative decoding. Vision retained.
|
|
| 36 |
- **Draft head — a cross-model graft.** The MoE MTP head (785 BF16 tensors, with per-expert weights) is
|
| 37 |
transplanted from a **sibling Qwen-MoE checkpoint**,
|
| 38 |
[`Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4-Vision`](https://huggingface.co/Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4-Vision)
|
| 39 |
-
(a DeepSeek-V4-Pro-Thinking distill of Qwen3.6-35B-A3B).
|
|
|
|
|
|
|
|
|
|
| 40 |
**same Qwen MoE architecture** — identical hidden size and expert layout — the head transfers cleanly
|
| 41 |
and drafts well on the Ornith trunk (acceptance below).
|
| 42 |
- **Speculative decoding** — vLLM native `mtp` method, `num_speculative_tokens=3`. **Lossless**: the
|
|
@@ -116,10 +119,11 @@ Step 1 (the MXFP4 quantize) is scripted in [`quantize_mxfp4.py`](./quantize_mxfp
|
|
| 116 |
script handles both a dense head and this MoE head.
|
| 117 |
|
| 118 |
## Credits
|
| 119 |
-
- **DeepReinforce** — [`Ornith-1.0-35B`](https://huggingface.co/
|
| 120 |
-
- **nerkyor** — [`Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill`](https://huggingface.co/nerkyor/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill),
|
| 121 |
- **DeepSeek** — DeepSeek-V4-Pro-Thinking, the model that distill learned from.
|
| 122 |
-
- **
|
|
|
|
| 123 |
- **vLLM** and **compressed-tensors** — serving stack and quantization format.
|
| 124 |
- **Rob Smith** (`tcclaviger`) — the [RDNA4 vLLM base image](https://hub.docker.com/r/tcclaviger/vllm-rocm-mxfp4-nvfp4) that made gfx1201 serving possible. The [`capicua25x/vllm-rocm-rdna4`](https://hub.docker.com/r/capicua25x/vllm-rocm-rdna4) image this model runs on is a forward-port of his work — without it, none of this runs.
|
| 125 |
|
|
|
|
| 25 |
|
| 26 |
# Ornith-1.0-35B-A3B — MXFP4 + MTP (vision), for AMD RDNA4 / vLLM
|
| 27 |
|
| 28 |
+
[`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/ornith-ai/Ornith-1.0-35B) — a
|
| 29 |
**Qwen MoE model (A3B, ~3 B active parameters per token)** — quantized to **MXFP4** with a **grafted MTP
|
| 30 |
(Multi-Token-Prediction) draft head**, packaged to run **out of the box on AMD Radeon RDNA4 (gfx1201)
|
| 31 |
under vLLM** with lossless self-speculative decoding. Vision retained.
|
|
|
|
| 36 |
- **Draft head — a cross-model graft.** The MoE MTP head (785 BF16 tensors, with per-expert weights) is
|
| 37 |
transplanted from a **sibling Qwen-MoE checkpoint**,
|
| 38 |
[`Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4-Vision`](https://huggingface.co/Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4-Vision)
|
| 39 |
+
(a DeepSeek-V4-Pro-Thinking distill of Qwen3.6-35B-A3B). The head weights originate with
|
| 40 |
+
[`Qwen/Qwen3.6-35B-A3B`](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) (Apache-2.0) and reach this build unchanged
|
| 41 |
+
via [`pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4`](https://huggingface.co/pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4) —
|
| 42 |
+
every checkpoint in that chain carries them byte-identical. Because the trunk and the donor share the
|
| 43 |
**same Qwen MoE architecture** — identical hidden size and expert layout — the head transfers cleanly
|
| 44 |
and drafts well on the Ornith trunk (acceptance below).
|
| 45 |
- **Speculative decoding** — vLLM native `mtp` method, `num_speculative_tokens=3`. **Lossless**: the
|
|
|
|
| 119 |
script handles both a dense head and this MoE head.
|
| 120 |
|
| 121 |
## Credits
|
| 122 |
+
- **DeepReinforce** — [`Ornith-1.0-35B`](https://huggingface.co/ornith-ai/Ornith-1.0-35B), the base/trunk model (MIT). *(The `deepreinforce-ai/` id now redirects to `ornith-ai/`.)*
|
| 123 |
+
- **nerkyor** — [`Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill`](https://huggingface.co/nerkyor/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill), the DeepSeek-V4-Pro-Thinking **trunk distill**. *(Corrected 2026-08-17: this card previously credited the distill with training the MTP head. It did not — the head is Qwen's, carried through unchanged. Verified by hash: `mtp.fc.weight` [2048, 4096] and the other `mtp.*` tensors are byte-identical between `Qwen/Qwen3.6-35B-A3B` and the distill.)*
|
| 124 |
- **DeepSeek** — DeepSeek-V4-Pro-Thinking, the model that distill learned from.
|
| 125 |
+
- **pahajokiconsulting** — [`Qwen3.6-35B-A3B-MXFP4`](https://huggingface.co/pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4) (Apache-2.0), the build the 785 `mtp.*` tensors were copied from verbatim: it quantizes the trunk but keeps `mtp.*` in BF16, with the experts relaid from stacked 3D into per-expert 2D.
|
| 126 |
+
- **Qwen / Alibaba** — the Qwen MoE (A3B) architecture shared by trunk and head, **and the trained MTP head weights themselves** ([`Qwen/Qwen3.6-35B-A3B`](https://huggingface.co/Qwen/Qwen3.6-35B-A3B), Apache-2.0). The grafted head is Qwen's own trained block, not a third party's.
|
| 127 |
- **vLLM** and **compressed-tensors** — serving stack and quantization format.
|
| 128 |
- **Rob Smith** (`tcclaviger`) — the [RDNA4 vLLM base image](https://hub.docker.com/r/tcclaviger/vllm-rocm-mxfp4-nvfp4) that made gfx1201 serving possible. The [`capicua25x/vllm-rocm-rdna4`](https://hub.docker.com/r/capicua25x/vllm-rocm-rdna4) image this model runs on is a forward-port of his work — without it, none of this runs.
|
| 129 |
|