Capicua25x commited on
Commit
9ae0b43
·
verified ·
1 Parent(s): b3be14b

credits: the MTP head is Qwen's own trained block, not nerkyor's — byte-verified, and name pahajokiconsulting as the donor it was copied from

Browse files
Files changed (1) hide show
  1. README.md +9 -5
README.md CHANGED
@@ -25,7 +25,7 @@ tags:
25
 
26
  # Ornith-1.0-35B-A3B — MXFP4 + MTP (vision), for AMD RDNA4 / vLLM
27
 
28
- [`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B) — a
29
  **Qwen MoE model (A3B, ~3 B active parameters per token)** — quantized to **MXFP4** with a **grafted MTP
30
  (Multi-Token-Prediction) draft head**, packaged to run **out of the box on AMD Radeon RDNA4 (gfx1201)
31
  under vLLM** with lossless self-speculative decoding. Vision retained.
@@ -36,7 +36,10 @@ under vLLM** with lossless self-speculative decoding. Vision retained.
36
  - **Draft head — a cross-model graft.** The MoE MTP head (785 BF16 tensors, with per-expert weights) is
37
  transplanted from a **sibling Qwen-MoE checkpoint**,
38
  [`Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4-Vision`](https://huggingface.co/Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4-Vision)
39
- (a DeepSeek-V4-Pro-Thinking distill of Qwen3.6-35B-A3B). Because the trunk and the donor share the
 
 
 
40
  **same Qwen MoE architecture** — identical hidden size and expert layout — the head transfers cleanly
41
  and drafts well on the Ornith trunk (acceptance below).
42
  - **Speculative decoding** — vLLM native `mtp` method, `num_speculative_tokens=3`. **Lossless**: the
@@ -116,10 +119,11 @@ Step 1 (the MXFP4 quantize) is scripted in [`quantize_mxfp4.py`](./quantize_mxfp
116
  script handles both a dense head and this MoE head.
117
 
118
  ## Credits
119
- - **DeepReinforce** — [`Ornith-1.0-35B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B), the base/trunk model (MIT).
120
- - **nerkyor** — [`Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill`](https://huggingface.co/nerkyor/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill), which trained the MoE MTP head grafted here.
121
  - **DeepSeek** — DeepSeek-V4-Pro-Thinking, the model that distill learned from.
122
- - **Qwen / Alibaba** — the Qwen MoE (A3B) architecture shared by both the trunk and the grafted head.
 
123
  - **vLLM** and **compressed-tensors** — serving stack and quantization format.
124
  - **Rob Smith** (`tcclaviger`) — the [RDNA4 vLLM base image](https://hub.docker.com/r/tcclaviger/vllm-rocm-mxfp4-nvfp4) that made gfx1201 serving possible. The [`capicua25x/vllm-rocm-rdna4`](https://hub.docker.com/r/capicua25x/vllm-rocm-rdna4) image this model runs on is a forward-port of his work — without it, none of this runs.
125
 
 
25
 
26
  # Ornith-1.0-35B-A3B — MXFP4 + MTP (vision), for AMD RDNA4 / vLLM
27
 
28
+ [`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/ornith-ai/Ornith-1.0-35B) — a
29
  **Qwen MoE model (A3B, ~3 B active parameters per token)** — quantized to **MXFP4** with a **grafted MTP
30
  (Multi-Token-Prediction) draft head**, packaged to run **out of the box on AMD Radeon RDNA4 (gfx1201)
31
  under vLLM** with lossless self-speculative decoding. Vision retained.
 
36
  - **Draft head — a cross-model graft.** The MoE MTP head (785 BF16 tensors, with per-expert weights) is
37
  transplanted from a **sibling Qwen-MoE checkpoint**,
38
  [`Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4-Vision`](https://huggingface.co/Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4-Vision)
39
+ (a DeepSeek-V4-Pro-Thinking distill of Qwen3.6-35B-A3B). The head weights originate with
40
+ [`Qwen/Qwen3.6-35B-A3B`](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) (Apache-2.0) and reach this build unchanged
41
+ via [`pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4`](https://huggingface.co/pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4) —
42
+ every checkpoint in that chain carries them byte-identical. Because the trunk and the donor share the
43
  **same Qwen MoE architecture** — identical hidden size and expert layout — the head transfers cleanly
44
  and drafts well on the Ornith trunk (acceptance below).
45
  - **Speculative decoding** — vLLM native `mtp` method, `num_speculative_tokens=3`. **Lossless**: the
 
119
  script handles both a dense head and this MoE head.
120
 
121
  ## Credits
122
+ - **DeepReinforce** — [`Ornith-1.0-35B`](https://huggingface.co/ornith-ai/Ornith-1.0-35B), the base/trunk model (MIT). *(The `deepreinforce-ai/` id now redirects to `ornith-ai/`.)*
123
+ - **nerkyor** — [`Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill`](https://huggingface.co/nerkyor/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill), the DeepSeek-V4-Pro-Thinking **trunk distill**. *(Corrected 2026-08-17: this card previously credited the distill with training the MTP head. It did not — the head is Qwen's, carried through unchanged. Verified by hash: `mtp.fc.weight` [2048, 4096] and the other `mtp.*` tensors are byte-identical between `Qwen/Qwen3.6-35B-A3B` and the distill.)*
124
  - **DeepSeek** — DeepSeek-V4-Pro-Thinking, the model that distill learned from.
125
+ - **pahajokiconsulting** — [`Qwen3.6-35B-A3B-MXFP4`](https://huggingface.co/pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4) (Apache-2.0), the build the 785 `mtp.*` tensors were copied from verbatim: it quantizes the trunk but keeps `mtp.*` in BF16, with the experts relaid from stacked 3D into per-expert 2D.
126
+ - **Qwen / Alibaba** — the Qwen MoE (A3B) architecture shared by trunk and head, **and the trained MTP head weights themselves** ([`Qwen/Qwen3.6-35B-A3B`](https://huggingface.co/Qwen/Qwen3.6-35B-A3B), Apache-2.0). The grafted head is Qwen's own trained block, not a third party's.
127
  - **vLLM** and **compressed-tensors** — serving stack and quantization format.
128
  - **Rob Smith** (`tcclaviger`) — the [RDNA4 vLLM base image](https://hub.docker.com/r/tcclaviger/vllm-rocm-mxfp4-nvfp4) that made gfx1201 serving possible. The [`capicua25x/vllm-rocm-rdna4`](https://hub.docker.com/r/capicua25x/vllm-rocm-rdna4) image this model runs on is a forward-port of his work — without it, none of this runs.
129