Model card: the exact command that made the file
Browse files
README.md
CHANGED
|
@@ -30,7 +30,7 @@ with the model's MTP head for speculative decoding and its vision tower.
|
|
| 30 |
| File | `qwen3.8-next-flash-fp8-iq4r-moe.rad` — 113.6 GiB |
|
| 31 |
| Routed experts | 4-bit codes into a 16-level non-uniform codebook after a 128-point Walsh–Hadamard rotation, one E4M3 scale per 64 weights; codes chosen by GPTQ (10M-token calibration, per-expert down-projection Hessians) and refined by three sweeps of coordinate descent |
|
| 32 |
| Protected experts | ten experts that carry most of their layer's down-projection energy, kept bf16 |
|
| 33 |
-
| Trunk | attention and shared-expert linears and the lm_head int8 (W8A8), a scale per 128 columns searched for least error; hyper-connection mixing matrices E4M3 |
|
| 34 |
| Speculator | the model's MTP head (depth 3 by default) |
|
| 35 |
| Vision | the vision tower, bf16: images and video in chat requests |
|
| 36 |
| Context | 262,144 tokens trained; served at 200K |
|
|
@@ -65,6 +65,63 @@ tokens per step, and ~6,400 tokens/s prefill on a 26K-token prompt.
|
|
| 65 |
The server speaks the OpenAI API (`/v1/chat/completions`, `/v1/completions`), with tool calls,
|
| 66 |
structured output, and `image_url` / video parts in chat messages.
|
| 67 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 68 |
## License
|
| 69 |
|
| 70 |
The Qwen community license of the base model; see `LICENSE`.
|
|
|
|
| 30 |
| File | `qwen3.8-next-flash-fp8-iq4r-moe.rad` — 113.6 GiB |
|
| 31 |
| Routed experts | 4-bit codes into a 16-level non-uniform codebook after a 128-point Walsh–Hadamard rotation, one E4M3 scale per 64 weights; codes chosen by GPTQ (10M-token calibration, per-expert down-projection Hessians) and refined by three sweeps of coordinate descent |
|
| 32 |
| Protected experts | ten experts that carry most of their layer's down-projection energy, kept bf16 |
|
| 33 |
+
| Trunk | attention, delta-net and shared-expert linears and the lm_head int8 (W8A8), a scale per 128 columns searched for least error; hyper-connection mixing matrices E4M3 |
|
| 34 |
| Speculator | the model's MTP head (depth 3 by default) |
|
| 35 |
| Vision | the vision tower, bf16: images and video in chat requests |
|
| 36 |
| Context | 262,144 tokens trained; served at 200K |
|
|
|
|
| 65 |
The server speaks the OpenAI API (`/v1/chat/completions`, `/v1/completions`), with tool calls,
|
| 66 |
structured output, and `image_url` / video parts in chat messages.
|
| 67 |
|
| 68 |
+
## How this file was made
|
| 69 |
+
|
| 70 |
+
The exact command, run on 2026-10-03 from a radiance build directory of the commit series merged
|
| 71 |
+
to radiance's main branch as 2531f14 (`radiance_home` is that build's plugin directory):
|
| 72 |
+
|
| 73 |
+
```sh
|
| 74 |
+
CALIB=~/models/calib/w4nl-calib OMP_NUM_THREADS=12 nice -n 10 ./bin/rad-convert \
|
| 75 |
+
~/models/Qwen/Qwen3.8-Flash-Next --recipe data/recipes/qwen4exp-w4nl64-i8-hc8m.recipe \
|
| 76 |
+
--home radiance_home --tokenizer ~/models/Qwen/Qwen3.8-Flash-Next/tokenizer.json \
|
| 77 |
+
--reuse ~/models/rad/w4nl-i8lm.rad -o ~/models/rad/w4nl64-i8lm.rad -v
|
| 78 |
+
```
|
| 79 |
+
|
| 80 |
+
- `--reuse` names an earlier container converted from the same checkpoint by an intermediate
|
| 81 |
+
recipe. Every weight it already held under the same rule -- the int8 trunk and lm_head, the E4M3
|
| 82 |
+
connection and MTP matrices, the ten bf16 experts -- was copied from it; the routed experts were
|
| 83 |
+
quantised by this run. Without `--reuse` the same command quantises those weights from the
|
| 84 |
+
checkpoint by the same rules.
|
| 85 |
+
- `CALIB` is the calibration directory the GPTQ rule reads: each layer's routed-expert input Gram
|
| 86 |
+
from ten million tokens of the bf16 model, and each expert's own down-projection Gram from one
|
| 87 |
+
million, recorded with the engine's calibration mode on a mixed chat, tool-use and code corpus.
|
| 88 |
+
It is not published.
|
| 89 |
+
- The published file is that output under this repository's name, with the local directory
|
| 90 |
+
removed from the `created by`, recipe and calibration-path strings in its header; the weights
|
| 91 |
+
are the converter's bytes.
|
| 92 |
+
|
| 93 |
+
The recipe, as `rad-info --recipe qwen3.8-next-flash-fp8-iq4r-moe.rad` prints it; the first
|
| 94 |
+
rule that matches a weight decides it, and everything no rule names is the checkpoint's bf16:
|
| 95 |
+
|
| 96 |
+
```
|
| 97 |
+
blk.34.ffn_*_exps.407.weight cast dtype=bf16
|
| 98 |
+
blk.34.ffn_*_exps.496.weight cast dtype=bf16
|
| 99 |
+
blk.44.ffn_*_exps.292.weight cast dtype=bf16
|
| 100 |
+
blk.44.ffn_*_exps.350.weight cast dtype=bf16
|
| 101 |
+
blk.46.ffn_*_exps.290.weight cast dtype=bf16
|
| 102 |
+
blk.46.ffn_*_exps.392.weight cast dtype=bf16
|
| 103 |
+
blk.47.ffn_*_exps.122.weight cast dtype=bf16
|
| 104 |
+
blk.47.ffn_*_exps.143.weight cast dtype=bf16
|
| 105 |
+
blk.47.ffn_*_exps.399.weight cast dtype=bf16
|
| 106 |
+
blk.47.ffn_*_exps.445.weight cast dtype=bf16
|
| 107 |
+
blk.*.ffn_*_exps.*.weight gptq table=w4nl group=64 scale=fp8_e4m3 scale2=f32 block2=*x* scale2_value=0.0001220703125 transform=fwht128 rule=search cd=3 calib=calib/w4nl-calib
|
| 108 |
+
*_hc_down.weight rtn codes=fp8_e4m3 group=128 scale=f32
|
| 109 |
+
*_hc_up.weight rtn codes=fp8_e4m3 group=80 scale=f32
|
| 110 |
+
mtp.fc_hidden.weight rtn codes=fp8_e4m3 group=128 scale=f32
|
| 111 |
+
mtp.fc_embedding.weight rtn codes=fp8_e4m3 group=128 scale=f32
|
| 112 |
+
blk.*.attn_qg.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 113 |
+
blk.*.attn_k.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 114 |
+
blk.*.attn_v.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 115 |
+
blk.*.attn_output.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 116 |
+
blk.*.ssm_inz.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 117 |
+
blk.*.ssm_out.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 118 |
+
blk.*.ffn_gate_up_shexp.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 119 |
+
blk.*.ffn_down_shexp.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 120 |
+
output.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 121 |
+
blk.*.ple_ngram.weight rtn codes=fp8_e4m3 block=*x* scale=bf16 rule=fixed scale_value=1.99317932128906e-4
|
| 122 |
+
mtp.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16
|
| 123 |
+
```
|
| 124 |
+
|
| 125 |
## License
|
| 126 |
|
| 127 |
The Qwen community license of the base model; see `LICENSE`.
|