Conversion recipe
Browse files- README.md +4 -2
- qwen4exp-w4nl64-i8-hc8m.recipe +80 -0
README.md
CHANGED
|
@@ -90,8 +90,10 @@ CALIB=~/models/calib/w4nl-calib OMP_NUM_THREADS=12 nice -n 10 ./bin/rad-convert
|
|
| 90 |
removed from the `created by`, recipe and calibration-path strings in its header; the weights
|
| 91 |
are the converter's bytes.
|
| 92 |
|
| 93 |
-
The recipe
|
| 94 |
-
|
|
|
|
|
|
|
| 95 |
|
| 96 |
```
|
| 97 |
blk.34.ffn_*_exps.407.weight cast dtype=bf16
|
|
|
|
| 90 |
removed from the `created by`, recipe and calibration-path strings in its header; the weights
|
| 91 |
are the converter's bytes.
|
| 92 |
|
| 93 |
+
The recipe file the command names is in this repository as `qwen4exp-w4nl64-i8-hc8m.recipe`, with
|
| 94 |
+
its comments; its `$CALIB` is the calibration directory. Its rules, as
|
| 95 |
+
`rad-info --recipe qwen3.8-next-flash-fp8-iq4r-moe.rad` prints them; the first rule that matches a
|
| 96 |
+
weight decides it, and everything no rule names is the checkpoint's bf16:
|
| 97 |
|
| 98 |
```
|
| 99 |
blk.34.ffn_*_exps.407.weight cast dtype=bf16
|
qwen4exp-w4nl64-i8-hc8m.recipe
ADDED
|
@@ -0,0 +1,80 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# qwen4exp-w4nl64-i8-hc8m.recipe -- Qwen3.8-Flash-Next from its bf16 checkpoint as it is served
|
| 2 |
+
# (the w4nl64-i8lm.rad container): the routed experts as rotated u4 codes into the w4nl
|
| 3 |
+
# codebook with one E4M3 scale a (row, 64 of K) under a fixed f32 of 2^-13 (libr4d's w4nl64a8h),
|
| 4 |
+
# the codes chosen by GPTQ and three sweeps of coordinate descent; ten protected experts plain
|
| 5 |
+
# bf16; the trunk's block linears and the lm_head int8, served W8A8; the connections' mixing
|
| 6 |
+
# matrices and the MTP head's fc matrices E4M3; the n-gram table at one fixed E4M3 scale.
|
| 7 |
+
#
|
| 8 |
+
# Measured with the engine's KL mode against the bf16 model (docs/MOE-W4.md, "What is served"):
|
| 9 |
+
# mean KL 0.0789 and top-1 agreement 92.85% over the whole corpus, where a second bf16
|
| 10 |
+
# implementation scores 0.0399 / 95.13%; 0.0239 / 94.57% on text and assistant turns.
|
| 11 |
+
#
|
| 12 |
+
# CALIB=DIR rad-convert <Qwen3.8-Flash-Next checkpoint> \
|
| 13 |
+
# --recipe data/recipes/qwen4exp-w4nl64-i8-hc8m.recipe -o w4nl64-i8lm.rad
|
| 14 |
+
#
|
| 15 |
+
# DIR holds each layer's pooled Grams, recorded from ten million tokens of the bf16 model, and each
|
| 16 |
+
# expert's own down Gram, from one million tokens with the per-expert tap, side by side; libquant
|
| 17 |
+
# takes an expert's own wherever it holds at least 4096 rows and the layer's otherwise
|
| 18 |
+
# (docs/MOE-W4.md, "Calibration"). A variant of this recipe converted with --reuse w4nl64-i8lm.rad
|
| 19 |
+
# re-quantises only the weights whose rule changed.
|
| 20 |
+
#
|
| 21 |
+
# Everything not named here is served as the checkpoint holds it; the first rule that matches a
|
| 22 |
+
# weight decides it.
|
| 23 |
+
|
| 24 |
+
# The protected experts, from a 1M-token recording of the bf16 model on the calibration corpus:
|
| 25 |
+
# an expert whose share of its layer's down-input energy is at least 3% and at least three times its
|
| 26 |
+
# share of the routing, and on each rank's half of the layer an even count -- the four-bit GEMM takes
|
| 27 |
+
# the rest as two equal tables by parity -- by adding that half's next most energetic expert
|
| 28 |
+
# (407/496, 350/292, 290/392 and 445,143/399,122 below). Layer 47's expert 445 alone carries 64% of
|
| 29 |
+
# its layer's down-input energy from 0.4% of the routing.
|
| 30 |
+
blk.34.ffn_*_exps.407.weight cast dtype=bf16
|
| 31 |
+
blk.34.ffn_*_exps.496.weight cast dtype=bf16
|
| 32 |
+
blk.44.ffn_*_exps.292.weight cast dtype=bf16
|
| 33 |
+
blk.44.ffn_*_exps.350.weight cast dtype=bf16
|
| 34 |
+
blk.46.ffn_*_exps.290.weight cast dtype=bf16
|
| 35 |
+
blk.46.ffn_*_exps.392.weight cast dtype=bf16
|
| 36 |
+
blk.47.ffn_*_exps.122.weight cast dtype=bf16
|
| 37 |
+
blk.47.ffn_*_exps.143.weight cast dtype=bf16
|
| 38 |
+
blk.47.ffn_*_exps.399.weight cast dtype=bf16
|
| 39 |
+
blk.47.ffn_*_exps.445.weight cast dtype=bf16
|
| 40 |
+
|
| 41 |
+
# Every other routed expert, trunk and MTP layer alike: u4 codes into the w4nl table, rotated by
|
| 42 |
+
# the 128-wide Hadamard, an E4M3 scale a group of 64 under the fixed 2^-13 searched for least error
|
| 43 |
+
# as stored, the codes chosen by GPTQ against $CALIB and then swept three times by coordinate
|
| 44 |
+
# descent. The experts' scales sit between 2^-15 and 2^-10.6, inside E4M3's normal range under
|
| 45 |
+
# 2^-13 with four binades either side; one past its top would be held there and said on stderr.
|
| 46 |
+
# The MTP layer has no recording and is rounded to nearest, which libquant says on stderr.
|
| 47 |
+
blk.*.ffn_*_exps.*.weight gptq table=w4nl group=64 scale=fp8_e4m3 scale2=f32 block2=*x* scale2_value=0.0001220703125 transform=fwht128 rule=search cd=3 calib=$CALIB
|
| 48 |
+
|
| 49 |
+
# Every connection's two mixing matrices -- the 96 trunk connections, the final mixer and the MTP
|
| 50 |
+
# head's three -- as E4M3 rows with an f32 scale a group of each row. A 128x128 block cannot cover
|
| 51 |
+
# them: lowrank is 320, 2.5 blocks. Down is [320, 10240] and takes a scale every 128 columns; up is
|
| 52 |
+
# [10240, 320] and takes four scales a row, 80 columns each.
|
| 53 |
+
*_hc_down.weight rtn codes=fp8_e4m3 group=128 scale=f32
|
| 54 |
+
*_hc_up.weight rtn codes=fp8_e4m3 group=80 scale=f32
|
| 55 |
+
|
| 56 |
+
# The MTP head's fc_hidden and fc_embedding, the same rows. They only ever change the drafts.
|
| 57 |
+
mtp.fc_hidden.weight rtn codes=fp8_e4m3 group=128 scale=f32
|
| 58 |
+
mtp.fc_embedding.weight rtn codes=fp8_e4m3 group=128 scale=f32
|
| 59 |
+
|
| 60 |
+
# The trunk's block linears as int8, a scale per 128 columns of a row searched for least error.
|
| 61 |
+
# The QSA indexer's projection is not one of them: its top-k is a hard choice between blocks, and
|
| 62 |
+
# keeping it as the checkpoint holds it costs 20 MB.
|
| 63 |
+
blk.*.attn_qg.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 64 |
+
blk.*.attn_k.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 65 |
+
blk.*.attn_v.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 66 |
+
blk.*.attn_output.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 67 |
+
blk.*.ssm_inz.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 68 |
+
blk.*.ssm_out.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 69 |
+
blk.*.ffn_gate_up_shexp.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 70 |
+
blk.*.ffn_down_shexp.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 71 |
+
|
| 72 |
+
# The lm_head, int8 with a bf16 scale a row per 128 columns (logits_gemm_i8): at E4M3 it costs
|
| 73 |
+
# ~0.6 points of top-1 on its own.
|
| 74 |
+
output.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search
|
| 75 |
+
|
| 76 |
+
# The n-gram table at one fixed E4M3 scale; qwen4exp-w4.recipe says where the value comes from.
|
| 77 |
+
blk.*.ple_ngram.weight rtn codes=fp8_e4m3 block=*x* scale=bf16 rule=fixed scale_value=1.99317932128906e-4
|
| 78 |
+
|
| 79 |
+
# The MTP head's 2-bit draft head.
|
| 80 |
+
mtp.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16
|