StillDeadcode commited on
Commit
e5b35b6
·
verified ·
1 Parent(s): 7c37006

Conversion recipe

Browse files
Files changed (2) hide show
  1. README.md +3 -1
  2. q36-35b-a3b-fp8.recipe +15 -0
README.md CHANGED
@@ -67,7 +67,9 @@ docker run --rm -v ~/models:/models --entrypoint rad-convert radiance \
67
  -o /models/qwen3.6-35b-a3b-fp8.rad
68
  ```
69
 
70
- The recipe (`rad-info --recipe qwen3.6-35b-a3b-fp8.rad`); everything it does not name is the checkpoint's own:
 
 
71
 
72
  ```
73
  output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
 
67
  -o /models/qwen3.6-35b-a3b-fp8.rad
68
  ```
69
 
70
+ The recipe file both commands name is in this repository as `q36-35b-a3b-fp8.recipe`, with its
71
+ comments. Its rules, as the container records them (`rad-info --recipe qwen3.6-35b-a3b-fp8.rad`);
72
+ everything they do not name is the checkpoint's own:
73
 
74
  ```
75
  output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
q36-35b-a3b-fp8.recipe ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # q36-35b-a3b-fp8.recipe -- Qwen3.6-35B-A3B-FP8 with its MTP head, as the q36-35b-a3b-fp8.rad
2
+ # container serves it.
3
+ #
4
+ # rad-convert <Qwen3.6-35B-A3B-FP8 checkpoint> --recipe data/recipes/q36-35b-a3b-fp8.recipe \
5
+ # -o q36-35b-a3b-fp8.rad
6
+ #
7
+ # The linears -- attention, delta net, shared and routed experts, the MTP layer's -- are
8
+ # block-scaled fp8 in the checkpoint and kept as they are. What the recipe adds:
9
+ #
10
+ # - the lm_head, block fp8: the widest GEMM in the model at half the bytes
11
+ # - the MTP head's 2-bit copy of the lm_head (u2 codes, an f16 scale and a u8 zero a group of
12
+ # 128), which the draft rounds read in place of the full head
13
+
14
+ output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
15
+ mtp.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16