Conversion recipe
Browse files- README.md +3 -1
- q36-35b-a3b-fp8.recipe +15 -0
README.md
CHANGED
|
@@ -67,7 +67,9 @@ docker run --rm -v ~/models:/models --entrypoint rad-convert radiance \
|
|
| 67 |
-o /models/qwen3.6-35b-a3b-fp8.rad
|
| 68 |
```
|
| 69 |
|
| 70 |
-
The recipe
|
|
|
|
|
|
|
| 71 |
|
| 72 |
```
|
| 73 |
output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
|
|
|
|
| 67 |
-o /models/qwen3.6-35b-a3b-fp8.rad
|
| 68 |
```
|
| 69 |
|
| 70 |
+
The recipe file both commands name is in this repository as `q36-35b-a3b-fp8.recipe`, with its
|
| 71 |
+
comments. Its rules, as the container records them (`rad-info --recipe qwen3.6-35b-a3b-fp8.rad`);
|
| 72 |
+
everything they do not name is the checkpoint's own:
|
| 73 |
|
| 74 |
```
|
| 75 |
output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
|
q36-35b-a3b-fp8.recipe
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# q36-35b-a3b-fp8.recipe -- Qwen3.6-35B-A3B-FP8 with its MTP head, as the q36-35b-a3b-fp8.rad
|
| 2 |
+
# container serves it.
|
| 3 |
+
#
|
| 4 |
+
# rad-convert <Qwen3.6-35B-A3B-FP8 checkpoint> --recipe data/recipes/q36-35b-a3b-fp8.recipe \
|
| 5 |
+
# -o q36-35b-a3b-fp8.rad
|
| 6 |
+
#
|
| 7 |
+
# The linears -- attention, delta net, shared and routed experts, the MTP layer's -- are
|
| 8 |
+
# block-scaled fp8 in the checkpoint and kept as they are. What the recipe adds:
|
| 9 |
+
#
|
| 10 |
+
# - the lm_head, block fp8: the widest GEMM in the model at half the bytes
|
| 11 |
+
# - the MTP head's 2-bit copy of the lm_head (u2 codes, an f16 scale and a u8 zero a group of
|
| 12 |
+
# 128), which the draft rounds read in place of the full head
|
| 13 |
+
|
| 14 |
+
output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
|
| 15 |
+
mtp.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16
|