--- license: other license_name: qwen-community-1.0 license_link: LICENSE base_model: - Qwen/Qwen3.8-Flash-Next pipeline_tag: image-text-to-text tags: - radiance - rocm - amd - rdna4 - moe - 4-bit - int8 - gptq - speculative-decoding - mtp - vision --- # Qwen3.8-Flash-Next, 4-bit codebook experts + int8 trunk — radiance container [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) quantised from its bf16 checkpoint into a single `.rad` container for the **radiance** inference engine (AMD RDNA4, ROCm), with the model's MTP head for speculative decoding and its vision tower. | | | |---|---| | File | `qwen3.8-next-flash-fp8-iq4r-moe.rad` — 113.6 GiB | | Routed experts | 4-bit codes into a 16-level non-uniform codebook after a 128-point Walsh–Hadamard rotation, one E4M3 scale per 64 weights; codes chosen by GPTQ (10M-token calibration, per-expert down-projection Hessians) and refined by three sweeps of coordinate descent | | Protected experts | ten experts that carry most of their layer's down-projection energy, kept bf16 | | Trunk | attention, delta-net and shared-expert linears and the lm_head int8 (W8A8), a scale per 128 columns searched for least error; hyper-connection mixing matrices E4M3 | | Speculator | the model's MTP head (depth 3 by default) | | Vision | the vision tower, bf16: images and video in chat requests | | Context | 262,144 tokens trained; served at 200K | ## Quality Full-vocabulary KL divergence against the bf16 model, teacher-forced over ~75K positions of chat, tool-use and code transcripts: | | mean KL | median | p99 | p99.9 | top-1 agreement | |---|---|---|---|---|---| | all tokens | 0.0789 | 0.0037 | 1.65 | 5.33 | 92.85% | | text and assistant turns | 0.0239 | 0.0029 | 0.32 | 1.03 | 94.57% | | reference: a second bf16 implementation, all tokens | 0.0399 | 0.0013 | 0.82 | 3.43 | 95.13% | ## Serve The routed experts do not fit two 32 GB cards; the engine keeps the hottest in VRAM and streams the rest from a pinned host pool, so the host needs about 32 GB of free RAM. ```sh radiance --model qwen3.8-next-flash-fp8-iq4r-moe.rad --tp 2 --tp-wire wht6 \ --max-model-len 200000 --max-num-seqs 8 --kv-cache-dtype fp8 \ --placement expert_tiered --host-pool-mib 12288 --gpu-headroom-mib 96 \ --expert-vs-cache-ratio 0.82 --num-speculative-tokens 3 --max-num-batched-tokens 2048 \ --host 0.0.0.0 --port 8000 ``` On 2× Radeon AI PRO R9700 (gfx1201), one stream: about 15 ms per decode step at ~2.4 accepted tokens per step, and ~6,400 tokens/s prefill on a 26K-token prompt. The server speaks the OpenAI API (`/v1/chat/completions`, `/v1/completions`), with tool calls, structured output, and `image_url` / video parts in chat messages. ## How this file was made ```sh CALIB=calib/w4nl-calib rad-convert Qwen/Qwen3.8-Flash-Next \ --tokenizer Qwen/Qwen3.8-Flash-Next/tokenizer.json \ --recipe qwen4exp-w4nl64-i8-hc8m.recipe -o qwen3.8-next-flash-fp8-iq4r-moe.rad ``` The recipe (`qwen4exp-w4nl64-i8-hc8m.recipe` in this repository; `$CALIB` is its calibration data, which is not published). The first rule that matches a weight decides it, and everything no rule names is the checkpoint's bf16: ``` blk.34.ffn_*_exps.407.weight cast dtype=bf16 blk.34.ffn_*_exps.496.weight cast dtype=bf16 blk.44.ffn_*_exps.292.weight cast dtype=bf16 blk.44.ffn_*_exps.350.weight cast dtype=bf16 blk.46.ffn_*_exps.290.weight cast dtype=bf16 blk.46.ffn_*_exps.392.weight cast dtype=bf16 blk.47.ffn_*_exps.122.weight cast dtype=bf16 blk.47.ffn_*_exps.143.weight cast dtype=bf16 blk.47.ffn_*_exps.399.weight cast dtype=bf16 blk.47.ffn_*_exps.445.weight cast dtype=bf16 blk.*.ffn_*_exps.*.weight gptq table=w4nl group=64 scale=fp8_e4m3 scale2=f32 block2=*x* scale2_value=0.0001220703125 transform=fwht128 rule=search cd=3 calib=calib/w4nl-calib *_hc_down.weight rtn codes=fp8_e4m3 group=128 scale=f32 *_hc_up.weight rtn codes=fp8_e4m3 group=80 scale=f32 mtp.fc_hidden.weight rtn codes=fp8_e4m3 group=128 scale=f32 mtp.fc_embedding.weight rtn codes=fp8_e4m3 group=128 scale=f32 blk.*.attn_qg.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search blk.*.attn_k.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search blk.*.attn_v.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search blk.*.attn_output.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search blk.*.ssm_inz.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search blk.*.ssm_out.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search blk.*.ffn_gate_up_shexp.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search blk.*.ffn_down_shexp.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search output.weight rtn codes=i8 group=128 scale=bf16 clamp=sym rule=search blk.*.ple_ngram.weight rtn codes=fp8_e4m3 block=*x* scale=bf16 rule=fixed scale_value=1.99317932128906e-4 mtp.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16 ``` ## License The Qwen community license of the base model; see `LICENSE`.