--- license: apache-2.0 base_model: - Qwen/Qwen3.6-35B-A3B-FP8 pipeline_tag: image-text-to-text tags: - radiance - rocm - amd - rdna4 - moe - fp8 - speculative-decoding - mtp - vision --- # Qwen3.6-35B-A3B FP8 — radiance container [Qwen/Qwen3.6-35B-A3B-FP8](https://huggingface.co/Qwen/Qwen3.6-35B-A3B-FP8) as a single `.rad` container for the **radiance** inference engine (AMD RDNA4, ROCm), with its MTP head for speculative decoding and its vision tower. | | | |---|---| | File | `qwen3.6-35b-a3b-fp8.rad` — 34.63 GiB | | Weights | the published FP8 checkpoint as it ships: every linear, the 256 routed experts included, E4M3 with a scale per 128×128 block; the lm_head made block FP8 by the recipe | | Speculator | the model's MTP head, which drafts with a 2-bit copy of the lm_head | | Vision | the 27-block vision tower, bf16: images and video in chat requests | | Context | 262,144 tokens trained; 200K tested | ## Serve The whole model fits two 32 GB cards. ```sh radiance --model qwen3.6-35b-a3b-fp8.rad --tp 2 --max-model-len 200000 --kv-cache-dtype fp8 \ --max-num-seqs 8 --host 0.0.0.0 --port 8000 ``` The server speaks the OpenAI API (`/v1/chat/completions`, `/v1/completions`), with tool calls and structured output, and `image_url` / `video_url` parts in chat messages. The MTP depth is chosen automatically (`--num-speculative-tokens N` states one, `0` turns speculation off). Built and tested on 2× Radeon AI PRO R9700 (gfx1201). ## How this file was made ```sh rad-convert Qwen/Qwen3.6-35B-A3B-FP8 --recipe q36-35b-a3b-fp8.recipe -o qwen3.6-35b-a3b-fp8.rad ``` The recipe (`q36-35b-a3b-fp8.recipe` in this repository); everything it does not name is the checkpoint's own: ``` output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 mtp.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16 ``` ## License Apache 2.0, as the base model; see `LICENSE`.