---
license: apache-2.0
base_model: nerkyor/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill
base_model_relation: quantized
language:
- en
- zh
pipeline_tag: text-generation
tags:
- mxfp4
- quantized
- compressed-tensors
- rocm
- rdna4
- moe
- qwen3
- distillation
- deepseek-v4
- reasoning
- thinking
- agentic
---
# Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill — MXFP4 (RDNA4 / R9700)
MXFP4 quantization of [`nerkyor/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill`](https://huggingface.co/nerkyor/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill)
— a DeepSeek-V4-Pro thinking-style LoRA distill of Qwen3.6-35B-A3B — built to run on **AMD
RDNA4** (Radeon AI PRO R9700 / RX 9070 XT) under the
[`tcclaviger/vllm-rocm-mxfp4-nvfp4`](https://hub.docker.com/r/tcclaviger/vllm-rocm-mxfp4-nvfp4)
vLLM container, **with multi-token-prediction (MTP) speculative decoding grafted back in.**
As far as I know this is the **first RDNA4-loadable MXFP4 of this distill.** The distill's own
NVFP4 build targets SGLang/Blackwell and crashes this container; this build loads cleanly and
benches **at parity with the production base** ([`pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4`](https://huggingface.co/pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4)).
- **69.3 GB BF16 → 21 GB MXFP4** (~30%)
- Recipe, tooling & full build log: **https://github.com/Capicua25x/qwen3.6-mxfp4-rdna4**
## Serving (vLLM, 2× R9700, TP2)
```bash
vllm serve Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4 \
--tensor-parallel-size 2 --gpu-memory-utilization 0.92 --max-model-len 262144 \
--language-model-only \
--enable-prefix-caching --max-num-seqs 64 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3 \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
```
⚠️ **`--language-model-only` is REQUIRED.** This is a text-only checkpoint, but its `config.json`
is wrapped in the multimodal `Qwen3_5MoeConfig` shape so the container's
`Qwen3_5MoeForConditionalGeneration` loader accepts it; the flag makes the (weightless) vision
tower a no-init stub. Without it, loading fails.
### Thinking mode is optional (verified both ways)
It's a *thinking* distill, but **thinking is not required** — it runs cleanly with thinking
**on or off**. This isn't obvious from the model; we tested to confirm it:
- **Off** — pass `chat_template_kwargs={"enable_thinking": false}` (the chat template injects a
pre-closed `` so the model answers directly). Validated on a 137-test
SQL/analytics regression suite — classification + SQL-gen + interpretation all thinking-off:
**136/137, 0 failures**.
- **On** (default; `temperature=0.6 / top_p=0.95`) — the DeepSeek-V4-Pro reasoning style.
Validated on a 27-scenario agentic tool-calling eval: **26/27**, ~4× faster per turn-chain
than the base model.
So one deployment serves both: **thinking-off** for deterministic / high-throughput paths,
**thinking-on** for agentic / reasoning paths. (Note: in thinking-on it reasons *inline in
`content`* — the qwen3 parser returns empty `reasoning_content`.)
## Performance (2× R9700, MTP-3)
Both columns measured on the same hardware/bench (2× R9700, TP2, MTP-3 each):
| Metric | This build | Base (pahajoki) |
|---|---|---|
| Single-stream, short prompt | **~107 tok/s** | ~101 |
| Single-stream, 6k prompt | ~82 tok/s | ~85 |
| Aggregate @128, short prompt | **~1875 tok/s** | ~1683 |
| Concurrency ceiling, short prompt | **~128** | ~128 |
| MTP draft acceptance (MTP-3, measured) | ~56% (grafted) | ~64% (native) |
Effectively **at parity**: the distill edges the base on short-prompt single-stream (107 vs 101)
and high-concurrency aggregate (1875 vs 1683 @128); 6k single-stream is a wash (82 vs 85). Where
it clearly wins is **agentic** — in a 27-scenario thinking-on tool-calling eval it ran **~4×
faster per turn-chain than the base** at equal task success (its DS-V4-Pro distillation makes it
*decisive* — fewer tokens to a decision).
## Reproduce — step by step
Full scripts in the [companion repo](https://github.com/Capicua25x/qwen3.6-mxfp4-rdna4). The
recipe is **generic for the Qwen `qwen3_5_moe` family** on RDNA4.
**0. Tooling.** Clone [olka/qstream](https://github.com/olka/qstream) `master` (includes the
merged Qwen3.6 fix) into a venv: `pip install -e qstream`.
**1. Quantize (CPU, RTN — no GPU).** The qstream defaults *are* the recipe:
```bash
qstream-quantize --model_dir --output_dir --workers 8 --format ct
```
Keeps BF16: `*self_attn* *.mlp.gate. *shared_expert* *lm_head* *embed_tokens* *visual* *mtp*`.
Quantizes to MXFP4 (compressed-tensors `mxfp4-pack-quantized`, per-expert, group-32 symmetric,
MSE-optimal scale): the `linear_attn` projections + all 256 routed experts.
**2. Wrap the config (text-only checkpoints only).** The RDNA4 container only registers the
multimodal loader, so rewrap `config.json` into `Qwen3_5MoeConfig` shape — nest the text fields
under `text_config`, graft a `vision_config` + vision/image token ids from a multimodal build
(e.g. pahajoki's base) — then always serve with `--language-model-only`. (`scripts/wrap-config.py`)
**3. Graft MTP (optional; restores speculative decoding).** This distill shipped
`mtp_num_hidden_layers: 0`. The base Qwen3.6 MTP block is dim-identical, so: copy the base's 785
BF16 `mtp.*` tensors into a new shard, set `mtp_num_hidden_layers: 1`, **and add the `mtp.*`
modules to `quantization_config.ignore`** (else vLLM loads them as quantized →
`fc.weight not found in params_dict`). (`scripts/graft-mtp.py`)
## Caveats
- **HF shows an "8-bit precision" badge — ignore it, this model is 4-bit.** MXFP4 packs two
4-bit FP4 values into each `uint8` byte (`weight_packed`), so HF reads the `uint8` *storage*
dtype and mislabels it. Source of truth: `config.json` → `num_bits: 4`,
`format: mxfp4-pack-quantized`. Every MXFP4 compressed-tensors model shows this (incl. the base).
- Built/tested only on **gfx1201 (RDNA4)** with `tcclaviger/vllm-rocm-mxfp4-nvfp4`.
- Reasons **inline in `content`** — the qwen3 reasoning-parser returns empty `reasoning_content`.
- `--language-model-only` required (see Serving).
- Grafted MTP acceptance ~56% vs ~64% native (both measured at MTP-3) — only ~8pp behind; a
native MTP retrain would close it. MTP is lossless either way.
## Credits & acknowledgments
This build stands entirely on others' work — full credit to:
- **The Qwen team (Alibaba)** — [Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B),
the base model and architecture (Apache-2.0).
- **nerkyor / "Lynn"** — the [DSV4Pro-Thinking distill](https://huggingface.co/nerkyor/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill)
this quantizes; the actual reasoning capability is theirs.
- **DeepSeek-AI** — DeepSeek-V4-Pro, the distillation teacher.
- **olka** — [qstream](https://github.com/olka/qstream), the streamable MXFP4 quantizer.
- **kallepahajoki** — the qstream Qwen3.6-family fix (PR #1: shared-expert + fused-passthrough)
and the reference MXFP4 recipe published as
[`pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4`](https://huggingface.co/pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4),
including the BF16/MXFP4 exclude set and the MTP layout this graft borrows.
- **tcclaviger** — the [RDNA4/gfx12 MXFP4 vLLM container & custom MoE kernel](https://hub.docker.com/r/tcclaviger/vllm-rocm-mxfp4-nvfp4)
that makes any of this run on consumer/pro AMD cards.
- Distillation technique lineage (from the source distill): ReAct (Yao et al., 2022,
arXiv:2210.03629), Self-Instruct/Baize, AgentTuning, ToolBench, DeepSeek-R1 distillation.
Quantization + RDNA4 packaging (config-wrap + MTP graft) by **Capicua25x**. Apache-2.0.