Capicua25x's picture
Add files using upload-large-folder tool
4cdd29e verified
|
Raw History Blame
5.02 kB
---
license: apache-2.0
base_model: nerkyor/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill
base_model_relation: quantized
language:
- en
- zh
pipeline_tag: image-text-to-text
tags:
- mxfp4
- quantized
- compressed-tensors
- rocm
- rdna4
- moe
- qwen3
- distillation
- deepseek-v4
- reasoning
- thinking
- agentic
- multimodal
- vision
---
# Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill — MXFP4 + Vision (RDNA4 / R9700)
The **vision-enabled sibling** of
[`Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4`](https://huggingface.co/Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4):
the same MXFP4 quant of nerkyor's DeepSeek-V4-Pro thinking distill, **with the Qwen3.6 base
vision tower grafted back in** — so it's **text + reasoning + vision + MTP**, all on the
RDNA4 [`tcclaviger/vllm-rocm-mxfp4-nvfp4`](https://hub.docker.com/r/tcclaviger/vllm-rocm-mxfp4-nvfp4) container.
The distill shipped text-only (nerkyor dropped the vision tower on export). This build restores it.
Recipe + tooling: **https://github.com/Capicua25x/qwen3.6-mxfp4-rdna4**
## How image support was added (the vision graft)
nerkyor's distill is `Qwen3_5MoeForCausalLM` (text-only) — no vision tower. But the base
`Qwen/Qwen3.6-35B-A3B` is multimodal, and its vision tower is **architecturally independent** of
the language model: the vision encoder → merger projects image features to `out_hidden_size: 2048`,
which is exactly the LM's `hidden_size`. So the tower bolts straight on — the same trick used to
graft MTP back:
1. **Copy the base's 333 `model.visual.*` tensors** (BF16, ~0.9 GB) from a known-good multimodal
build into a new shard.
2. **Add `model.visual.*` to `quantization_config.ignore`** so vLLM loads them as BF16 (the vision
tower is never quantized), and keep the multimodal `Qwen3_5MoeConfig` (with `vision_config`).
3. **Serve without `--language-model-only`** (that flag had been stubbing the tower out), with the
base's `chat_template` + `preprocessor_config` for image-token handling.
**Why it works:** the DeepSeek-V4-Pro distill is a *light* LoRA (text-only reasoning traces); it
shifted the LM but didn't break its ability to interpret the base vision tower's embedding space.
**Verified empirically** — fed a generated image (text "VISION 7", a red circle, a blue rectangle),
the model read the **text, shapes, colors, and spatial positions** correctly, while still reasoning.
(Vision is *not* lossless like MTP — a heavier-finetuned LM might fail to ground the embeddings — so
this is a measured result, not a guarantee for arbitrary distills.)
## Serving (vLLM, 2× R9700, TP2)
```bash
vllm serve Capicua25x/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill-MXFP4-Vision \
--tensor-parallel-size 2 --gpu-memory-utilization 0.92 --max-model-len 262144 \
--enable-prefix-caching --max-num-seqs 64 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3 \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
```
**Do NOT pass `--language-model-only`** here (that disables vision). Send images via the standard
OpenAI `image_url` content blocks. Thinking-on: `temperature=0.6 / top_p=0.95`.
## Validation — identical to the text-only build, plus vision
| Gate | This (vision) build | Text-only sibling |
|---|---|---|
| SQL regression | **136/137, 0 FAIL** | 136/137, 0 FAIL |
| Agent eval (tool-calling, thinking-on) | **27/27, 0 FAIL** | 26/27 |
| Single-stream / ceiling (short) | **108.9 tok/s / ~128** | 107 / ~128 |
| MTP draft acceptance (MTP-3) | **~57%** | ~56% |
| Image understanding | ✅ | ❌ (none) |
| Cost of vision | **+0.9 GB VRAM, 0 text-perf** | — |
The vision tower only fires on image input, so text/agentic throughput is unchanged.
## Notes
- The "8-bit precision" badge is a HuggingFace artifact of MXFP4's `uint8` packing — this model
is genuinely **4-bit** (`config.json` → `num_bits: 4`, `mxfp4-pack-quantized`).
- Built/tested only on **gfx1201 (RDNA4, R9700)** with `tcclaviger/vllm-rocm-mxfp4-nvfp4`.
## Credits
Same chain as the text-only build, plus the vision tower from Qwen's base:
- **Qwen team (Alibaba)** — [Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) base + the vision tower (Apache-2.0).
- **nerkyor / "Lynn"** — the [DSV4Pro-Thinking distill](https://huggingface.co/nerkyor/Qwen3.6-35B-A3B-DSV4Pro-Thinking-Distill) (the reasoning).
- **DeepSeek-AI** — DeepSeek-V4-Pro, the distillation teacher.
- **olka** — [qstream](https://github.com/olka/qstream), the MXFP4 quantizer.
- **kallepahajoki** — the qstream Qwen3.6 fix + reference MXFP4 recipe ([`pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4`](https://huggingface.co/pahajokiconsulting/Qwen3.6-35B-A3B-MXFP4)), the source of the BF16 vision-tower weights grafted here.
- **tcclaviger** — the [RDNA4/gfx12 MXFP4 vLLM container & kernel](https://hub.docker.com/r/tcclaviger/vllm-rocm-mxfp4-nvfp4).
Quantization + RDNA4 packaging (config-wrap, MTP graft, vision graft) by **Capicua25x**. Apache-2.0.