Qwen3.5-9B-DFlash / README.md
jianchen0311's picture
Update README.md
9b697fb verified
|
Raw
History Blame
2.13 kB
---
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags:
- speculative-decoding
- diffusion
- efficiency
- flash-decoding
- qwen
- diffusion-language-model
---
# Qwen3.5-9B-DFlash
[**Paper**](https://arxiv.org/abs/2602.06036) | [**GitHub**](https://github.com/z-lab/dflash) | [**Blog**](https://z-lab.ai/projects/dflash/)
**This model is still under training.**
**DFlash** is a novel speculative decoding method that utilizes a lightweight **block diffusion** model for drafting. It enables efficient, high-quality parallel drafting that pushes the limits of inference speed.
This model is the **drafter** component. It must be used in conjunction with the target model `Qwen/Qwen3.5-9B`. It was trained with a context length of 4096 tokens.
<div align="center">
<img src="assets/dflash_system.png" alt="DFlash Architecture" width="100%">
</div>
## 🚀 Quick Start
### SGLang
#### Installation
```bash
uv pip install "git+https://github.com/sgl-project/sglang.git@refs/pull/16818/head#subdirectory=python"
```
#### Inference
```bash
python -m sglang.launch_server \
--model-path Qwen/Qwen3.5-9B \
--speculative-algorithm DFLASH \
--speculative-draft-model-path z-lab/Qwen3.5-9B-DFlash \
--speculative-num-draft-tokens 16 \
--tp-size 1 \
--dtype bfloat16 \
--attention-backend fa3 \
--mem-fraction-static 0.75 \
--trust-remote-code \
--mamba-scheduler-strategy extra_buffer \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder
```
> **Note:** For long-context or agentic usage (such as OpenClaw or Claude Code), consider adding `--speculative-dflash-draft-window-size WINDOW_SIZE` to enable sliding-window attention for the draft model. Because the draft model is only trained on 4K context, this often improves performance on very long context (50K+ tokens).
#### Early Results
- Thinking: enabled
- Max new tokens: 4096
- Block size: 16
| Dataset | Accept Length |
|-----------|---------------|
| GSM8K | 6.709 |
| Math500 | 7.388 |
| HumanEval | 7.888 |
| MBPP | 6.617 |
| MT-Bench | 5.506 |
| Alpaca | 5.079 |