Qwen3.5-9B-DFlash / README.md
jianchen0311's picture
Update README.md
9b697fb verified
|
Raw
History Blame
2.13 kB
metadata
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags:
  - speculative-decoding
  - diffusion
  - efficiency
  - flash-decoding
  - qwen
  - diffusion-language-model

Qwen3.5-9B-DFlash

Paper | GitHub | Blog

This model is still under training.

DFlash is a novel speculative decoding method that utilizes a lightweight block diffusion model for drafting. It enables efficient, high-quality parallel drafting that pushes the limits of inference speed.

This model is the drafter component. It must be used in conjunction with the target model Qwen/Qwen3.5-9B. It was trained with a context length of 4096 tokens.

DFlash Architecture

๐Ÿš€ Quick Start

SGLang

Installation

uv pip install "git+https://github.com/sgl-project/sglang.git@refs/pull/16818/head#subdirectory=python"

Inference

python -m sglang.launch_server \
    --model-path Qwen/Qwen3.5-9B \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path z-lab/Qwen3.5-9B-DFlash \
    --speculative-num-draft-tokens 16 \
    --tp-size 1 \
    --dtype bfloat16 \
    --attention-backend fa3 \
    --mem-fraction-static 0.75 \
    --trust-remote-code \
    --mamba-scheduler-strategy extra_buffer \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder

Note: For long-context or agentic usage (such as OpenClaw or Claude Code), consider adding --speculative-dflash-draft-window-size WINDOW_SIZE to enable sliding-window attention for the draft model. Because the draft model is only trained on 4K context, this often improves performance on very long context (50K+ tokens).

Early Results

  • Thinking: enabled
  • Max new tokens: 4096
  • Block size: 16
    Dataset Accept Length
    GSM8K 6.709
    Math500 7.388
    HumanEval 7.888
    MBPP 6.617
    MT-Bench 5.506
    Alpaca 5.079