Anbeeld commited on
Commit
bed9ec3
·
verified ·
1 Parent(s): a535937

Restore README image asset

Browse files
Files changed (1) hide show
  1. README.md +124 -124
README.md CHANGED
@@ -1,125 +1,125 @@
1
- ---
2
- base_model: z-lab/Qwen3.6-27B-DFlash
3
- tags:
4
- - transformers
5
- - safetensors
6
- - qwen3
7
- - feature-extraction
8
- - dflash
9
- - speculative-decoding
10
- - diffusion
11
- - efficiency
12
- - flash-decoding
13
- - qwen
14
- - diffusion-language-model
15
- - text-generation
16
- - custom_code
17
- - arxiv:2602.06036
18
- - license:mit
19
- - text-generation-inference
20
- - endpoints_compatible
21
- - region:us
22
- ---
23
-
24
- # Qwen 3.6 27B DFlash GGUF
25
-
26
- GGUF quantizations of [**z-lab DFlash draft model**](https://huggingface.co/z-lab/Qwen3.6-27B-DFlash) for [**Qwen 3.6 27B**](https://huggingface.co/Qwen/Qwen3.6-27B).
27
-
28
- Use with [BeeLlama.cpp](https://github.com/Anbeeld/beellama.cpp), a llama.cpp fork with advanced quantization features.
29
-
30
- ---
31
-
32
-
33
- # Qwen3.6-27B-DFlash
34
- [**Paper**](https://arxiv.org/abs/2602.06036) | [**GitHub**](https://github.com/z-lab/dflash) | [**Blog**](https://z-lab.ai/projects/dflash/)
35
-
36
- **This model is still under training, and inference engine support may not be fully available yet due to architectural changes, including causal SWA layers.**
37
-
38
- **DFlash** is a novel speculative decoding method that utilizes a lightweight **block diffusion** model for drafting. It enables efficient, high-quality parallel drafting that pushes the limits of inference speed.
39
-
40
- This model is the **drafter** component. It must be used in conjunction with the target model `Qwen/Qwen3.6-27B`.
41
-
42
- <div align="center">
43
- <img src="https://huggingface.co/z-lab/Qwen3.6-27B-DFlash/resolve/main/assets/dflash_system.png" alt="DFlash Architecture" width="100%">
44
- </div>
45
-
46
- ## Quick Start
47
-
48
- ### Installation
49
-
50
- vLLM (We temporarily modify the installation through this PR to support interleaved SWA and ensure correct handling of target hidden states for optimal performance):
51
- ```bash
52
- uv pip install vllm
53
- uv pip install -U --torch-backend=auto "vllm @ git+https://github.com/vllm-project/vllm.git@refs/pull/40898/head"
54
- ```
55
-
56
- SGLang:
57
- ```bash
58
- uv pip install "git+https://github.com/sgl-project/sglang.git@refs/pull/23000/head#subdirectory=python"
59
- ```
60
-
61
- ### Launch Server
62
-
63
- vLLM:
64
- ```bash
65
- vllm serve Qwen/Qwen3.6-27B \
66
- --speculative-config '{"method": "dflash", "model": "z-lab/Qwen3.6-27B-DFlash", "num_speculative_tokens": 15}' \
67
- --attention-backend flash_attn \
68
- --max-num-batched-tokens 32768
69
- ```
70
-
71
- SGLang:
72
- ```bash
73
- # Optional: enable schedule overlapping (experimental, may not be stable)
74
- # export SGLANG_ENABLE_SPEC_V2=1
75
- # export SGLANG_ENABLE_DFLASH_SPEC_V2=1
76
- # export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1
77
-
78
- python -m sglang.launch_server \
79
- --model-path Qwen/Qwen3.6-27B \
80
- --speculative-algorithm DFLASH \
81
- --speculative-draft-model-path z-lab/Qwen3.6-27B-DFlash \
82
- --speculative-num-draft-tokens 16 \
83
- --tp-size 1 \
84
- --attention-backend fa3 \
85
- --mem-fraction-static 0.75 \
86
- --mamba-scheduler-strategy extra_buffer \
87
- --trust-remote-code
88
- ```
89
-
90
- ### Usage
91
-
92
- ```python
93
- from openai import OpenAI
94
-
95
- client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
96
-
97
- response = client.chat.completions.create(
98
- model="Qwen/Qwen3.6-27B",
99
- messages=[{"role": "user", "content": "Write a quicksort in Python."}],
100
- max_tokens=4096,
101
- temperature=0.0
102
- )
103
- print(response.choices[0].message.content)
104
- ```
105
-
106
- ## Benchmark Results
107
-
108
- N/A
109
-
110
- ## Acknowledgements
111
-
112
- Special thanks to [David Wang](https://davidwa.ng/) for his outstanding engineering support on this project. We are also grateful to [Modal](https://modal.com/), [InnoMatrix](https://innomatrix.ai), and [Yotta Labs](https://www.yottalabs.ai/) for providing the compute resources used to train this draft model.
113
-
114
- ## Citation
115
-
116
- If you find DFlash useful, please cite our work. To share feedback on DFlash or request new model support, please fill out this form: [DFlash Feedback](https://forms.gle/4YNwfqb4nJdqn6hq9).
117
-
118
- ```bibtex
119
- @article{chen2026dflash,
120
- title = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
121
- author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
122
- journal = {arXiv preprint arXiv:2602.06036},
123
- year = {2026}
124
- }
125
  ```
 
1
+ ---
2
+ base_model: z-lab/Qwen3.6-27B-DFlash
3
+ tags:
4
+ - transformers
5
+ - safetensors
6
+ - qwen3
7
+ - feature-extraction
8
+ - dflash
9
+ - speculative-decoding
10
+ - diffusion
11
+ - efficiency
12
+ - flash-decoding
13
+ - qwen
14
+ - diffusion-language-model
15
+ - text-generation
16
+ - custom_code
17
+ - arxiv:2602.06036
18
+ - license:mit
19
+ - text-generation-inference
20
+ - endpoints_compatible
21
+ - region:us
22
+ ---
23
+
24
+ # Qwen 3.6 27B DFlash GGUF
25
+
26
+ GGUF quantizations of [**z-lab DFlash draft model**](https://huggingface.co/z-lab/Qwen3.6-27B-DFlash) for [**Qwen 3.6 27B**](https://huggingface.co/Qwen/Qwen3.6-27B).
27
+
28
+ Use with [BeeLlama.cpp](https://github.com/Anbeeld/beellama.cpp), a llama.cpp fork with advanced quantization features.
29
+
30
+ ---
31
+
32
+
33
+ # Qwen3.6-27B-DFlash
34
+ [**Paper**](https://arxiv.org/abs/2602.06036) | [**GitHub**](https://github.com/z-lab/dflash) | [**Blog**](https://z-lab.ai/projects/dflash/)
35
+
36
+ **This model is still under training, and inference engine support may not be fully available yet due to architectural changes, including causal SWA layers.**
37
+
38
+ **DFlash** is a novel speculative decoding method that utilizes a lightweight **block diffusion** model for drafting. It enables efficient, high-quality parallel drafting that pushes the limits of inference speed.
39
+
40
+ This model is the **drafter** component. It must be used in conjunction with the target model `Qwen/Qwen3.6-27B`.
41
+
42
+ <div align="center">
43
+ <img src="assets/dflash_system.png" alt="DFlash Architecture" width="100%">
44
+ </div>
45
+
46
+ ## Quick Start
47
+
48
+ ### Installation
49
+
50
+ vLLM (We temporarily modify the installation through this PR to support interleaved SWA and ensure correct handling of target hidden states for optimal performance):
51
+ ```bash
52
+ uv pip install vllm
53
+ uv pip install -U --torch-backend=auto "vllm @ git+https://github.com/vllm-project/vllm.git@refs/pull/40898/head"
54
+ ```
55
+
56
+ SGLang:
57
+ ```bash
58
+ uv pip install "git+https://github.com/sgl-project/sglang.git@refs/pull/23000/head#subdirectory=python"
59
+ ```
60
+
61
+ ### Launch Server
62
+
63
+ vLLM:
64
+ ```bash
65
+ vllm serve Qwen/Qwen3.6-27B \
66
+ --speculative-config '{"method": "dflash", "model": "z-lab/Qwen3.6-27B-DFlash", "num_speculative_tokens": 15}' \
67
+ --attention-backend flash_attn \
68
+ --max-num-batched-tokens 32768
69
+ ```
70
+
71
+ SGLang:
72
+ ```bash
73
+ # Optional: enable schedule overlapping (experimental, may not be stable)
74
+ # export SGLANG_ENABLE_SPEC_V2=1
75
+ # export SGLANG_ENABLE_DFLASH_SPEC_V2=1
76
+ # export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1
77
+
78
+ python -m sglang.launch_server \
79
+ --model-path Qwen/Qwen3.6-27B \
80
+ --speculative-algorithm DFLASH \
81
+ --speculative-draft-model-path z-lab/Qwen3.6-27B-DFlash \
82
+ --speculative-num-draft-tokens 16 \
83
+ --tp-size 1 \
84
+ --attention-backend fa3 \
85
+ --mem-fraction-static 0.75 \
86
+ --mamba-scheduler-strategy extra_buffer \
87
+ --trust-remote-code
88
+ ```
89
+
90
+ ### Usage
91
+
92
+ ```python
93
+ from openai import OpenAI
94
+
95
+ client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
96
+
97
+ response = client.chat.completions.create(
98
+ model="Qwen/Qwen3.6-27B",
99
+ messages=[{"role": "user", "content": "Write a quicksort in Python."}],
100
+ max_tokens=4096,
101
+ temperature=0.0
102
+ )
103
+ print(response.choices[0].message.content)
104
+ ```
105
+
106
+ ## Benchmark Results
107
+
108
+ N/A
109
+
110
+ ## Acknowledgements
111
+
112
+ Special thanks to [David Wang](https://davidwa.ng/) for his outstanding engineering support on this project. We are also grateful to [Modal](https://modal.com/), [InnoMatrix](https://innomatrix.ai), and [Yotta Labs](https://www.yottalabs.ai/) for providing the compute resources used to train this draft model.
113
+
114
+ ## Citation
115
+
116
+ If you find DFlash useful, please cite our work. To share feedback on DFlash or request new model support, please fill out this form: [DFlash Feedback](https://forms.gle/4YNwfqb4nJdqn6hq9).
117
+
118
+ ```bibtex
119
+ @article{chen2026dflash,
120
+ title = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
121
+ author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
122
+ journal = {arXiv preprint arXiv:2602.06036},
123
+ year = {2026}
124
+ }
125
  ```