AlexAtomic's picture
Upload README.md with huggingface_hub
9017c28 verified
|
Raw
History Blame
4.31 kB
metadata
license: mit
base_model:
  - z-lab/Qwen3-8B-DFlash-b16
base_model_relation: quantized
quantized_by: AlexAtomic
pipeline_tag: text-generation
library_name: gguf
tags:
  - atomic-chat
  - dflash
  - speculative-decoding
  - draft-model
  - qwen
  - gguf
  - llama.cpp
Atomic Chat Join Discord GitHub

DFlash

Qwen3 8B DFlash, the DFlash speculative-decoding draft converted to GGUF by Atomic Chat. Built straight from z-lab's original weights. Runs fully offline.

What this is

DFlash is a speculative-decoding method that drafts a whole block of candidate tokens in a single forward pass using a lightweight block-diffusion model, instead of one token at a time. This repo is the draft component only — it does nothing on its own. You run it alongside the target model Qwen/Qwen3-8B, which verifies the drafted block and keeps the longest correct prefix. Output is identical to running the target alone, just faster.

These GGUFs are converted from z-lab's original weights, not a repack of someone else's GGUF. Q8_0 and bf16 give the same draft acceptance, so Q8_0 is the pick.

Run in llama.cpp

Needs a build of llama.cpp with DFlash speculative decoding (PR #22105). You supply the target as -m and this draft as -md:

./llama-server \
    -m   Qwen3-8B.gguf \
    -md  Qwen3-8B-DFlash.Q8_0.gguf \
    --spec-type draft-dflash --spec-draft-n-max 15 \
    -ngl 99 -fa on --jinja -c 8192

DFlash is trained for non-thinking generation — pass enable_thinking=false in the chat template for best acceptance.

Choosing a quant

Quant Size Notes
Q8_0 1.12 GB Recommended. Same acceptance as bf16, half the size and slightly faster drafting.
bf16 2.10 GB Full-precision draft (reference). No acceptance gain over Q8_0.

Performance

z-lab report up to 6.17x lossless acceleration for Qwen3-8B on their reference stack (vLLM / SGLang / Transformers). In llama.cpp today the DFlash port is newer: on our RTX 4090 test (Q8_0 target + Q8_0 draft, code generation) it delivered about 2.3x end-to-end at roughly 26% draft acceptance. Acceptance is set by the implementation and the content, not by the quantization (bf16 and Q8_0 measure the same). Speedups grow on structured/code output and shrink on free-form prose.

How this was made

  1. Download the DFlash draft z-lab/Qwen3-8B-DFlash-b16 (original weights).
  2. Convert to GGUF with llama.cpp convert_hf_to_gguf.py --target-model-dir (the target supplies the tokenizer; its weights are not needed).
  3. Quantize the draft to Q8_0.

License

Released by z-lab under the MIT license. Converted to GGUF by Atomic Chat. See the DFlash paper and project page.