Qwen3.8-27B-DSpark (mirror)

This is an unmodified mirror of RadixArk/Qwen3.8-27B-DSpark, re-hosted so the Qwen3.8-27B-abliterated-MLX-{8,4}bit models are self-contained for DSpark speculative decoding on Apple Silicon. All credit for the drafter goes to RadixArk. If the original repo is available, prefer it — mlx-dspark resolves it automatically.

Why it's here

DSpark speculative decoding needs two weights: the target model (the abliterated Qwen3.8-27B) and this ~1.36B drafter. mlx-dspark auto-downloads the drafter from RadixArk by default, so no manual assembly is needed. This mirror exists only as a fallback so the package keeps working even if the upstream repo moves.

Use with the abliterated MLX target

pip install mlx-dspark
# auto (uses RadixArk upstream): 
mlx-dspark generate --model ./Qwen3.8-27B-abliterated-MLX-8bit --mode dspark --prompt "..."
# explicit (uses this mirror): 
mlx-dspark generate --model ./Qwen3.8-27B-abliterated-MLX-8bit \
  --drafter ./Qwen3.8-27B-DSpark-drafter --mode dspark --prompt "..."

Note: this drafter was trained against the original Qwen3.8-27B. On the abliterated target it still produces correct output (the target verifies every token — speculative decoding is lossless), but the accepted-token rate may be a little lower on prompts the original model would have refused.


Original model card (RadixArk/Qwen3.8-27B-DSpark)

A DSpark speculator for Qwen/Qwen3.8-27B-FP8. DSpark extends DFlash with target-model auxiliary features and a confidence head that dynamically chooses the number of draft tokens. Trained with SpecForge, served with SGLang.

  • Draft parameters: 1,359,284,737 (1.36B) · BF16 · hidden 5,120 · 5 full-attention layers · GQA 40Q/8KV
  • Target auxiliary feature layers: 4, 16, 28, 40, 52 · confidence head: Markov, rank 256
  • DSpark block size: 7 draft tokens (verify width 8) · max positions 262,144
  • Mean acceptance length across 11 workloads (SGLang, FP8 target): 3.35–3.39
Downloads last month
892
Safetensors
Model size
1B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jialinyyzz/Qwen3.8-27B-DSpark-drafter

Base model

Qwen/Qwen3.8-27B
Finetuned
(7)
this model