YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Qwen3-8B Speculators (DFlash & FLM)

Speculator models for the Qwen/Qwen3-8B target model.

Architectures

Model File mlp_type Draft Layers Style
dflash_5layer_mix.pt dflash 5 DFlash-style
flm_5layer_uni.pt flm 5 FLM-style

Training Data

These models were trained on the regenerated dataset (DFlash-style) derived from Qwen3-8B's own responses to the following sources:

  • Nemotron Post-Training V2
  • CodeAlpaca
  • ShareGPT
  • UltraChat-200K

Specifically:

  • dflash_5layer_mix: Uses a mixture of datasets and a custom Tau schedule (30% pure noise, 20% grid-aligned, 50% uniform).
  • flm_5layer_uni: Uses a uniform Tau schedule and standard dataset mixture.

Details

  • Target Model: Qwen/Qwen3-8B
  • Sequence Length: 2048
  • Batch Size: 5
  • Draft Layers: 5 (despite some folder naming suggesting 4, internal config confirms 5)

Trained at Princeton AILab.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support