# Qwen3-8B Speculators (DFlash & FLM) Speculator models for the [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) target model. ## Architectures | Model File | `mlp_type` | Draft Layers | Style | |------------|------------|--------------|-------| | `dflash_5layer_mix.pt` | `dflash` | 5 | DFlash-style | | `flm_5layer_uni.pt` | `flm` | 5 | FLM-style | ## Training Data These models were trained on the **regenerated dataset** (DFlash-style) derived from Qwen3-8B's own responses to the following sources: - **Nemotron Post-Training V2** - **CodeAlpaca** - **ShareGPT** - **UltraChat-200K** Specifically: - **dflash_5layer_mix**: Uses a mixture of datasets and a custom Tau schedule (30% pure noise, 20% grid-aligned, 50% uniform). - **flm_5layer_uni**: Uses a uniform Tau schedule and standard dataset mixture. ## Details - **Target Model**: Qwen/Qwen3-8B - **Sequence Length**: 2048 - **Batch Size**: 5 - **Draft Layers**: 5 (despite some folder naming suggesting 4, internal config confirms 5) --- *Trained at Princeton AILab.*