Qwen3-32B-Dense-3B-D2D

Research artifact, not a general-purpose model. These weights are released so that the experiments in the paper Pruning and Distilling Mixture-of-Experts into Dense Language Models can be reproduced. Each student is distilled on a fixed, small token budget (~4B tokens for Qwen3, 0.3B tokens for DeepSeek-V2-Lite and GPT-OSS) purely so that expert scoring and grouping methods can be compared under an equal budget. Absolute quality is therefore far below the teacher and below pretrained models of the same size, and no instruction tuning or alignment was applied. Please do not use this as an off-the-shelf assistant.

What this is

A dense student obtained by pruning and distilling a Mixture-of-Experts teacher, from the paper Pruning and Distilling Mixture-of-Experts into Dense Language Models. Code: https://github.com/krafton-ai/moe-to-dense

Dense-to-dense (D2D) pruning baseline. Unlike the other students in this collection its teacher is the dense Qwen3-32B, not the MoE Qwen3-30B-A3B. It is included so the MoE-to-dense route can be compared against dense pruning at a matched student size and a matched distillation budget.

Results

Full comparison group for this architecture, extended training after ~4B tokens (paper Table 6). This model's row is in bold, and rows whose weights are also released link to them.

Configuration Wino Hella ARC-E ARC-C MMLU Avg
DO-ACP, K=8 63.1 60.3 75.6 45.4 46.1 58.10
SF, K=16 61.2 56.3 74.0 43.1 32.7 53.46
D2D pruning (Qwen3-32B to 3.4B) 60.5 57.5 73.1 41.5 26.6 51.84
Random FFN + teacher attn 54.4 45.4 66.0 34.2 27.1 45.44
Qwen3-1.7B (pretrained reference) 66.1 67.1 81.9 55.5 62.6 66.63
Qwen3-4B (pretrained reference) 72.0 75.8 86.2 64.6 73.1 74.34

Downstream accuracy is Winogrande 5-shot, HellaSwag 10-shot, ARC-Easy 25-shot, ARC-Challenge 25-shot and MMLU 5-shot. Avg is the unweighted mean of the five benchmarks.

Configuration

Field Value
Teacher Qwen/Qwen3-32B
Student parameters 3.44B
Distillation data FineWeb-Edu (sample-10BT), ~4B tokens

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("jhyuckkim/Qwen3-32B-Dense-3B-D2D", dtype="bfloat16")
tok = AutoTokenizer.from_pretrained("jhyuckkim/Qwen3-32B-Dense-3B-D2D")

The student is a standard Qwen3ForCausalLM, so it loads with stock transformers and runs in vLLM without extra code.

Citation

@article{kim2026pruning,
  title={Pruning and Distilling Mixture-of-Experts into Dense Language Models},
  author={Kim, Junhyuck and Yun, Jihun and Kim, Haechan and Kim, Gyeongman and Bae, Joonghyun and Cho, Jaewoong},
  journal={arXiv preprint arXiv:2605.28207},
  year={2026}
}
Downloads last month
98
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jhyuckkim/Qwen3-32B-Dense-3B-D2D

Base model

Qwen/Qwen3-32B
Finetuned
(591)
this model

Collection including jhyuckkim/Qwen3-32B-Dense-3B-D2D

Paper for jhyuckkim/Qwen3-32B-Dense-3B-D2D