Instructions to use jhyuckkim/Qwen3-32B-Dense-3B-D2D with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jhyuckkim/Qwen3-32B-Dense-3B-D2D with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jhyuckkim/Qwen3-32B-Dense-3B-D2D") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("jhyuckkim/Qwen3-32B-Dense-3B-D2D") model = AutoModelForCausalLM.from_pretrained("jhyuckkim/Qwen3-32B-Dense-3B-D2D", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jhyuckkim/Qwen3-32B-Dense-3B-D2D with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jhyuckkim/Qwen3-32B-Dense-3B-D2D" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jhyuckkim/Qwen3-32B-Dense-3B-D2D", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jhyuckkim/Qwen3-32B-Dense-3B-D2D
- SGLang
How to use jhyuckkim/Qwen3-32B-Dense-3B-D2D with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jhyuckkim/Qwen3-32B-Dense-3B-D2D" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jhyuckkim/Qwen3-32B-Dense-3B-D2D", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jhyuckkim/Qwen3-32B-Dense-3B-D2D" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jhyuckkim/Qwen3-32B-Dense-3B-D2D", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use jhyuckkim/Qwen3-32B-Dense-3B-D2D with Docker Model Runner:
docker model run hf.co/jhyuckkim/Qwen3-32B-Dense-3B-D2D
Qwen3-32B-Dense-3B-D2D
Research artifact, not a general-purpose model. These weights are released so that the experiments in the paper Pruning and Distilling Mixture-of-Experts into Dense Language Models can be reproduced. Each student is distilled on a fixed, small token budget (~4B tokens for Qwen3, 0.3B tokens for DeepSeek-V2-Lite and GPT-OSS) purely so that expert scoring and grouping methods can be compared under an equal budget. Absolute quality is therefore far below the teacher and below pretrained models of the same size, and no instruction tuning or alignment was applied. Please do not use this as an off-the-shelf assistant.
What this is
A dense student obtained by pruning and distilling a Mixture-of-Experts teacher, from the paper Pruning and Distilling Mixture-of-Experts into Dense Language Models. Code: https://github.com/krafton-ai/moe-to-dense
Dense-to-dense (D2D) pruning baseline. Unlike the other students in this collection its teacher is the dense Qwen3-32B, not the MoE Qwen3-30B-A3B. It is included so the MoE-to-dense route can be compared against dense pruning at a matched student size and a matched distillation budget.
Results
Full comparison group for this architecture, extended training after ~4B tokens (paper Table 6). This model's row is in bold, and rows whose weights are also released link to them.
| Configuration | Wino | Hella | ARC-E | ARC-C | MMLU | Avg |
|---|---|---|---|---|---|---|
| DO-ACP, K=8 | 63.1 | 60.3 | 75.6 | 45.4 | 46.1 | 58.10 |
| SF, K=16 | 61.2 | 56.3 | 74.0 | 43.1 | 32.7 | 53.46 |
| D2D pruning (Qwen3-32B to 3.4B) | 60.5 | 57.5 | 73.1 | 41.5 | 26.6 | 51.84 |
| Random FFN + teacher attn | 54.4 | 45.4 | 66.0 | 34.2 | 27.1 | 45.44 |
| Qwen3-1.7B (pretrained reference) | 66.1 | 67.1 | 81.9 | 55.5 | 62.6 | 66.63 |
| Qwen3-4B (pretrained reference) | 72.0 | 75.8 | 86.2 | 64.6 | 73.1 | 74.34 |
Downstream accuracy is Winogrande 5-shot, HellaSwag 10-shot, ARC-Easy 25-shot, ARC-Challenge 25-shot and MMLU 5-shot. Avg is the unweighted mean of the five benchmarks.
Configuration
| Field | Value |
|---|---|
| Teacher | Qwen/Qwen3-32B |
| Student parameters | 3.44B |
| Distillation data | FineWeb-Edu (sample-10BT), ~4B tokens |
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("jhyuckkim/Qwen3-32B-Dense-3B-D2D", dtype="bfloat16")
tok = AutoTokenizer.from_pretrained("jhyuckkim/Qwen3-32B-Dense-3B-D2D")
The student is a standard Qwen3ForCausalLM, so it loads with stock
transformers and runs in vLLM without extra code.
Citation
@article{kim2026pruning,
title={Pruning and Distilling Mixture-of-Experts into Dense Language Models},
author={Kim, Junhyuck and Yun, Jihun and Kim, Haechan and Kim, Gyeongman and Bae, Joonghyun and Cho, Jaewoong},
journal={arXiv preprint arXiv:2605.28207},
year={2026}
}
- Downloads last month
- 98
Model tree for jhyuckkim/Qwen3-32B-Dense-3B-D2D
Base model
Qwen/Qwen3-32B