DN-MOPD-Qwen3.5-9B-teacher-math

The Qwen3.5-9B mathematics expert used as a frozen teacher in the DN-MOPD paper: Qwen3.5-9B trained with GRPO on mathematics prompts. It is one of three same-size experts (math, code, IF) that the Qwen3.5-9B students learn from.

Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD

Model details

Base model Qwen/Qwen3.5-9B
Role Mathematics expert; frozen teacher of the 9B students
Training GRPO from the base model, 250 updates, seed 42
Precision bfloat16
Chat format non-thinking (enable_thinking=False)
License Apache-2.0 (same as the base model)

Training recipe

  • Algorithm: GRPO on mathematics prompts with a verifiable reward; no KL or entropy term.
  • Batching: 128 prompts per rollout, 8 responses per prompt, 256 responses per optimizer step. Dynamic sampling drops prompt groups without reward variation (at most 8 generation batches per rollout).
  • Lengths: prompt ≤ 2,048 tokens, response ≤ 8,192 tokens, temperature 1.0.
  • Optimizer: Adam, learning rate 1e-6 (constant after 10 warm-up updates), betas (0.9, 0.98), weight decay 0.1, gradient clipping 1.0.
  • Length of training: 250 updates (runs were capped at 400), seed 42.

The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md.

Usage

This model was trained and evaluated with the non-thinking chat format. Pass enable_thinking=False to the chat template. Qwen3.5-9B's chat template enables thinking by default, so this argument is required. The evaluation settings in the paper were temperature 1.0 and top-p 1.0, with up to 16,384 new tokens (8,192 in the appendix).

vLLM (the paper used vLLM 0.18.0):

from vllm import LLM, SamplingParams

llm = LLM(model="XINLI1997/DN-MOPD-Qwen3.5-9B-teacher-math", max_model_len=32768)
params = SamplingParams(temperature=1.0, top_p=1.0, max_tokens=16384, seed=42)
messages = [{"role": "user", "content": "Find the sum of all positive divisors of 36. Put the final answer in \\boxed{}."}]
outputs = llm.chat(messages, params, chat_template_kwargs={"enable_thinking": False})
print(outputs[0].outputs[0].text)

Transformers (Qwen3.5 needs transformers>=5; the paper's training environment used 5.12.1):

import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-9B-teacher-math")
model = AutoModelForImageTextToText.from_pretrained("XINLI1997/DN-MOPD-Qwen3.5-9B-teacher-math", dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "Write a Python function that returns the n-th Fibonacci number."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=4096, do_sample=True, temperature=1.0, top_p=1.0)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Evaluation

Paper Table 1 (Qwen3.5-9B):

Model AIME25 AIME26 LCB v5 LCB v6 IFEval IFBench Total
DN-MOPD-Qwen3.5-9B-teacher-math (Math expert) 59.7 69.4 53.4 49.4 82.3 34.9 58.2
Initial student (Qwen3.5-9B) 57.7 62.6 54.9 51.4 82.4 33.8 57.1

Scores (%) from the paper; training seed 42; 16,384-token evaluation cap; non-thinking chat template; temperature 1.0, top-p 1.0, generation seed 42. AIME25/AIME26: avg@64. LiveCodeBench v5/v6 (167/175 disjoint problems): avg@6. IFEval/IFBench: strict prompt accuracy, avg@16. Total: mean of the six task scores.

Files

  • Weights in Hugging Face format (Qwen3_5ForConditionalGeneration, bfloat16), exported from the FSDP training checkpoint.
  • The export omits the 15 multi-token-prediction tensors (mtp.*) of the base model. All other tensors have the base model's names and shapes. MTP-based speculative decoding is therefore not available with this checkpoint. Ordinary decoding is unaffected: the paper's evaluations used exactly these files.
  • config.json, the tokenizer files and chat_template.jinja are the base model's, unchanged.
  • The vision encoder is carried over from the base model. Training and evaluation used text only.
  • LICENSE is the base model's Apache-2.0 license.

Limitations

  • A specialist: it was trained for one domain only. In the table above it scores below the initial model on LCB v5, LCB v6, IFEval.
  • Trained with responses of at most 8,192 tokens and evaluated only in non-thinking mode; thinking mode, multimodal inputs, other languages and safety behaviour were not evaluated beyond the base model.

Citation

@article{li2026dnmopd,
  title   = {Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation},
  author  = {Li, Xin and Jiang, Hao and Gao, Xin and Wang, Annan and Xie, Yuchen and Guo, Jinghao and Qu, Xingwei and Zhang, Yichi and Yuen, Chau},
  journal = {arXiv preprint arXiv:2609.35347},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.35347}
}

This model is a fine-tuned derivative of Qwen/Qwen3.5-9B by the Qwen team, released under the Apache License 2.0.

Downloads last month
14
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for XINLI1997/DN-MOPD-Qwen3.5-9B-teacher-math

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(991)
this model

Collection including XINLI1997/DN-MOPD-Qwen3.5-9B-teacher-math

Paper for XINLI1997/DN-MOPD-Qwen3.5-9B-teacher-math