--- license: other license_name: qwen-research license_link: https://huggingface.co/Qwen/Qwen2.5-3B/blob/main/LICENSE base_model: Qwen/Qwen2.5-3B tags: - distillation - reproduction - distilldetect - arxiv:2607.09692 --- # DistillDetect-Qwen2.5-3B-from-gpt-oss-120b-s1 **Unofficial reproduction** of a distilled student model from the paper *Reference-Based Distillation Detection in LLMs* (Rawat et al., [arXiv:2607.09692](https://arxiv.org/abs/2607.09692)), retrained with the authors' released code and teacher-generated data ([github.com/RajatRawat-creator/DistillDetect](https://github.com/RajatRawat-creator/DistillDetect), MIT). The original authors did not release student checkpoints; this repo is an independent reproduction and is **not** affiliated with the authors. - **Base (student) model:** [Qwen/Qwen2.5-3B](https://huggingface.co/Qwen/Qwen2.5-3B) - **Teacher:** openai/gpt-oss-120b - **Training data:** s1 (1K prompts) — 1000 teacher-generated responses, shipped verbatim in the authors' repo (`data/training/Teacher=GPT-OSS-120B_Data=S1_Template=Chat.jsonl`) - **Prompt template:** plain `Problem:\n{question}\n\nSolution:\n` ## Training SFT with the authors' released `training/` scripts (paper Appendix A recipe): 3 epochs, LR 1e-5, cosine schedule, 5% warmup, per-device batch 4 x grad-accum 4 (effective batch 16), block size 4096, bf16, gradient checkpointing, loss on response tokens only (prompt masked -100). Teacher responses were pre-truncated to 2,048 tokens in the released data. Trained on 1x H100 (paper used 2x H200; hyperparameters identical), transformers 4.55.4 / trl 0.19.1, seed 42 (HF default; paper seed unknown). One compatibility patch: trl renamed `max_seq_length` to `max_length`, so the block size is passed explicitly (no behavioral change). ## Evaluation (ours vs. paper Table 9) Greedy decoding, template matched to training, scored with `math_verify`. GSM8K 4-shot; MATH500 zero-shot. Few-shot counts were calibrated so the *base* models reproduce their Table 9 baselines. The paper does not document its eval protocol, so treat cross-paper comparisons as approximate. | Benchmark | This reproduction | Paper Table 9 | Gen. budget | Hit budget cap | |---|---|---|---|---| | GSM8K | 82.49 | 79.90 | 4096 tok | 1.0% | | MATH500 | 44.00 | 53.40 | 16384 tok | 49.8% | A high "hit budget cap" fraction means the model was still generating when the token budget ran out, so that accuracy is a **lower bound** — these students are SFT'd on teacher traces that were themselves truncated at 2,048 tokens, which makes some of them generate very long self-checking traces. Base-model reference (our protocol / paper): GSM8K 67.63 / 75.82, MATH500 40.20 / 39.80. ## License The base model's license applies: Qwen Research License. Teacher-generated training data redistributed by the paper authors under their repo's MIT license.