DistillDetect-Qwen2.5-1.5B-from-Llama-3.3-70B-Instruct-s1
Unofficial reproduction of a distilled student model from the paper Reference-Based Distillation Detection in LLMs (Rawat et al., arXiv:2607.09692), retrained with the authors' released code and teacher-generated data (github.com/RajatRawat-creator/DistillDetect, MIT). The original authors did not release student checkpoints; this repo is an independent reproduction and is not affiliated with the authors.
- Base (student) model: Qwen/Qwen2.5-1.5B
- Teacher: nvidia/Llama-3.3-70B-Instruct-NVFP8
- Training data: s1 (1K prompts) — 1000 teacher-generated responses, shipped verbatim in the authors' repo (
data/training/Teacher=Nvidia-Llama-3.3-70B-Instruct_Data=S1_Template=Chat.jsonl) - Prompt template: plain
Problem:\n{question}\n\nSolution:\n
Training
SFT with the authors' released training/ scripts (paper Appendix A recipe):
3 epochs, LR 1e-5, cosine schedule, 5% warmup, per-device batch 4 x grad-accum 4
(effective batch 16), block size 4096, bf16, gradient checkpointing, loss on
response tokens only (prompt masked -100). Teacher responses were pre-truncated
to 2,048 tokens in the released data. Trained on 1x H100
(paper used 2x H200; hyperparameters identical), transformers 4.55.4 / trl 0.19.1,
seed 42 (HF default; paper seed unknown). One compatibility patch: trl renamed
max_seq_length to max_length, so the block size is passed explicitly
(no behavioral change).
Evaluation (ours vs. paper Table 9)
Greedy decoding, template matched to training, scored with math_verify.
GSM8K 4-shot; MATH500 zero-shot. Few-shot counts were calibrated so the
base models reproduce their Table 9 baselines. The paper does not document its
eval protocol, so treat cross-paper comparisons as approximate.
| Benchmark | This reproduction | Paper Table 9 | Gen. budget | Hit budget cap |
|---|---|---|---|---|
| GSM8K | 69.22 | 69.44 | 4096 tok | 0.6% |
| MATH500 | 42.80 | 34.00 | 16384 tok | 11.8% |
A high "hit budget cap" fraction means the model was still generating when the token budget ran out, so that accuracy is a lower bound — these students are SFT'd on teacher traces that were themselves truncated at 2,048 tokens, which makes some of them generate very long self-checking traces.
Base-model reference (our protocol / paper): GSM8K 58.07 / 67.25, MATH500 35.40 / 32.60.
License
The base model's license applies: apache-2.0. Teacher-generated training data redistributed by the paper authors under their repo's MIT license.
- Downloads last month
- 23
Model tree for francescortu/DistillDetect-Qwen2.5-1.5B-from-Llama-3.3-70B-Instruct-s1
Base model
Qwen/Qwen2.5-1.5B