Qwen3.5-2B - dr_grpo/threshold (adamw)
vs base Qwen3.5-2B: InD acc 83.8→26.0, total output tokens 3240→76 (-98%)
gpqa_diamond (OOD) acc 7.6→17.0 (+9.4 pp, +124%)
Trained via GRPO with dr_grpo loss, threshold reward shape (alpha=0.1), adamw optimizer, lr=2.0e-06, G=8, max_steps=200, max_completion_length=8000, evaluated over 3 seeds.
Accuracy vs base Qwen3.5-2B
| Dataset | Base | Tuned (mean ± std) | Δ (pp, rel %) |
|---|---|---|---|
| gsm8k | 81.2 | 0.5 ± 0.9 | -80.7 pp, -99% |
| arc_challenge | 85.2 | 27.7 ± 4.9 | -57.5 pp, -68% |
| arc_easy | 97.7 | 39.0 ± 6.0 | -58.7 pp, -60% |
| commonsenseqa | 69.8 | 25.7 ± 7.5 | -44.2 pp, -63% |
| openbookqa | 82.3 | 19.0 ± 4.8 | -63.3 pp, -77% |
| qasc | 76.5 | 33.0 ± 6.9 | -43.5 pp, -57% |
| sciq | 93.8 | 37.2 ± 5.3 | -56.7 pp, -60% |
| mmlu_pro(OOD) | 33.8 | 11.0 ± 0.5 | -22.8 pp, -67% |
| mmlu_redux(OOD) | 52.5 | 22.0 ± 3.5 | -30.5 pp, -58% |
| gpqa_diamond(OOD) | 7.6 | 17.0 ± 0.3 | +9.4 pp, +124% |
| InD Average | 83.8 | 26.0 ± 3.8 | -57.8 pp, -69% |
| OOD | 31.4 | 16.7 ± 1.0 | -14.7 pp, -47% |
| ALL | 68.1 | 23.2 ± 2.9 | -44.9 pp, -66% |
Δ shows the absolute change in accuracy points (pp) and the relative percent change (tuned − base) / base × 100 (rel %, shown as n/a when base accuracy is 0).
Output tokens (total) vs base Qwen3.5-2B
| Dataset | Base | Tuned (mean ± std) | Reduction % |
|---|---|---|---|
| gsm8k | 4450 | 158 ± 38 | -96% |
| arc_challenge | 3157 | 63 ± 9 | -98% |
| arc_easy | 1871 | 58 ± 17 | -97% |
| commonsenseqa | 3949 | 58 ± 24 | -99% |
| openbookqa | 3378 | 55 ± 21 | -98% |
| qasc | 3932 | 87 ± 44 | -98% |
| sciq | 1944 | 54 ± 16 | -97% |
| mmlu_pro(OOD) | 6582 | 85 ± 13 | -99% |
| mmlu_redux(OOD) | 5589 | 64 ± 15 | -99% |
| gpqa_diamond(OOD) | 8001 | 73 ± 14 | -99% |
| InD Average | 3240 | 76 ± 23 | -98% |
| OOD | 6720 | 74 ± 11 | -99% |
| ALL | 4281 | 76 ± 20 | -98% |
Output tokens = total generated tokens (full completion), 3-seed mean.
Reduction = percentage decrease in mean output tokens vs base Qwen3.5-2B (negative reduction, i.e. +, means the tuned model generates more tokens).
- Downloads last month
- 12