ahmed-3m commited on
Commit
6325288
·
verified ·
1 Parent(s): 0c0c1b6

Add model card with fair evaluation results

Browse files
Files changed (1) hide show
  1. README.md +39 -0
README.md ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen2.5-1.5B-Instruct
3
+ pipeline_tag: text-generation
4
+ library_name: transformers
5
+ tags:
6
+ - qwen2.5
7
+ - gsm8k
8
+ - svamp
9
+ - sdpo
10
+ - online-dpo
11
+ - self-training
12
+ ---
13
+
14
+ # Phase 2 Final Model
15
+
16
+ Merged full model produced by the Phase 2 online-DPO / SDPO-style experiment in
17
+ the Lab 0 self-training project.
18
+
19
+ ## Fair evaluation
20
+
21
+ Deterministic greedy exact-match evaluation with the required `#### <number>`
22
+ answer parser:
23
+
24
+ - Full-set eval: **45.41% GSM8K / 57.67% SVAMP**
25
+ - Second fixed-subset eval: **48% GSM8K / 50% SVAMP**
26
+
27
+ These results currently outperform the old Phase 3 REINFORCE checkpoints under
28
+ the same fair-eval setup.
29
+
30
+ ## Notes
31
+
32
+ - Base model: `Qwen/Qwen2.5-1.5B-Instruct`
33
+ - Hardware/training constraints: fp16 + SDPA, no bf16, no Flash Attention 2
34
+ - This is a research artifact, not a production math model
35
+
36
+ ## Extra files
37
+
38
+ - `phase2_final_full_eval.json`
39
+ - `phase2_vs_phase3_independent_shuffle.json`