grpo-llama3.1-8b-0623-final
Base model: meta-llama/Llama-3.1-8B-Instruct
Method: GRPO (Group Relative Policy Optimization, TRL) with dual-RM reward signal.
Final checkpoint = load_best_model_at_end=True로 골라진 best step 5,185 가중치.
Reward models used during training (frozen)
Training
| Hyperparameter |
Value |
| Algorithm |
GRPO (TRL) |
| Step at best ckpt (saved as final) |
5,185 |
| Epoch at best ckpt |
1.6989 / 2 |
| Total training step |
6,100 (≈ 2 epochs completed) |
| Group size (G) |
8 |
| Per-device batch |
8 |
| Gradient accumulation |
2 |
| Learning rate |
1e-6 |
| Warmup ratio |
0.1 |
| KL coefficient (β) |
0.04 |
| Clip ε |
0.2 |
| λ (verifiable penalty) |
0.75 |
| Format penalty |
0.5 |
| Answer-tag c |
0.5 |
| Max completion length |
1024 |
| Optimizer |
DeepSpeed ZeRO-2 + CPU offload |
| Generation backend |
vLLM (server mode, port 8001) |
| Seed |
42 |
Data (in-distribution 90/10 split of train_subset.json)
| Split |
Source |
Rows |
| Train |
train_subset_90pct.json (90% stratified by error_type) |
6,105 |
| Eval (during training) |
train_subset_10pct.json (in-distribution 10% held-out, stratified, seed=42) |
678 |
Performance
| Metric |
Value |
Step |
| eval_reward (best, saved as final) |
1.2560 |
5,185 |
| eval_loss (at best) |
0.0187 |
5,185 |
| eval_reward (last) |
1.2236 |
6,100 |
| eval_loss (last) |
0.0308 |
6,100 |
| train reward (last 100-step window) |
≈ 1.20 |
— |
License
Apache-2.0. Base model is governed by the Llama 3.1 Community License — verify compliance for your use case.