YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

grpo-llama3.1-8b-0623-final

Base model: meta-llama/Llama-3.1-8B-Instruct Method: GRPO (Group Relative Policy Optimization, TRL) with dual-RM reward signal. Final checkpoint = load_best_model_at_end=True로 골라진 best step 5,185 가중치.

Reward models used during training (frozen)

Slot Model
RM 1 WooYoungSeok/rm-qwen2.5-math-7b-0622 — verifier-set acc 0.8354
RM 2 WooYoungSeok/rm-deepseek-r1-qwen3-8b-0622 — verifier-set acc 0.8497

Training

Hyperparameter Value
Algorithm GRPO (TRL)
Step at best ckpt (saved as final) 5,185
Epoch at best ckpt 1.6989 / 2
Total training step 6,100 (≈ 2 epochs completed)
Group size (G) 8
Per-device batch 8
Gradient accumulation 2
Learning rate 1e-6
Warmup ratio 0.1
KL coefficient (β) 0.04
Clip ε 0.2
λ (verifiable penalty) 0.75
Format penalty 0.5
Answer-tag c 0.5
Max completion length 1024
Optimizer DeepSpeed ZeRO-2 + CPU offload
Generation backend vLLM (server mode, port 8001)
Seed 42

Data (in-distribution 90/10 split of train_subset.json)

Split Source Rows
Train train_subset_90pct.json (90% stratified by error_type) 6,105
Eval (during training) train_subset_10pct.json (in-distribution 10% held-out, stratified, seed=42) 678

Performance

Metric Value Step
eval_reward (best, saved as final) 1.2560 5,185
eval_loss (at best) 0.0187 5,185
eval_reward (last) 1.2236 6,100
eval_loss (last) 0.0308 6,100
train reward (last 100-step window) ≈ 1.20

License

Apache-2.0. Base model is governed by the Llama 3.1 Community License — verify compliance for your use case.

Downloads last month
1
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support