qwen2.5-3b-deepscaler-critic-priv0
Pretrained scalar value critic for the short-horizon GRPO-vs-PPO study (SkyRL). Qwen2.5-3B backbone (model.safetensors) plus a
scalar value head (value_head.safetensors, value_head_config.json); skyrl_critic_config.json is the SkyRL critic contract
(schema 3: action-chain credit, gamma 1, prompt contract). Actor-visible critic (priv0): sees only what the policy sees.
Trained offline on DeepScaleR easy10k rollouts (deepscaler_critic_pretrain_priv0_easy10k). Use as CRITIC_PATH for ALGO=ppo
with PRIVILEGED_CRITIC=0 in SkyRL/scale/train/examples/deepscaler/run_rl.sh.
- Downloads last month
- 32
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for JWei05/qwen2.5-3b-deepscaler-critic-priv0
Base model
Qwen/Qwen2.5-3B