qwen2.5-3b-deepscaler-critic-priv0

Pretrained scalar value critic for the short-horizon GRPO-vs-PPO study (SkyRL). Qwen2.5-3B backbone (model.safetensors) plus a scalar value head (value_head.safetensors, value_head_config.json); skyrl_critic_config.json is the SkyRL critic contract (schema 3: action-chain credit, gamma 1, prompt contract). Actor-visible critic (priv0): sees only what the policy sees. Trained offline on DeepScaleR easy10k rollouts (deepscaler_critic_pretrain_priv0_easy10k). Use as CRITIC_PATH for ALGO=ppo with PRIVILEGED_CRITIC=0 in SkyRL/scale/train/examples/deepscaler/run_rl.sh.

Downloads last month
32
Safetensors
Model size
3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JWei05/qwen2.5-3b-deepscaler-critic-priv0

Base model

Qwen/Qwen2.5-3B
Finetuned
(543)
this model