Duplicate repository. The weights here are byte-identical to Qwen3-4B-Critic-SFT, the 4B SFT critic from Steer, Don't Solve. Use that repository; it has the full model card. This copy is kept because the DPO training runs in this organization reference it as their base_model.

qwen3-4b-sft-prm

Qwen3-4B-Instruct-2507 fine-tuned on critic-sft-cwm-qwen. The auto-generated card that used to live here named the wrong training run; the weights were always the CWM + Qwen corpus model.

Downloads last month
1,795
Safetensors
Model size
4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for code-critic-model/qwen3-4b-sft-prm

Finetuned
(2130)
this model
Finetunes
4 models

Paper for code-critic-model/qwen3-4b-sft-prm