--- license: apache-2.0 base_model: Qwen/Qwen3-VL-8B-Instruct library_name: peft tags: - lora - process-reward-model - multimodal - chart-reasoning - em-prm --- # EM-PRM v2 — `a3_claims_s0` LoRA adapter for **Qwen/Qwen3-VL-8B-Instruct** from the paper *EM-PRM: Evidence-Mediated Process Rewards for Robust Multimodal Reasoning* (EM-PRM v2 experiment ladder). A3 claims-stage checkpoint (see planning/V2_PLAN.md in the dataset mirror for its pre-registration). ## Training - LoRA rank 64, alpha 128, dropout 0.05, target modules down_proj, gate_proj, k_proj, o_proj, q_proj, up_proj, v_proj; vision tower frozen; bfloat16. - Seed 0, learning rate 5e-05, effective batch 2×4, one epoch. - Training data, pair sets and every gate artifact are in the mirror `RESEARCH-EMPRM/emprm-v2` (dataset repo; `results/runs_v2/train/a3_claims_s0/`) and the paper bundle under `backdata/`. ## Load ```python from transformers import AutoModelForImageTextToText, AutoProcessor from peft import PeftModel base = AutoModelForImageTextToText.from_pretrained("Qwen/Qwen3-VL-8B-Instruct", dtype="bfloat16", device_map="cuda") model = PeftModel.from_pretrained(base, "RESEARCH-EMPRM/emprm-v2-a3_claims_s0") processor = AutoProcessor.from_pretrained("Qwen/Qwen3-VL-8B-Instruct") ``` `adapter_config.json` records the local path the adapter was trained from; pass the base model explicitly as above. Scoring prompts (bank extraction, claim extraction, claim support, ranking) are the ones in `work/scripts/eval_bon.py` of the mirror. ## Provenance Trained in the EM-PRM v2 repository; every number quoted in the paper is traceable to `planning/V2_PLAN.md` and the generated tables in the mirror.