Instructions to use RESEARCH-EMPRM/emprm-v2-stageB_rung5_noev_s0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use RESEARCH-EMPRM/emprm-v2-stageB_rung5_noev_s0 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("/home/ingyu/.cache/huggingface/hub/models--Qwen--Qwen3-VL-8B-Instruct/snapshots/0c351dd01ed87e9c1b53cbc748cba10e6187ff3b") model = PeftModel.from_pretrained(base_model, "RESEARCH-EMPRM/emprm-v2-stageB_rung5_noev_s0") - Notebooks
- Google Colab
- Kaggle
EM-PRM v2 β stageB_rung5_noev_s0 β R+G arm (E2 ablation), seed 0
LoRA adapter for Qwen/Qwen3-VL-8B-Instruct (snapshot 0c351dd) from the EM-PRM v2 experiment ladder (EM-PRM: Evidence-Mediated Process Rewards for Robust Multimodal Reasoning). Trained 2026-09-10 at git commit 4438aea.
The rung-5 recipe with the evidence line removed from the ranker prompt at training and inference (GPRM_RANKER_NO_EVIDENCE=1): same 20,000 task records, the same 8,353 pairs rendered without the evidence line, same A2 initialisation, same schedule, pair_sees_image: true. Its product pass is the R+G cell of the E2 factorial (record gate x evidence-free ranker); its verify_only pass is the R cell.
Training
- LoRA rank 64, alpha 128, dropout 0.05, target modules down_proj, gate_proj, k_proj, o_proj, q_proj, up_proj, v_proj (174,587,904 trainable parameters); vision tower frozen; bfloat16.
- Seed 0, learning rate 5e-05, micro-batch 2 x grad_accum 4, 1 epoch, 2,500 optimiser steps over 20,000 task records; pair records 8,353, pair micro-batches 5,000, lambda_pair 1.0, pair_sees_image True.
- Initialised from
a2_support_s0(RESEARCH-EMPRM/emprm-v2-a2_support_s0). Task-data sha256ddd7be2d3ddfβ¦, pair-dataruns/v2/data/stageB_pairs_rung5_noev/pairs.jsonl. - Wall time 5.4 h on one NVIDIA A100-PCIE-40GB; training-pair accuracy mean 0.8504, final 0.98.
Pre-registered gates and reads (development halves; test halves unread)
- Held-out relational FlipAcc (operation, deployed text-only pass): 0.0000 [0.0000, 0.0000] β gate (>= 0.74) lost; forced by the architecture (no evidence line and no image in the ranker), verified on 800/800 pairs. With the ranker shown the chart: 0.0037 [0.0000, 0.0088].
- Forced-evidence acceptance at 0.5 (legend binding, 400 per cell): true 0.0275 / false 0.000 deployed; 0.920 / 0.005 with the chart shown (true gate >= 0.95 lost in both passes).
- Chart-disjoint pair gain over the v1 head: +0.1330 [0.0874, 0.1787] deployed; +0.1139 [0.0733, 0.1545] with the chart shown.
- Controlled premise-adversarial pools (dev, Best-of-5, InternVL3.5-8B / Qwen3-VL-8B): 0.3273 / 0.4318 against 0.8369 / 0.8475 for the final arm.
- External dev halves (deployed pass): VisualProcessBench macro-F1@0.5 0.3894, VLRMBench 0.3757, VL-RewardBench 0.5365, Multimodal RewardBench 0.5056 (no target met).
Status
Not a deployment candidate (chart gates lost). Kept, with the two other seeds, as the matched "E removed" ablation of the final EM-PRM.
Where the artifacts are
- Result files, per-example dumps, config and prompts: dataset
RESEARCH-EMPRM/emprm-sync-20260910βresults/**/runs/v2/train/stageB_rung5_noev_s0/,results/**/runs/v2/e2/stageB_rung5_noev_s0__*.json,configs/ablations/stageB_rung5_noev_s0.yaml,EXPERIMENT_REGISTRY.csv(rows tagged with this adapter),CURRENT.mdandWRITER_SYNC_BUNDLE.md(what the deployed scorer computes; which arm is which). - The frozen 2026-09-09 tree backup
RESEARCH-EMPRM/emprm-v2predates this arm and does not contain it.
Load
from transformers import AutoModelForImageTextToText, AutoProcessor
from peft import PeftModel
base = AutoModelForImageTextToText.from_pretrained("Qwen/Qwen3-VL-8B-Instruct", dtype="bfloat16", device_map="cuda")
model = PeftModel.from_pretrained(base, "RESEARCH-EMPRM/emprm-v2-stageB_rung5_noev_s0")
processor = AutoProcessor.from_pretrained("Qwen/Qwen3-VL-8B-Instruct")
adapter_config.json records the local path the adapter was trained from; pass the base model explicitly as above. The scorer (scoring.Scorer.score_grounded, family grounded, aggregation product) and its prompts are in code/ of the sync dataset.
- Downloads last month
- 11
Model tree for RESEARCH-EMPRM/emprm-v2-stageB_rung5_noev_s0
Base model
Qwen/Qwen3-VL-8B-Instruct