YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Co-GRPO heter Qwen2.5-VL-3B x InternVL3.5-2B 路 OpenR1 路 mmupt 路 group A (Qwen side)
Recipe: mmupt (beta 0.01, K 10, T 0.7, cap 2048, lr 1e-6, warmup 0, weight_decay 0.01, max_grad_norm 1.0, bnpo, scale_rewards group, 12 prompts/step = EB 120), same as big-tier InternVL-8B column. Co-learning (both sides mmupt); 640 steps. Record only: the Qwen column uses the old-recipe pairing. best/ = best-by-val (MathVista-150) step 450; endpoint/ = checkpoint-640 (1 epoch); training/ = train.log, trainer_state, best_metric. Note: endpoint/ = checkpoint-640 weights (complete); the step-640 trainer_state was not written because the launcher stopped group A right after group B finished (fixed for later runs); training/trainer_state_checkpoint-600.json is the latest full state. Eval protocol for the paper tables: v2 (T=0, 16k, boxed prompt, rule + Qwen2.5-32B judge), endpoint selection. Local source: /weka/scratch/jhu/dssg2026-ext-rghani1/yyang331/mllm-repro-out/mllm-co-grpo-dp/openr1_heter_qwen25vl3b_x_internvl35_2b_mmuptv2_20260902_110444/model_a Trained 2026-09-02 on JHU a100 (8 GPUs, 4 per model).