YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Co-GRPO heter Qwen2.5-VL-3B x InternVL3.5-2B 路 OpenR1 路 old recipe 路 group B (InternVL side)
Recipe: old (beta 0, K 8, T 1.0, cap 1024, lr 1e-6, warmup 0.03, 8 prompts/step = EB 64), same as big-tier Qwen-7B column. Co-learning; 961 steps. best/ = best-by-val (MathVista-150) step 500; endpoint/ = checkpoint-961 (1 epoch); training/ = train.log, trainer_state, best_metric.
Eval protocol for the paper tables: v2 (T=0, 16k, boxed prompt, rule + Qwen2.5-32B judge), endpoint selection. Local source: /weka/scratch/jhu/dssg2026-ext-rghani1/yyang331/mllm-repro-out/mllm-co-grpo-dp/openr1_heter_qwen25vl3b_x_internvl35_2b_oldv2_20260901_213002/model_b Trained 2026-09-01/02 on JHU a100.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support