YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

InternVL3.5-2B 路 TTRL 路 OpenR1 路 mmupt

Recipe: mmupt (beta 0.01, K 10, T 0.7, cap 2048, lr 1e-6, warmup 0, weight_decay 0.01, max_grad_norm 1.0, bnpo, scale_rewards group, 12 prompts/step = EB 120), same as big-tier InternVL-8B column. TTRL = self-labeling majority vote. best/ = best-by-val (MathVista-150) step 320; endpoint/ = checkpoint-640 (1 epoch); training/ = train.log, trainer_state, best_metric.

Eval protocol for the paper tables: v2 (T=0, 16k, boxed prompt, rule + Qwen2.5-32B judge), endpoint selection. Local source: /weka/scratch/jhu/dssg2026-ext-rghani1/yyang331/mllm-repro-out/mllm-co-grpo-dp/openr1_internvl35_2b_ttrl_mmupt Trained 2026-09-01/02 on JHU a100.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support