Qwen2.5-VL-7B GUI grounding — arm bblr5e7, checkpoint-13500
ScreenSpot-Pro hit_top1 All-avg = 47.19 (n=1581, binomial SE ≈ 1.25pp).
全臂最高分。该副本由最佳点守护脚本从原目录复制而来,只含权重与 trainer_state.json。
What this run is
A pointer-head GUI grounding model supervised-finetuned from a warmup checkpoint that trained the pointer head alone (SS-Pro 42.57). The visual encoder stays frozen throughout; only the LLM backbone, the pointer head and the new token embeddings are trained.
The one variable that separates this arm from three siblings is the backbone learning rate, set here to 5e-7. Two arms at 5e-6 peaked around step 2000 and then collapsed to 36-40; an arm at 1e-6 held near 45 but drifted down after step 14000. This arm rose over the full 19613 steps and finished without degrading, so its last checkpoint is directly usable rather than something you have to search for.
| arm | backbone lr | best 3-point mean | last checkpoint |
|---|---|---|---|
| A | 5e-6 | 45.79 | 40.35 @ 13500 |
| B | 5e-6 | 46.39 | 36.50 @ 10500 |
| 3 | 1e-6 | 46.05 | 43.64 @ 17500 |
| 5 (this) | 5e-7 | 46.74 | 46.05 @ 19613 |
Training configuration: 4 nodes x 8 RTX 4090, DeepSpeed ZeRO-3, global batch 64, 19613 steps, visual encoder frozen.
Contents
Weights plus trainer_state.json, which holds the full loss / learning-rate /
gradient-norm history for the run.
The DeepSpeed optimizer state is not included. It is 4 x 23 GB sharded across the four training nodes and can only be reloaded on the identical 32-GPU topology, so it would not be usable to anyone else. Consequently this checkpoint is for inference and evaluation, not for resuming training.
- Downloads last month
- 17
Model tree for stephenfan1101/qwen25vl7b_detr_sft_bblr5e7_checkpoint-13500
Base model
Qwen/Qwen2.5-VL-7B-Instruct