Qwen2.5-VL-7B coordinate-text baseline, checkpoint-4000
ScreenSpot-Pro hit_top1 All-avg = 21.57 (n=1581).
OOM 前最后一个完整 checkpoint。曲线在 2000 之后回落,尚未跑过 step 6000 的预定止损点。
What this run is
The pointer-free reference arm for the 7B SFT suite. Same data and global
batch as the pointer-head arms, but the model is stock Qwen2.5-VL-7B-Instruct
trained to emit pyautogui.click(x=..., y=...) in smart-resized pixels. There
is no pointer head. The visual encoder is trained (unfreeze_all_parameters).
| setting | value |
|---|---|
| learning rate | 1e-6 |
| weight decay | 0.01 |
| warmup | 0.03 |
| schedule | cosine |
| global batch | 64 (4 x 8 x micro-1 x accum-2) |
| save every | 2000 steps |
| planned length | 19613 steps (one epoch) |
Sibling pointer-head peaks on the same data sit near 46-47. This arm is the coordinate-text ceiling, not a pointer-head competitor.
Contents
Consolidated weights, tokenizer, trainer_state.json. DeepSpeed ZeRO-3
optimizer shards are omitted: they are 24G per node and only reloadable on
the original 32-GPU layout.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for stephenfan1101/qwen25vl7b_baseline_checkpoint-4000
Base model
Qwen/Qwen2.5-VL-7B-Instruct