AdithyaSK's picture
AdithyaSK HF Staff
Write a real model card: scores, training config, behaviour and links
81a179c verified
|
Raw
History Blame Contribute Delete
3.24 kB
---
base_model: Qwen/Qwen3.5-4B
library_name: peft
license: apache-2.0
pipeline_tag: image-text-to-text
tags:
- openenv
- trl
- grpo
- reinforcement-learning
- lora
- visual-geolocation
- geoguessr
- ablation
datasets:
- HuggingEnvs/geoguesser-tasks
model-index:
- name: geoguesser-qwen3.5-4b-grpo-v3
results:
- task:
type: image-text-to-text
name: Visual geolocation
dataset:
name: GeoGuesser eval split (200 held-out tasks)
type: HuggingEnvs/geoguesser-tasks
metrics:
- type: reward
value: 0.5526
name: mean-of-4 reward (best checkpoint, step 175)
---
# GeoGuesser 路 Qwen3.5-4B 路 GRPO run 3 (the ablation)
A LoRA adapter for [GeoGuessr](https://huggingface.co/spaces/HuggingEnvs/geoguesser-env), trained to
answer a question rather than to be the best model: **why did run 1 work?**
If you want the model that scores, use
**[`geoguesser-qwen3.5-4b-grpo`](https://huggingface.co/HuggingEnvs/geoguesser-qwen3.5-4b-grpo)**
(run 1, 0.6445). This one is here so the ablation is reproducible.
## What it answers
Run 1's training dynamics looked alarming: entropy collapsed, the reward spread inside each group
went to almost nothing, and grad norm peaked above 11. Run 2 was designed to suppress exactly that
(`scale_rewards="none"`, `beta=0.02`, two tasks per optimizer step). It trained cleanly and gained a
fifth as much. Run 3 reverted only the first two of those settings.
| | run 1 | run 2 | run 3 (this) |
|---|---:|---:|---:|
| `scale_rewards` | `group` | `none` | `group` |
| `beta` | 0 | 0.02 | 0 |
| tasks per step | 1 | 2 | 2 |
| action cost scale | 1.0 | 0.2 | 0.2 |
| median within-group spread | 0.016 | 0.193 | 0.078 |
| peak grad norm | 11.25 | 0.16 | 6.77 |
| **paired gain over its own base** | **+0.1620** | +0.0326 | **+0.0717** |
So the instability was not a bug to suppress: it was where most of the learning came from. Run 3
recovers about 44% of run 1's gain by putting it back, which narrows the cause to those two settings
plus the action cost, and is the reason the next experiment is a cost sweep.
Best checkpoint here is step 175 at **0.5526** mean-of-4 (its own base arm scored 0.4809). Scores
are on the training reward curve, recomputed from raw distance; see the project README.
## Training
| | |
|---|---|
| base | `Qwen/Qwen3.5-4B` |
| method | GRPO (TRL), `environment_factory` multi-turn tool calling |
| LoRA | r=16, 伪=32, dropout 0.05, on `q/k/v/o_proj` |
| steps | 300, two tasks per optimizer step (`ACCUM=4`), `NUM_GENERATIONS=8` |
| turns | 12 max 路 image 448 px |
| optimiser | LR 3e-5, temperature 1.0, `beta=0`, `scale_rewards="group"`, `COST_SCALE=0.2` |
| environment | [`HuggingEnvs/geoguesser-env`](https://huggingface.co/spaces/HuggingEnvs/geoguesser-env), over HTTP |
## Everything else
- **[The write-up](https://huggingface.co/spaces/HuggingEnvs/geoguesser-article)**, including why these two settings mattered so much
- **[All four runs](https://huggingface.co/spaces/HuggingEnvs/geoguesser-trackio)** on one axis
- **[Code and exact commands](https://github.com/adithya-s-k/HuggingEnvs/tree/main/03-geoguesser)**
Imagery is Mapillary, CC BY-SA 4.0.