GeoGuesser · Qwen3.5-4B · GRPO run 3 (the ablation)

A LoRA adapter for GeoGuessr, trained to answer a question rather than to be the best model: why did run 1 work?

If you want the model that scores, use geoguesser-qwen3.5-4b-grpo (run 1, 0.6445). This one is here so the ablation is reproducible.

What it answers

Run 1's training dynamics looked alarming: entropy collapsed, the reward spread inside each group went to almost nothing, and grad norm peaked above 11. Run 2 was designed to suppress exactly that (scale_rewards="none", beta=0.02, two tasks per optimizer step). It trained cleanly and gained a fifth as much. Run 3 reverted only the first two of those settings.

run 1 run 2 run 3 (this)
scale_rewards group none group
beta 0 0.02 0
tasks per step 1 2 2
action cost scale 1.0 0.2 0.2
median within-group spread 0.016 0.193 0.078
peak grad norm 11.25 0.16 6.77
paired gain over its own base +0.1620 +0.0326 +0.0717

So the instability was not a bug to suppress: it was where most of the learning came from. Run 3 recovers about 44% of run 1's gain by putting it back, which narrows the cause to those two settings plus the action cost, and is the reason the next experiment is a cost sweep.

Best checkpoint here is step 175 at 0.5526 mean-of-4 (its own base arm scored 0.4809). Scores are on the training reward curve, recomputed from raw distance; see the project README.

Training

base Qwen/Qwen3.5-4B
method GRPO (TRL), environment_factory multi-turn tool calling
LoRA r=16, α=32, dropout 0.05, on q/k/v/o_proj
steps 300, two tasks per optimizer step (ACCUM=4), NUM_GENERATIONS=8
turns 12 max · image 448 px
optimiser LR 3e-5, temperature 1.0, beta=0, scale_rewards="group", COST_SCALE=0.2
environment HuggingEnvs/geoguesser-env, over HTTP

Everything else

Imagery is Mapillary, CC BY-SA 4.0.

Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HuggingEnvs/geoguesser-qwen3.5-4b-grpo-v3

Finetuned
Qwen/Qwen3.5-4B
Adapter
(551)
this model

Dataset used to train HuggingEnvs/geoguesser-qwen3.5-4b-grpo-v3

Collection including HuggingEnvs/geoguesser-qwen3.5-4b-grpo-v3

Evaluation results

  • mean-of-4 reward (best checkpoint, step 175) on GeoGuesser eval split (200 held-out tasks)
    self-reported
    0.553