dmitchelljackson's picture
Upload Qwen3.5 AndroidWorld RL harness milestone 2026-06-09
b7df156 verified
|
Raw History Blame Contribute Delete
7.26 kB
metadata
language: en
license: apache-2.0
base_model: Qwen/Qwen3.5-0.8B
tags:
  - android
  - ui-automation
  - accessibility
  - vision-language
  - lora
  - dora
  - peft

Cerebellum

Cerebellum is a local Android UI action model. It takes the current screen, accessibility tree, task goal, and recent action history, then emits one compact action code for the next UI step.

The current milestone is a Qwen/Qwen3.5-0.8B adapter trained for constrained history-aware action prediction. It is intended as the supervised checkpoint before moving into a narrow RL curriculum.

Current Milestone

Adapter: checkpoints/qwen35_rl_taskcool_sft268_env5_accum4_20260605/current

Base model: Qwen/Qwen3.5-0.8B

Live AndroidWorld/APK eval: configs/android_world_mixed100_eval_infraclean_20260609.json, 100 mixed cases, 5 emulators, target app preopened, collector APK state source.

Metric Result
Success 51/100 = 51.0%
Infra skips 0
Average steps 11.9

Strong task families: contacts (11/11), audio record (4/4), clock (9/10), WiFi/Bluetooth/system toggles except brightness (16/21 system overall), camera photo (3/3).

Unsolved task families in this checkpoint: Markor/file workflows, calendar delete, browser maze, OsmAnd, camera video, and brightness sliders. Expense* and SimpleSms* were excluded from the current live eval because local AndroidWorld app setup currently leaves those apps in broken first-run states.

Previous SFT Milestone

Adapter: checkpoints/qwen35_history_actions_eosK_aw2_accum4_from_step140/current

Base model: Qwen/Qwen3.5-0.8B

Held-out eval: shards 18,19, n=100, seed 44103

Metric Result
Exact action string 78/100 = 78.0%
Action type 85/100 = 85.0%
Tap/long-press element label 14/21 = 66.7%

Per group:

Group Result
Tap / long press 14/21 = 66.7%
Type text 20/25 = 80.0%
Scroll 16/21 = 76.2%
Wait 28/31 = 90.3%
System 0/2 = 0.0%

Scroll misses were mostly action selection errors, not direction errors:

scroll_failures: {'wrong_action': 4, 'wrong_direction': 1}
scroll_confusion: {'D->T': 1, 'D->W': 1, 'U->T': 1, 'D->U': 1, 'L->W': 1}

Latency from the eval run was p50=0.47s, avg=0.86s, with the first warmup sample causing a max=25.55s.

Input Format

Training and eval use a chat prompt with up to four history frames and one current frame.

Task: {goal}

Step 1:
<image>
Action taken: {history action text}
{optional compact subtree for acted element}

...

Current screen:
<image>
{compact accessibility tree with randomized SoM labels}

Actions: T <label>=tap, P <label>=long-press, K {text}=type text, U/D/L/R=scroll, B=back, H=home, W=wait, F=done, I=impossible.
Next action:

History tap and type actions include an outlined image region plus a compact subtree for the element acted on. Current-screen tappable elements use randomized uppercase bigram labels. The output is constrained at inference:

  • First token is constrained to valid action codes.
  • T and P element labels are constrained to labels visible on the current screen.
  • K text is generated freely until EOS.
  • Single-token actions (U/D/L/R/B/H/W/F/I) terminate immediately.

Output Format

Code Meaning Example
T <label> Tap labelled element T UE
P <label> Long-press labelled element P JD
K {text} Type text into focused field K hello world
U Scroll up U
D Scroll down D
L Scroll left L
R Scroll right R
B Back B
H Home H
W Wait W
F Done F
I Impossible I

Training Setup

Current trainer:

scripts/train_qwen35_history_actions.py

Important settings for the current milestone:

model_name: Qwen/Qwen3.5-0.8B
init_adapter: checkpoints/qwen35_history_actions_eosK_aw2_accum3_from_step210/current
optimizer: paged_adamw_8bit
learning_rate: 1e-4
grad_accum: 4
effective_batch: 48
quota: tp=5,k=2,scroll=2,wait=2,system=1
action_weight: 2.0
margin_weight: 0.25
margin_target: 8.0
history_len: 4
min_history: 4
history image width: 232
current image width: 464
max_raw_tree_chars: 4000
max_seq_len: 2600

The trainer uses focused loss computation on the answer positions rather than materializing full-prompt vocab logits for every token. This made larger effective batches practical on a 12 GB GPU.

Evaluation

Current held-out eval command shape:

docker run --rm --gpus all `
  -v "${PWD}:/workspace" `
  -v "$env:USERPROFILE\.cache\huggingface:/root/.cache/huggingface" `
  -w /workspace cerebellum:torch27-fla bash -lc `
  "python scripts/eval_qwen35_history_actions.py \
    --adapter checkpoints/qwen35_history_actions_eosK_aw2_accum4_from_step140/current \
    --n 100 \
    --seed 44103 \
    --shards 18 19 \
    --quota tp=5,k=2,scroll=2,wait=2,system=1 \
    --hist-width 232 --hist-gutter 16 \
    --curr-width 464 --curr-gutter 32 \
    --max-raw-tree-chars 4000 \
    --max-seq-len 2600 \
    --max-pix-rows 12000 \
    --max-mm-tokens 2600"

Repository Layout

data/
  datasets/android_control_a11y.py   AndroidControl a11y loading, tree compression, history formatting
  som.py                             randomized set-of-mark label pool and rendering
model/
  cerebellum.py                      legacy Gemma model wrapper and helpers
scripts/
  train_qwen35_grounding.py          Qwen grounding trainer
  train_qwen35_history_tap.py        Qwen tap-history trainer
  train_qwen35_history_actions.py    current Qwen all-action history trainer
  eval_qwen35_grounding.py           Qwen grounding eval
  eval_qwen35_history_tap.py         Qwen tap-history eval
  eval_qwen35_history_actions.py     current Qwen all-action eval

Docker Runtime

The cerebellum:latest image contains the CUDA/PyTorch model stack, Android SDK, Android emulator, and a pinned AndroidWorld checkout. Source files are mounted at runtime, so dependency changes need a rebuild but code edits do not.

docker compose build
docker compose run --rm shell

For AndroidWorld/emulator work, KVM must be available. On Windows/WSL:

powershell -ExecutionPolicy Bypass -File scripts\ensure_wsl_kvm.ps1
docker compose up androidworld

See docs/docker_androidworld.md for the full shareable setup and smoke tests.

Loading The Adapter

from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoProcessor

base = AutoModelForImageTextToText.from_pretrained(
    "Qwen/Qwen3.5-0.8B",
    torch_dtype="auto",
    device_map="auto",
)
model = PeftModel.from_pretrained(
    base,
    "dmitchelljackson/cerebellum-qwen35-history-actions-lora",
)
processor = AutoProcessor.from_pretrained("Qwen/Qwen3.5-0.8B")
model.eval()

Next Work

  • Start a narrow RL curriculum from this milestone.
  • Keep constrained decoding for action and visible-label selection.
  • Use short tasks first so wrong taps can be recovered rather than treated as terminal failure.
  • Preserve the current held-out eval as a regression check while tuning RL.