--- language: en license: apache-2.0 base_model: Qwen/Qwen3.5-0.8B tags: - android - ui-automation - accessibility - vision-language - lora - dora - peft --- # Cerebellum Cerebellum is a local Android UI action model. It takes the current screen, accessibility tree, task goal, and recent action history, then emits one compact action code for the next UI step. The current milestone is a Qwen/Qwen3.5-0.8B adapter trained for constrained history-aware action prediction. It is intended as the supervised checkpoint before moving into a narrow RL curriculum. ## Current Milestone **Adapter:** `checkpoints/qwen35_rl_taskcool_sft268_env5_accum4_20260605/current` **Base model:** `Qwen/Qwen3.5-0.8B` **Live AndroidWorld/APK eval:** `configs/android_world_mixed100_eval_infraclean_20260609.json`, 100 mixed cases, 5 emulators, target app preopened, collector APK state source. | Metric | Result | | --- | ---: | | Success | 51/100 = 51.0% | | Infra skips | 0 | | Average steps | 11.9 | Strong task families: contacts (`11/11`), audio record (`4/4`), clock (`9/10`), WiFi/Bluetooth/system toggles except brightness (`16/21` system overall), camera photo (`3/3`). Unsolved task families in this checkpoint: Markor/file workflows, calendar delete, browser maze, OsmAnd, camera video, and brightness sliders. `Expense*` and `SimpleSms*` were excluded from the current live eval because local AndroidWorld app setup currently leaves those apps in broken first-run states. ## Previous SFT Milestone **Adapter:** `checkpoints/qwen35_history_actions_eosK_aw2_accum4_from_step140/current` **Base model:** `Qwen/Qwen3.5-0.8B` **Held-out eval:** shards `18,19`, `n=100`, seed `44103` | Metric | Result | | --- | ---: | | Exact action string | 78/100 = 78.0% | | Action type | 85/100 = 85.0% | | Tap/long-press element label | 14/21 = 66.7% | Per group: | Group | Result | | --- | ---: | | Tap / long press | 14/21 = 66.7% | | Type text | 20/25 = 80.0% | | Scroll | 16/21 = 76.2% | | Wait | 28/31 = 90.3% | | System | 0/2 = 0.0% | Scroll misses were mostly action selection errors, not direction errors: ```text scroll_failures: {'wrong_action': 4, 'wrong_direction': 1} scroll_confusion: {'D->T': 1, 'D->W': 1, 'U->T': 1, 'D->U': 1, 'L->W': 1} ``` Latency from the eval run was `p50=0.47s`, `avg=0.86s`, with the first warmup sample causing a `max=25.55s`. ## Input Format Training and eval use a chat prompt with up to four history frames and one current frame. ```text Task: {goal} Step 1: Action taken: {history action text} {optional compact subtree for acted element} ... Current screen: {compact accessibility tree with randomized SoM labels} Actions: T