--- base_model: Qwen/Qwen3.5-0.8B library_name: peft pipeline_tag: image-text-to-text license: apache-2.0 tags: - android - android-world - ui-automation - accessibility - vision-language - peft - lora - reinforcement-learning --- # Cerebellum Qwen3.5 AndroidWorld RL Adapter This is a PEFT adapter for `Qwen/Qwen3.5-0.8B` trained for fast local Android UI control. ## What this is for This model is not meant to replace a remote frontier computer-use agent. It is meant to sit underneath one, as a small, fast, cheap execution layer between a remote frontier model and an Android computer-use harness. The intended split is: - The **frontier model** does the higher-level reasoning and planning: interpreting the task, deciding what should happen next, recovering when something goes wrong. - **This 0.8B model**, running locally, executes the low-level Android UI interaction the frontier model delegates to it -- find and tap the right element, type into the right field, scroll the right list. No screenshot round-trip to a remote model is needed for those steps. The hypothesis being tested is that a large fraction of Android UI interaction does not actually require frontier-model intelligence, so delegating it to a specialized local model can cut inference latency and cost while reserving the frontier model for the parts where its capability is doing real work. Whether that hypothesis holds is what the evaluation below is trying to measure; it is not settled. The whole project is developed and trained locally on constrained consumer hardware -- a single RTX 3060 with 12GB of VRAM. That constraint is deliberate rather than incidental: the point is to find out how far a small, cheap, locally trainable model gets on this job, and a model that needs a datacenter to train is not the thing being investigated. ## Interface It consumes a task goal, the current screenshot with randomized set-of-mark labels, the compact accessibility tree, and up to four recent history frames/actions. It emits one compact next-action command. The intended runtime is the Cerebellum AndroidWorld/APK harness included in this upload. The model is not a general chat model and should be decoded with the constrained action grammar used by `rl_harness.policy_qwen35`. The pre-RL supervised checkpoint is preserved in this same repository under `sft_freeze/`. The root adapter is the RL milestone; load `sft_freeze/` directly when a clean SFT fallback is needed. ## Status This checkpoint is a milestone in active work, not a finished result. The 51% held-out AndroidWorld number below is where this particular training run landed, and several whole task families are still at zero. A second iteration is underway, built on a stronger SFT dataset and a revised approach to the online-learning stage; the numbers here should be read as a snapshot of the first iteration. ## Action Grammar ```text T