--- tags: - microduck - robotics - reinforcement-learning - onnx - microduck-slot:walk library_name: microduck pipeline_tag: robotics --- # hill-climb-1m5-reference Pollen's reference for the Microduck Arena's 1.5 m Hill Climb: the hill_climb_1m5 challenge of microduck-challenges trained unchanged with its recipe (4096 envs, 3000 iterations, seed 1); checkpoint 1000, its fastest on the Arena. A **perpetual** policy for the [microduck](https://github.com/pollen-robotics/microduck) (61-D observation, 14 actions, 50 Hz). Runs until told otherwise — a gait for the `walk` slot. ## Run it on a robot ```bash sudo robotctl policy load walk pollen-robotics/microduck-hill-climb-1m5-reference ``` The observation normalizer is baked into `policy.onnx`; feed raw observations. `manifest.json` follows schema 2 of the microduck policy manifest (`docs/policy-manifest.md` in the daemon repo). ## Training - **task_id**: `Mjlab-HillClimb1m5-MicroDuck` - **repo**: `https://github.com/pollen-robotics/microduck-challenges.git` - **branch**: `heading-hold` - **commit**: `7a668f11c` - **checkpoint**: `1000` - **seed**: `1` - **base**: `mjlab-microduck 0.1.0 @ 981a279c6` - **started**: `2026-09-28T15:02:36Z` ## Reproduce Same code, same `uv.lock`, same command, same seed. Training it again yields a comparable policy, not the same weights: GPU reinforcement learning is not bit-reproducible across machines. ```bash git clone https://github.com/pollen-robotics/microduck-challenges.git cd microduck-challenges git checkout 7a668f11c uv sync uv run train Mjlab-HillClimb1m5-MicroDuck --env.scene.num-envs 4096 --agent.max-iterations 3000 --agent.seed 1 --agent.logger tensorboard --agent.run-name reference ```