--- tags: - microduck - robotics - reinforcement-learning - onnx library_name: microduck pipeline_tag: robotics --- # groove In-place groove: rhythmic knee bounce with a small head sway (sim-trained, untested on hardware) A **episodic** policy for the [microduck](https://github.com/pollen-robotics/microduck) (61-D observation, 14 actions, 50 Hz). Runs 10.0 s and returns itself to a standing pose. ## Run it on a robot ```bash sudo robotctl policy add groove kramp/microduck-groove robotctl robot do groove ``` The observation normalizer is baked into `policy.onnx`; feed raw observations. `manifest.json` follows schema 2 of the microduck policy manifest (`docs/policy-manifest.md` in the daemon repo). ## Training - **task_id**: `Mjlab-Dance-Flat-MicroDuck` - **repo**: `pollen-robotics/microduck_rl` - **branch**: `dance` - **commit**: `2f3f4ea1d` - **checkpoint**: `3999` ## What it actually does (measured in simulation) 20 s headless rollouts in CPU MuJoCo with the BAM XL330 actuator model (`scripts/render_policy.py`), starting from the standing pose, all command slots at zero: - **Never falls**: max trunk tilt 1.6°, drift 1 cm. - **Knee bounce**: hip/knee/ankle bounce in sync on both legs, at 87% of the reference amplitude, about 2.1 Hz (≈126 BPM rather than the 100 BPM target). - **Head**: a small head-yaw sway at 30% of the reference amplitude. The planned head nod, head tilt and hip sway were **not** learned. `replay.mp4` shows it. ## How it was trained - Task `Mjlab-Dance-Flat-MicroDuck` (mjlab, PPO), built on the velocity recipe: domain randomization, pushes, and BAM actuators, with the twist command always zero. - The actor gets no clock (the standard 61-D contract), so rhythm has to come from state feedback. The reward runs a phase-locked reference: it advances at the target tempo and may slip only a little, so standing still never pays. - Reference-state initialization: 70% of episodes start mid-dance. - 4000 iterations × 4096 envs on one L4 (HF Jobs, about 2 h 50 min). - The training code is a local fork of `pollen-robotics/microduck_rl` (branch `dance`, not published), so the commit hash below does not exist upstream. ## Caveats - **Sim-only. It has never run on a real robot.** Test it held or on a soft surface first. - The policy grooves indefinitely. When `duration_s` (10 s) ends, the daemon hands back to the stand policy mid-bounce rather than from a settled pose.