Robert policies

RL policies for Robert, a Microduck (~800 g, ~25 cm biped, 14 XL330 servos). Trained in simulation (mjlab / MuJoCo Warp, PPO). Same 61D obs β†’ 14D action contract at 50 Hz as the official policies, so they hot-swap.

Shoe push

Push a shoe-sized box along the floor with the body. A/B experiment: forward vs backward, identical except the walking direction and the box side.

Forward Backward

Each GIF: left = starting policy, right = after 1500 PPO iterations. Command: straight, 0.2 m/s.

Wide view

Starting point: warm-started from velstand.onnx (Pollen's default walk/stand policy): actor + normalizer copied, critic from scratch, action std 0.25. It already walked; it learned to walk into the box and keep it against its body.

Setup: box 28Γ—10Γ—9 cm, 50–300 g, 20 cm ahead/behind, long side facing, yaw Β±45Β°. Commands 0.1–0.3 m/s. Actor blind to the box; critic sees it. Reward = walking recipe + exp(-(v_box βˆ’ v_robot)Β²/0.1Β²) + box-acceleration penalty (anti-kick). Episode ends if the box gets > 40 cm away. 4096 envs, 1500 iterations, one RTX 5060 Ti (~1.5–2 h).

Results

512 robots, 10 s, straight 0.2 m/s.

Forward before β†’ after Backward before β†’ after
Falls 7.2% β†’ 0% 0% β†’ 0%
Box lost 1.4% β†’ 0% 0% β†’ 0%
Box travel 0.22 β†’ 0.96 m 0.00 (stuck) β†’ 0.99 m
Contact (share of force) feet 71%, hips 20%, head 9% β†’ hips 88%, feet 12% trunk 70%, feet 27% β†’ trunk 92%, feet 6%
Contact force p99 33 (head) β†’ 6 N 13 β†’ 8 N
Leg servo load mean / p99 8 / 42% β†’ 8 / 33% 3 / 13% β†’ 9 / 34%
Heading drift, mean (abs) +4Β° (21Β°) β†’ +19Β° (30Β°) βˆ’3Β° (4Β°) β†’ βˆ’23Β° (45Β°)
Iteration where push reward β‰₯ 1 496 175
  • Both learned a clean body push; the box travels ~90% as far as the robot.
  • Backward learns ~3Γ— faster, although its start policy got stuck against the box. It trained with a bug (mjlab's default rel_forward_envs=0.2 turned 20% of its commands into "walk forward"), so its command obedience is suspect; fixed for later runs.
  • Backward pushes with the rigid trunk (battery); forward with hip brackets hanging from servos. Caveat: the front trunk and thigh servos have no collision geometry in this model, so real forward contacts will differ.
  • Forward keeps the box in the head camera's view; backward doesn't.

Limitations: no heading hold (Β±30–50Β° per 10 s, random sign per robot; steer with ang_vel_z); head-pose commands trained up to Β±80Β° (keep small when pushing); box, not a shoe; sim only.

Files

shoe_push/{forward,backward}/: policy.onnx (obs [1, 61] β†’ [1, 14], normalizer baked in, joint/obs metadata), model_1499.pt (rsl_rl checkpoint for further training), media/.

Credits

Pollen Robotics (Microduck, microduck_rl, base policy; Apache-2.0), mjlab, MuJoCo Warp, rsl_rl.

Downloads last month
4
Video Preview
loading