Instructions to use pisdeburra/robert-policies with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Microduck
How to use pisdeburra/robert-policies with Microduck:
# Replace SLOT with the slot specified in the model card (walk, stand, sitstand, ground_pick, kick_left, kick_right, roulade). sudo robotctl policy load SLOT pisdeburra/robert-policies
- Notebooks
- Google Colab
- Kaggle
Robert policies
RL policies for Robert, a Microduck (~800 g, ~25 cm biped, 14 XL330 servos). Trained in simulation (mjlab / MuJoCo Warp, PPO). Same 61D obs β 14D action contract at 50 Hz as the official policies, so they hot-swap.
Shoe push
Push a shoe-sized box along the floor with the body. A/B experiment: forward vs backward, identical except the walking direction and the box side.
Each GIF: left = starting policy, right = after 1500 PPO iterations. Command: straight, 0.2 m/s.
Starting point: warm-started from velstand.onnx (Pollen's default walk/stand policy): actor + normalizer copied, critic from scratch, action std 0.25. It already walked; it learned to walk into the box and keep it against its body.
Setup: box 28Γ10Γ9 cm, 50β300 g, 20 cm ahead/behind, long side facing, yaw Β±45Β°. Commands 0.1β0.3 m/s. Actor blind to the box; critic sees it.
Reward = walking recipe + exp(-(v_box β v_robot)Β²/0.1Β²) + box-acceleration penalty (anti-kick). Episode ends if the box gets > 40 cm away. 4096 envs, 1500 iterations, one RTX 5060 Ti (~1.5β2 h).
Results
512 robots, 10 s, straight 0.2 m/s.
| Forward before β after | Backward before β after | |
|---|---|---|
| Falls | 7.2% β 0% | 0% β 0% |
| Box lost | 1.4% β 0% | 0% β 0% |
| Box travel | 0.22 β 0.96 m | 0.00 (stuck) β 0.99 m |
| Contact (share of force) | feet 71%, hips 20%, head 9% β hips 88%, feet 12% | trunk 70%, feet 27% β trunk 92%, feet 6% |
| Contact force p99 | 33 (head) β 6 N | 13 β 8 N |
| Leg servo load mean / p99 | 8 / 42% β 8 / 33% | 3 / 13% β 9 / 34% |
| Heading drift, mean (abs) | +4Β° (21Β°) β +19Β° (30Β°) | β3Β° (4Β°) β β23Β° (45Β°) |
| Iteration where push reward β₯ 1 | 496 | 175 |
- Both learned a clean body push; the box travels ~90% as far as the robot.
- Backward learns ~3Γ faster, although its start policy got stuck against the box. It trained with a bug (mjlab's default
rel_forward_envs=0.2turned 20% of its commands into "walk forward"), so its command obedience is suspect; fixed for later runs. - Backward pushes with the rigid trunk (battery); forward with hip brackets hanging from servos. Caveat: the front trunk and thigh servos have no collision geometry in this model, so real forward contacts will differ.
- Forward keeps the box in the head camera's view; backward doesn't.
Limitations: no heading hold (Β±30β50Β° per 10 s, random sign per robot; steer with ang_vel_z); head-pose commands trained up to Β±80Β° (keep small when pushing); box, not a shoe; sim only.
Files
shoe_push/{forward,backward}/: policy.onnx (obs [1, 61] β [1, 14], normalizer baked in, joint/obs metadata), model_1499.pt (rsl_rl checkpoint for further training), media/.
Credits
Pollen Robotics (Microduck, microduck_rl, base policy; Apache-2.0), mjlab, MuJoCo Warp, rsl_rl.
- Downloads last month
- 4



