File size: 2,243 Bytes
59dd36c
 
 
 
 
 
be91970
 
 
59dd36c
 
 
 
fade060
59dd36c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fade060
16bb2aa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
---
tags:
- microduck
- robotics
- reinforcement-learning
- onnx
- microduck-slot:walk
library_name: microduck
pipeline_tag: robotics
---

# walk-flat

v2: the v1 flat-ground walk fine-tuned with +-1 deg of gear play in every servo (Mjlab-Velocity-Flat-Backlash-MicroDuck), simulation only: v1's 4,000 iterations plus 3,000 with the slop, 4096 envs, 57 min on one RTX 5090. With the slop it stayed on its feet for 20 s in 4 of 5 takes; v1 did 2 of 5 in the same sim (small sample, the training fall rate did not improve). Not yet tested on a real Microduck.

A **perpetual** policy for the [microduck](https://github.com/pollen-robotics/microduck) (61-D observation, 14 actions, 50 Hz). Runs until told otherwise — a gait for the `walk` slot.

## Run it on a robot

```bash
sudo robotctl policy load walk witcheer/microduck-walk-flat
```

The observation normalizer is baked into `policy.onnx`; feed raw observations.
`manifest.json` follows schema 2 of the microduck policy manifest (`docs/policy-manifest.md` in the daemon repo).

## Training

- **repo**: `pollen-robotics/microduck_rl`
- **branch**: `develop`
- **commit**: `53b8971b6`
- exported from a checkout with uncommitted changes

## Results in simulation (v2, 2026-09-26)

Proof takes with `headless_play` in `Mjlab-Velocity-Flat-Backlash-MicroDuck` (±1° of gear play in every servo), 20 s each, standing start, trunk height and tilt sampled every 0.1 s. A fall is a mid-episode environment termination (the robot tipping over).

- v2: 4 of 5 takes with no fall; the fall-free takes kept the trunk between 104 and 135 mm. One take tipped sideways at 5.8 s.
- v1 in the same gear-play sim: 2 of 5 takes with no fall.
- v1 in its original sim without gear play: 1 of 3.
- Training, last 100 iterations: fell_over 0.496 for v2 against 0.453 for v1, so the take gap is a small-sample hint, not a measured gain.

The training checkout had one local change (the task registration for an unrelated get-up task in `tasks/__init__.py`); the backlash task itself is unmodified at the commit above.

## v1

The first release: 4,000 iterations on `Mjlab-Velocity-Flat-MicroDuck`, no gear play. Still installable:

```bash
sudo robotctl policy load walk witcheer/microduck-walk-flat@v1
```