HannesVonEssen commited on
Commit
108df8b
·
verified ·
1 Parent(s): 9153205

Prepare public jump/backflip release with public source links and fresh behavior evidence

Browse files
README.md CHANGED
@@ -36,7 +36,7 @@ uses the same collision box as the visible mat.
36
 
37
  ## Reproduce in simulation
38
 
39
- [Source and instructions](https://github.com/Vottivott/microduck-secret-playground/tree/release/parkour/experiments/parkour/backflip)
40
 
41
  ```bash
42
  uv sync --locked
@@ -52,9 +52,12 @@ parity and a small fresh diagnostic of this checkpoint, `manifest.json` for
52
  provenance and `SHA256SUMS` for integrity. Load the pickle checkpoint only if
53
  you trust its source.
54
 
55
- Related: [long jump](https://huggingface.co/HannesVonEssen/microduck-long-jump)
56
- · [pillar jumps](https://huggingface.co/HannesVonEssen/microduck-parkour).
57
 
58
  ## Continue training
59
 
60
  [TRAINING.md](TRAINING.md) documents the full-checkpoint continuation recipe, distinct from render profiles. A clean locked installation passed a 64-environment/five-update resume, exact learning-state restoration, finite-step checks and normalized ONNX parity. Recipe provenance and reconstruction limits are explicit; this is not a claim of bit-identical historical replay or improved behavior. Original policy, checkpoint and media are unchanged.
 
 
 
 
 
36
 
37
  ## Reproduce in simulation
38
 
39
+ [Source and instructions](https://github.com/Vottivott/microduck-playground/tree/b6aed572587b52e41a452a9b2be517c95d2b3d96/experiments/parkour/backflip)
40
 
41
  ```bash
42
  uv sync --locked
 
52
  provenance and `SHA256SUMS` for integrity. Load the pickle checkpoint only if
53
  you trust its source.
54
 
55
+ Related: [long jump](https://huggingface.co/HannesVonEssen/microduck-long-jump).
 
56
 
57
  ## Continue training
58
 
59
  [TRAINING.md](TRAINING.md) documents the full-checkpoint continuation recipe, distinct from render profiles. A clean locked installation passed a 64-environment/five-update resume, exact learning-state restoration, finite-step checks and normalized ONNX parity. Recipe provenance and reconstruction limits are explicit; this is not a claim of bit-identical historical replay or improved behavior. Original policy, checkpoint and media are unchanged.
60
+
61
+ ## Fresh release check — 2026-09-29
62
+
63
+ Completes the backward rotation and recovers to standing on the simulated mat, but later drifts off and falls at 7.62 seconds. **This is not a stable idle policy.** One CPU rollout at the documented default seed (28), using the exact packaged checkpoint; not a success-rate estimate or hardware validation. See [source verification](https://github.com/Vottivott/microduck-playground/blob/b6aed572587b52e41a452a9b2be517c95d2b3d96/experiments/parkour/VERIFICATION.md) and `eval/render-20260929.json`.
SHA256SUMS CHANGED
@@ -1,20 +1,21 @@
1
  2d3fd9ffffa70a4307498ec630b14a6b4731e010f8ddda92686eaf65733c2dfe LICENSE
2
- 0e710c240920284a1afae0ef449f475ed165354db46f688b9e7ce2492cf0feac README.md
3
- 1202ec927bd5828316461e58e1e66dbc3d749e833b0e659a203334ae86e4dfb9 TRAINING.md
4
  f7e523c19de5cd5e2ced430f3d9bb46bef1ed69ad0003cb45955d56169929cdd checkpoint.pt
5
- 0f5574c1b044ece6df04567e8d469690cb9e83a46ddb96eeae3ef5c75c20e9e3 config.json
6
  e5b5bfb4e5fe8320584a20d12d729081a8c29980919be9540e8256d0a3d0fee5 eval/backflip.json
7
  d2478a39f6d03018d830be78c9d72a7322e755932f21e0c307305220ab890396 eval/onnx-parity.json
 
8
  a0c92f27287602cd7983d16957890584ce8b9d323a10ffc1aad1b1a60283a52a eval/summary.json
9
- cf3b8d9a65d079142c218c1c4c597280805e22b0484dfe0346aa66e1dc53fd0b manifest.json
10
  e0b3e1e17f68151dff05dc019999c0053bf6820cad6e44b2e80e305d799a36bd media/preview.mp4
11
  841e4eebfbe9946145f88cf01ce3f8f8a501b6b9312b0f037f2b4e0a5f91a48c media/social-preview.jpg
12
  b4d780ee2a2bf7aab5b33358124fd079b17a329412d496574296919dbc5b6ee0 media/social-preview.png
13
  59dd61ea1e8566a3bad9f8b24dd18530664906c33964b1c05e6ed6e821e16a24 policy.onnx
14
  880cbe85329c8dd1524997eb9049afaec361ad90514ba71e18290f97db22e4d8 training/provenance/agent.yaml
15
  ba13ebeb26b8437a862970aceec88f5370d94eb493c1bc6594716ebde3141085 training/provenance/env.yaml
16
- 2fc78f00e9199c2b206b6bcd8d1a7f7bbb95012176b55c9bdddec64036c679b8 training/resume.json
17
- bc414ea910af3f0e0e623af26a1c1db51e28f5f8c2469a1e6c0b42a60de1d145 training/source.json
18
  d013ef788d7789af07f3fce4d2344c25c074f0a2b187367881fba6461fd12d38 training/validation/params/agent.yaml
19
  c3de77796cdac808fce58bdfce75c2a9c8c8d4ff5bac438a1c8a3021e33f7503 training/validation/params/env.yaml
20
  e51de08bb6c0cfc8b87ce386b1451a94f504d4d574779446ddd277a7cf7a76d8 training/validation/params/events-after-restore.yaml
 
1
  2d3fd9ffffa70a4307498ec630b14a6b4731e010f8ddda92686eaf65733c2dfe LICENSE
2
+ 2a64ad6f7d86cb958cd3430f59dec3afed8dbf9fb77f8ede0ab675ea9d8bc77b README.md
3
+ f75e93c537ac4368d947ca3335f94cbe7cefc07fd04900e4e9b2dfcbf59c2573 TRAINING.md
4
  f7e523c19de5cd5e2ced430f3d9bb46bef1ed69ad0003cb45955d56169929cdd checkpoint.pt
5
+ b4ef411ee0c1357763390596beb8cd8c4cd64d0cbe98c11b9d7daf8e0371727b config.json
6
  e5b5bfb4e5fe8320584a20d12d729081a8c29980919be9540e8256d0a3d0fee5 eval/backflip.json
7
  d2478a39f6d03018d830be78c9d72a7322e755932f21e0c307305220ab890396 eval/onnx-parity.json
8
+ 8e3b3fc4530883b403c5f6bec48371e04985b6a3c3957a02bcedd99896c38ba3 eval/render-20260929.json
9
  a0c92f27287602cd7983d16957890584ce8b9d323a10ffc1aad1b1a60283a52a eval/summary.json
10
+ c3de8687b897e3dc50944c76d72f30c1597937bb81d1aab0563d9ef9f4cbeca2 manifest.json
11
  e0b3e1e17f68151dff05dc019999c0053bf6820cad6e44b2e80e305d799a36bd media/preview.mp4
12
  841e4eebfbe9946145f88cf01ce3f8f8a501b6b9312b0f037f2b4e0a5f91a48c media/social-preview.jpg
13
  b4d780ee2a2bf7aab5b33358124fd079b17a329412d496574296919dbc5b6ee0 media/social-preview.png
14
  59dd61ea1e8566a3bad9f8b24dd18530664906c33964b1c05e6ed6e821e16a24 policy.onnx
15
  880cbe85329c8dd1524997eb9049afaec361ad90514ba71e18290f97db22e4d8 training/provenance/agent.yaml
16
  ba13ebeb26b8437a862970aceec88f5370d94eb493c1bc6594716ebde3141085 training/provenance/env.yaml
17
+ 309e9211d115a9d6b3f7ed58e85135538e7e7f165570003d52583dbe5e7c3b57 training/resume.json
18
+ cf4a426315f06e2db9908078a8a1b63a50ec029bf42c3835956017e1ea3f7556 training/source.json
19
  d013ef788d7789af07f3fce4d2344c25c074f0a2b187367881fba6461fd12d38 training/validation/params/agent.yaml
20
  c3de77796cdac808fce58bdfce75c2a9c8c8d4ff5bac438a1c8a3021e33f7503 training/validation/params/env.yaml
21
  e51de08bb6c0cfc8b87ce386b1451a94f504d4d574779446ddd277a7cf7a76d8 training/validation/params/events-after-restore.yaml
TRAINING.md CHANGED
@@ -1,6 +1,6 @@
1
  # Continue the released parkour policies
2
 
3
- All three HF repos contain original full PPO checkpoints: actor, critic,
4
  observation normalizers, populated Adam state and environment step counter.
5
  Use the source revision in each repo's `training/source.json`. Install with
6
  `uv sync --locked` on Linux/CUDA. MuJoCo also needs system EGL or OSMesa libraries;
@@ -10,12 +10,10 @@ Only load trusted checkpoints.
10
  | Profile | HF repo | Stored iteration |
11
  | --- | --- | ---: |
12
  | `long-jump` | `HannesVonEssen/microduck-long-jump` | 15,250 |
13
- | `pillar-jumps` | `HannesVonEssen/microduck-parkour` | 34,000 |
14
  | `backflip` | `HannesVonEssen/microduck-backflip` | 12,000 |
15
 
16
  ```bash
17
  hf download HannesVonEssen/microduck-long-jump checkpoint.pt --local-dir policies/long-jump
18
- hf download HannesVonEssen/microduck-parkour checkpoint.pt --local-dir policies/pillar-jumps
19
  hf download HannesVonEssen/microduck-backflip checkpoint.pt --local-dir policies/backflip
20
 
21
  # From the source root, without inherited MICRODUCK_* environment settings:
@@ -23,10 +21,6 @@ uv run python scripts/resume_release.py \
23
  --recipe experiments/parkour/resume.json --profile long-jump \
24
  --checkpoint policies/long-jump/checkpoint.pt \
25
  --output logs/long-jump-smoke --num-envs 64 --iterations 5
26
- uv run python scripts/resume_release.py \
27
- --recipe experiments/parkour/resume.json --profile pillar-jumps \
28
- --checkpoint policies/pillar-jumps/checkpoint.pt \
29
- --output logs/pillar-jumps-smoke --num-envs 64 --iterations 5
30
  uv run python scripts/resume_release.py \
31
  --recipe experiments/parkour/resume.json --profile backflip \
32
  --checkpoint policies/backflip/checkpoint.pt \
@@ -65,14 +59,6 @@ performance, robustness or hardware-safety result.
65
  resumes at the saved counter (late range: gap 0.15–0.40 m, drop 0.15–0.35 m).
66
  This source includes later research changes, so it is a supported continuation
67
  rather than an exact reconstruction of p1's original implementation.
68
- - **Pillar jumps:** the c29 recipe is reconstructed, not a recovered launch.
69
- It explicitly enables **informed course observations** (`MICRODUCK_PK_BLIND=0`),
70
- five pillars, 12 cm gap, 14 cm drop, and a 75%-of-standing-height hold gate.
71
- Seed42 and remaining pinned values are documented continuation choices.
72
- The bare task defaults to blind observations and a disabled height gate;
73
- using it unconfigured would silently change the problem. Fixed gap/drop
74
- disable per-environment difficulty promotion, whose old state was not saved.
75
- Nine pillars remain a transfer/evaluation profile, not another checkpoint.
76
  - **Backflip:** f21 launcher and recorded YAML were recovered: 0.7–0.9 m drop,
77
  entropy0.003, takeoff weight45, minimum takeoff vz0, curriculum shift3000.
78
  The initial midflip probability0.25 becomes **0.15** at the restored mature
@@ -91,4 +77,4 @@ command slots and action bounds, not reward totals alone. The original montage's
91
  exact checkpoint/seed mapping remains unverified; training documentation does
92
  not change that limitation.
93
 
94
- Pinned source: [release/parkour@58d56df0](https://github.com/Vottivott/microduck-secret-playground/tree/58d56df0e863df552fac7caf409aaa968cfde34d).
 
1
  # Continue the released parkour policies
2
 
3
+ Both HF repos contain original full PPO checkpoints: actor, critic,
4
  observation normalizers, populated Adam state and environment step counter.
5
  Use the source revision in each repo's `training/source.json`. Install with
6
  `uv sync --locked` on Linux/CUDA. MuJoCo also needs system EGL or OSMesa libraries;
 
10
  | Profile | HF repo | Stored iteration |
11
  | --- | --- | ---: |
12
  | `long-jump` | `HannesVonEssen/microduck-long-jump` | 15,250 |
 
13
  | `backflip` | `HannesVonEssen/microduck-backflip` | 12,000 |
14
 
15
  ```bash
16
  hf download HannesVonEssen/microduck-long-jump checkpoint.pt --local-dir policies/long-jump
 
17
  hf download HannesVonEssen/microduck-backflip checkpoint.pt --local-dir policies/backflip
18
 
19
  # From the source root, without inherited MICRODUCK_* environment settings:
 
21
  --recipe experiments/parkour/resume.json --profile long-jump \
22
  --checkpoint policies/long-jump/checkpoint.pt \
23
  --output logs/long-jump-smoke --num-envs 64 --iterations 5
 
 
 
 
24
  uv run python scripts/resume_release.py \
25
  --recipe experiments/parkour/resume.json --profile backflip \
26
  --checkpoint policies/backflip/checkpoint.pt \
 
59
  resumes at the saved counter (late range: gap 0.15–0.40 m, drop 0.15–0.35 m).
60
  This source includes later research changes, so it is a supported continuation
61
  rather than an exact reconstruction of p1's original implementation.
 
 
 
 
 
 
 
 
62
  - **Backflip:** f21 launcher and recorded YAML were recovered: 0.7–0.9 m drop,
63
  entropy0.003, takeoff weight45, minimum takeoff vz0, curriculum shift3000.
64
  The initial midflip probability0.25 becomes **0.15** at the restored mature
 
77
  exact checkpoint/seed mapping remains unverified; training documentation does
78
  not change that limitation.
79
 
80
+ Pinned public source: [microduck-playground@b6aed572](https://github.com/Vottivott/microduck-playground/tree/b6aed572587b52e41a452a9b2be517c95d2b3d96).
config.json CHANGED
@@ -87,9 +87,9 @@
87
  "format": "onnx",
88
  "policy_file": "policy.onnx",
89
  "checkpoint_file": "checkpoint.pt",
90
- "source_repository": "Vottivott/microduck-secret-playground",
91
- "source_branch": "release/parkour",
92
- "source_revision": "26557d0a59537ce9d99f8fe1ca8f7e9bd8ebb6cb",
93
  "observation_layout": {
94
  "base_ang_vel": [
95
  0,
 
87
  "format": "onnx",
88
  "policy_file": "policy.onnx",
89
  "checkpoint_file": "checkpoint.pt",
90
+ "source_repository": "Vottivott/microduck-playground",
91
+ "source_branch": "main",
92
+ "source_revision": "b6aed572587b52e41a452a9b2be517c95d2b3d96",
93
  "observation_layout": {
94
  "base_ang_vel": [
95
  0,
eval/render-20260929.json ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "task": "Mjlab-Flip-MicroDuck",
3
+ "profile": "backflip",
4
+ "seed": 28,
5
+ "num_envs": 1,
6
+ "duration_s": 12.0,
7
+ "first_episode_end_events": [
8
+ {
9
+ "env": 0,
10
+ "time_s": 7.62,
11
+ "terms": [
12
+ "fell"
13
+ ]
14
+ }
15
+ ],
16
+ "best_link_lower_bound": [
17
+ 0.0
18
+ ],
19
+ "note": "Small fixed-profile diagnostic, not a hardware or robustness validation. Terminal-step state may reset before the link diagnostic is read."
20
+ }
manifest.json CHANGED
@@ -9,10 +9,10 @@
9
  "policy_sha256": "59dd61ea1e8566a3bad9f8b24dd18530664906c33964b1c05e6ed6e821e16a24",
10
  "checkpoint_sha256": "f7e523c19de5cd5e2ced430f3d9bb46bef1ed69ad0003cb45955d56169929cdd",
11
  "source": {
12
- "repository": "Vottivott/microduck-secret-playground",
13
- "branch": "release/parkour",
14
- "revision": "58d56df0e863df552fac7caf409aaa968cfde34d",
15
- "public_base": "c594e4457983ed1c1c36193e3df60e16de66b819"
16
  },
17
  "profiles": {
18
  "backflip": {
@@ -43,12 +43,12 @@
43
  "additional_policy_switching": false,
44
  "onnx_export_path": "scripts/export.py",
45
  "simulation_contact_capacity": 200,
46
- "source_note": "Scoped import of retained research modules; shared public dependencies retained.",
47
  "warning": "Landing may damage or break the real robot. No hardware safety validation.",
48
  "training_continuation": {
49
- "repository": "Vottivott/microduck-secret-playground",
50
- "branch": "release/parkour",
51
- "revision": "58d56df0e863df552fac7caf409aaa968cfde34d",
52
  "profile": "backflip",
53
  "recipe": "training/resume.json",
54
  "original_release_binaries_unchanged": true,
 
9
  "policy_sha256": "59dd61ea1e8566a3bad9f8b24dd18530664906c33964b1c05e6ed6e821e16a24",
10
  "checkpoint_sha256": "f7e523c19de5cd5e2ced430f3d9bb46bef1ed69ad0003cb45955d56169929cdd",
11
  "source": {
12
+ "repository": "Vottivott/microduck-playground",
13
+ "branch": "main",
14
+ "revision": "b6aed572587b52e41a452a9b2be517c95d2b3d96",
15
+ "public_base": "4289c236864ea7d8f9a72951667b31c4fcd4b935"
16
  },
17
  "profiles": {
18
  "backflip": {
 
43
  "additional_policy_switching": false,
44
  "onnx_export_path": "scripts/export.py",
45
  "simulation_contact_capacity": 200,
46
+ "source_note": "Public two-policy split; original task physics and per-policy continuation settings unchanged from validated preparation. Pillar release excluded.",
47
  "warning": "Landing may damage or break the real robot. No hardware safety validation.",
48
  "training_continuation": {
49
+ "repository": "Vottivott/microduck-playground",
50
+ "branch": "main",
51
+ "revision": "b6aed572587b52e41a452a9b2be517c95d2b3d96",
52
  "profile": "backflip",
53
  "recipe": "training/resume.json",
54
  "original_release_binaries_unchanged": true,
training/resume.json CHANGED
@@ -1,38 +1,6 @@
1
  {
2
  "schema_version": 1,
3
  "profiles": {
4
- "long-jump": {
5
- "task": "Mjlab-PlatformJump-MicroDuck",
6
- "checkpoint_sha256": "114ab38724b1bea789b0d3f12a1a18a5981b40572863117e212a2a0601f9ed4a",
7
- "seed": 42,
8
- "provenance": "p1_v1 stored iteration 15250. Original launch_p1.sh and recorded agent/env YAML recovered. Entropy 0.005; no evaluation gap/drop pin. Uses the published platform-jump source, which contains later research changes, not the original p1 implementation.",
9
- "limitations": "Supported continuation from original PPO state under explicit current-source settings, not bit-exact p1 replay. Stage schedule is applied at the restored counter; new episode/RNG state. No external reset bank required.",
10
- "environment": {
11
- "MICRODUCK_PJ_ENTROPY": "0.005",
12
- "MICRODUCK_PJ_CURRICULUM_SHIFT": "0",
13
- "MICRODUCK_PJ_FLIGHT_PROB": "0.30"
14
- }
15
- },
16
- "pillar-jumps": {
17
- "task": "Mjlab-Parkour-MicroDuck",
18
- "checkpoint_sha256": "549ff41b9d892c7244252b7238bb0bd127fd3ec1665e811fb89435f099919828",
19
- "seed": 42,
20
- "provenance": "c29 stored iteration 34000. Reconstructed continuation from the retained research log: informed actor, five pillars, 12 cm pinned gap and hold requiring 75% standing height. Original c29 launcher/complete resolved run configuration not recovered; seed 42 and remaining pinned settings are explicit continuation choices.",
21
- "limitations": "Not exact historical replay. Per-environment promotion state was not checkpointed; fixed gap/drop deliberately disable promotion. Nine pillars are evaluation/transfer, not a separate checkpoint or the default continuation curriculum. No external reset bank required.",
22
- "environment": {
23
- "MICRODUCK_PK_BLIND": "0",
24
- "MICRODUCK_PK_PILLARS": "5",
25
- "MICRODUCK_PK_BASE_TOP": "1.0",
26
- "MICRODUCK_PK_DEPTH": "0.15",
27
- "MICRODUCK_PK_GAP": "0.12",
28
- "MICRODUCK_PK_DROP": "0.14",
29
- "MICRODUCK_PK_EPISODE_S": "11",
30
- "MICRODUCK_PK_HOLD_STAND_FRAC": "0.75",
31
- "MICRODUCK_PK_MAX_SPAWN_LINK": "3",
32
- "MICRODUCK_PK_ENTROPY": "0.004",
33
- "MICRODUCK_PK_CURRICULUM_SHIFT": "0"
34
- }
35
- },
36
  "backflip": {
37
  "task": "Mjlab-Flip-MicroDuck",
38
  "checkpoint_sha256": "f7e523c19de5cd5e2ced430f3d9bb46bef1ed69ad0003cb45955d56169929cdd",
 
1
  {
2
  "schema_version": 1,
3
  "profiles": {
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  "backflip": {
5
  "task": "Mjlab-Flip-MicroDuck",
6
  "checkpoint_sha256": "f7e523c19de5cd5e2ced430f3d9bb46bef1ed69ad0003cb45955d56169929cdd",
training/source.json CHANGED
@@ -1,7 +1,7 @@
1
  {
2
- "repository": "Vottivott/microduck-secret-playground",
3
- "branch": "release/parkour",
4
- "revision": "58d56df0e863df552fac7caf409aaa968cfde34d",
5
  "profile": "backflip",
6
  "recipe": "training/resume.json",
7
  "original_release_binaries_unchanged": true
 
1
  {
2
+ "repository": "Vottivott/microduck-playground",
3
+ "branch": "main",
4
+ "revision": "b6aed572587b52e41a452a9b2be517c95d2b3d96",
5
  "profile": "backflip",
6
  "recipe": "training/resume.json",
7
  "original_release_binaries_unchanged": true