SjohnU commited on
Commit
5d4f6b8
Β·
verified Β·
1 Parent(s): 747c33d

r16lat_g250: deploy the BEST checkpoint (best-of selection), not model_9999

Browse files
r16lat_g250/README.md CHANGED
@@ -1,48 +1,61 @@
1
- # r16lat_g250 β€” Report 17 leg-gain sweep (gain=250)
2
 
3
- leg-gain point 250 in the r17 sweep. Measured **0.760 m/s at command 1.2**.
 
 
4
 
5
  | | |
6
  |---|---|
7
  | Task | `ETHRC-humanoidv1-walk-r17-g250-psu-v0` |
 
 
8
  | Branch | `test/SjohnU_remodel_actuators` |
9
- | Checkpoint | `model_9999.pt` (10,000 iters) |
10
- | Envs / seed | 4096 / 42 |
11
  | rsl-rl | 5.0.1 |
12
- | wandb | `6g30agt5` (project rc_humanoid_rl) |
13
- | Hardware | AWS us-east-2, 1x L40S (g6e.xlarge) |
14
 
15
- ## Report-17 leg-gain sweep (r16-latency lineage) β€” measured across the command sweep
16
 
17
- Five leg-gain points (g210..g250) off the `r16_latency` walker, each an otherwise-identical clean
18
- ablation, 10k iters, 4096 envs, seed 42, branch `test/SjohnU_remodel_actuators`. Measured with
19
- `scripts/rsl_rl/measure_gait.py` (1000 steps x 64 envs) at three forward commands.
20
- **Physical m/s, not reward-term tvel.**
21
 
22
- | leg gain | @0.5 | @0.8 | @1.2 | lean@1.2 |
23
  |---|---|---|---|---|
24
- | g210 | 0.365 | 0.579 | 0.708 | 3.3 deg fwd |
25
- | **g220** | 0.414 | 0.691 | **0.981** | 1.8 deg back |
26
- | g230 | 0.393 | 0.641 | 0.884 | 5.8 deg back |
27
- | g240 | 0.422 | 0.648 | 0.883 | 3.5 deg back |
28
- | g250 | 0.397 | 0.655 | 0.760 | 1.3 deg fwd |
29
-
30
- ### Findings
31
-
32
- 1. **g220 is the sweet spot: 0.981 m/s @ command 1.2, nearly upright (1.8 deg).** It ties the
33
- campaign's fastest walker (`final_act`, 0.977) and has the cleanest command scaling
34
- (0.414 -> 0.691 -> 0.981).
35
-
36
- 2. **The gain vs speed relationship is an INVERTED-U, not monotonic.** 210 -> 220 (peak) -> 230
37
- ~= 240 -> 250. Both softer (g210, 0.708) and stiffer (g250, 0.760) legs are slower at the top
38
- command than g220. Stiffer is not simply better.
39
-
40
- 3. **The effect is invisible at command 0.5.** All five cluster at 0.37-0.42 m/s there -- the
41
- 0.5 training cap. The 29% spread (g220 0.981 vs g210 0.708) only appears at command 1.2.
42
- A single-point measurement at 0.5 would wrongly conclude the gain does not matter; the sweep
43
- is what surfaces it. (Same lesson as the r16 slate: measure across commands, in physical units.)
 
 
 
 
 
 
 
 
 
 
 
 
 
44
 
45
  ## Files
46
- - `policy.pt`, `policy.onnx` β€” exported via `scripts/rsl_rl/play_export.py`
47
  - `deploy.yaml` β€” deploy config (50 Hz)
48
- - `model_9999.pt` β€” raw rsl-rl checkpoint
 
 
1
+ # r16lat_g250 β€” BEST-checkpoint bundle (model_8000)
2
 
3
+ **`policy.pt` / `policy.onnx` here are model_8000, the BEST checkpoint by measured velocity
4
+ (0.925 m/s @ cmd 1.2), NOT the final model_9999 (0.760 m/s).** `model_9999.pt` is retained
5
+ in this folder for reference only.
6
 
7
  | | |
8
  |---|---|
9
  | Task | `ETHRC-humanoidv1-walk-r17-g250-psu-v0` |
10
+ | Deployed checkpoint | `model_8000.pt` (best of the run) |
11
+ | Measured @ cmd 1.2 | 0.925 m/s |
12
  | Branch | `test/SjohnU_remodel_actuators` |
 
 
13
  | rsl-rl | 5.0.1 |
 
 
14
 
15
+ ## Best-checkpoint selection (model_9999 is the LAST, not the BEST)
16
 
17
+ Each candidate's late checkpoints were measured at command 1.2 (deploy speed) and the highest
18
+ selected. `model_9999` won for only 3 of 7 β€” for the other 4 an earlier checkpoint is better,
19
+ by up to +22%.
 
20
 
21
+ | run | best ckpt | best m/s @1.2 | model_9999 @1.2 | delta |
22
  |---|---|---|---|---|
23
+ | r16_final_act | **model_9000** | **1.004** | 0.977 | +0.027 (first >1 m/s) |
24
+ | r16_h2_noar | model_9999 | 0.888 | 0.888 | β€” |
25
+ | r16lat_g210 | **model_8000** | **0.838** | 0.708 | +0.130 (+18%) |
26
+ | r16lat_g220 | model_9999 | 0.981 | 0.981 | β€” |
27
+ | r16lat_g230 | model_9999 | 0.884 | 0.884 | β€” |
28
+ | r16lat_g240 | **model_8000** | **0.958** | 0.883 | +0.075 (+8%) |
29
+ | r16lat_g250 | **model_8000** | **0.925** | 0.760 | +0.165 (+22%) |
30
+
31
+ Per-checkpoint measurements (m/s @ cmd 1.2):
32
+
33
+ | run | 5000 | 6000 | 7000 | 8000 | 9000 | 9999 |
34
+ |---|---|---|---|---|---|---|
35
+ | final_act | 0.514 | 0.870 | 0.973 | 0.941 | **1.004** | 0.977 |
36
+ | h2_noar | 0.840 | 0.857 | 0.870 | 0.845 | 0.818 | **0.888** |
37
+ | g210 | 0.749 | 0.717 | 0.803 | **0.838** | 0.777 | 0.708 |
38
+ | g220 | 0.775 | 0.801 | 0.958 | 0.893 | 0.947 | **0.981** |
39
+ | g230 | 0.803 | 0.857 | 0.792 | 0.865 | 0.833 | **0.884** |
40
+ | g240 | 0.933 | 0.865 | 0.957 | **0.958** | 0.956 | 0.883 |
41
+ | g250 | 0.642 | 0.745 | 0.781 | **0.925** | 0.794 | 0.760 |
42
+
43
+ ### This corrects the gain-sweep conclusion
44
+
45
+ The gain curve is FLAT, not a sharp inverted-U. The earlier "sharp peak at g220, dropoff by g250"
46
+ was an artifact of comparing final (model_9999) checkpoints, which land at noisy points:
47
+
48
+ ```
49
+ model_9999 only (WRONG): g210 .708 g220 .981 g230 .884 g240 .883 g250 .760
50
+ best-of (CORRECT): g210 .838 g220 .981 g230 .884 g240 .958 g250 .925
51
+ ```
52
+
53
+ On best checkpoints, g220-g250 are a broad plateau (~0.88-0.98 m/s) -- leg stiffness in this range
54
+ barely matters. The apparent g250 collapse (0.760) was purely that run sitting at a poor iter-9999;
55
+ its true best is 0.925. **Judge policies on their best measured checkpoint, not their last.**
56
 
57
  ## Files
58
+ - `policy.pt`, `policy.onnx` β€” the BEST checkpoint (model_8000), exported for deploy
59
  - `deploy.yaml` β€” deploy config (50 Hz)
60
+ - `model_8000.pt` β€” the best raw checkpoint
61
+ - `model_9999.pt` β€” the final checkpoint, kept for reference
r16lat_g250/deploy.yaml CHANGED
@@ -422,4 +422,4 @@ command_ranges:
422
  ang_vel_z:
423
  - -0.1
424
  - 0.1
425
- checkpoint: /root/rc/logs/rsl_rl/humanoid_v1_locomotion/r16lat_g250/model_9999.pt
 
422
  ang_vel_z:
423
  - -0.1
424
  - 0.1
425
+ checkpoint: /root/rc/logs/rsl_rl/humanoid_v1_locomotion/r16lat_g250/model_8000.pt
r16lat_g250/model_8000.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4d050e96747eafcb831f71f45177c1b92a5e019ac8f7b7d82ea4ae43143aace4
3
+ size 7475573
r16lat_g250/policy.onnx CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:5a8cca4950a35782b7bd1a1e7a41dce2ebd8f6e336310464e8a475a862059875
3
  size 1584867
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b4ce6f43a6ece7a6a44a36aeafa5d333e12c341abb91c55d23fb0accaa6cba15
3
  size 1584867
r16lat_g250/policy.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:0d8dab117480ad66ea7fb9ba8ff711357f4238ba9713e6e9f4e7e81c56d6382f
3
  size 1596822
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f1adcaca5264906bb4b63bff7d42b0f2cdb5b47a97c646a0d1ce1bfddb639d42
3
  size 1596822