SjohnU commited on
Commit
0feb9e1
·
verified ·
1 Parent(s): 8bf1692

Sync policy checkpoints (2026-06-25T14:03:13Z)

Browse files
! ADDED
File without changes
ETHRC-humanoidv1-suspended-imu-v0_u4_imu_arm_lg_es500/README.md ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Suspended IMU arm-pointing policy — pure-IMU, low-gain, stabilized (`u4_imu_arm_lg_es500`)
2
+
3
+ A bring-up policy that validates the IMU pipeline end-to-end on ETHRC Humanoid V1 (URDF_4). The
4
+ suspended (fixed-base) robot is tilted continuously and the **arms point down the felt-gravity
5
+ vector**. The policy reads **only the IMU**, so correct pointing proves the IMU observation path is
6
+ wired and consumed correctly — before the IMU is trusted by the walker.
7
+
8
+ This run supersedes `u4_imu_arm_pure`: it trains at the **updated (lower) arm drive gains** and is
9
+ **early-stopped at the peak** so it does not erode (the long 8000-iter run drifted from ~9.5 down to
10
+ 8.2 as entropy crept back up). Exported checkpoint: **`model_175`** (the reward peak).
11
+
12
+ ## Observation (actor) — IMU only
13
+ A real IMU provides exactly these two quantities, and the deployable policy sees nothing else:
14
+
15
+ | term | meaning | dims | scale |
16
+ |---|---|---|---|
17
+ | `base_ang_vel` | gyro: torso angular velocity in the base frame | 3 | 0.2 |
18
+ | `projected_gravity` | felt-gravity direction in the base frame = `quat_rotate_inverse(base_quat, [0,0,-1])` | 3 | 1.0 |
19
+
20
+ `history_length = 5` → flattened obs **30**. No joint positions/velocities or last-action, so the
21
+ policy cannot take proprioceptive shortcuts and must derive the pose from the IMU. (The critic uses
22
+ full privileged obs during training only — asymmetric actor-critic.)
23
+
24
+ ## Action — 6 arm joints, position targets
25
+ `shoulder_flexion_{l,r}`, `shoulder_abduction_{l,r}`, `elbow_flexion_{l,r}` (scale 0.5, default
26
+ offset). All other joints are PD-held at default. The action is a **position target**, so the joint
27
+ PD closes the loop and the policy is a pure **IMU → arm-pose** map (sim-agnostic: same IMU → same
28
+ commanded pose in Isaac, MuJoCo, and on hardware). `shoulder_rotation` is intentionally excluded (it
29
+ twists about the arm axis without changing the pointing direction).
30
+
31
+ ## Drive gains (matches the control-repo "test lower gains" update)
32
+ Trained and deployed at the lowered maxon JVPT gains (JVPT = sim gain × 1000):
33
+
34
+ | group | joints | kp | kd |
35
+ |---|---|---|---|
36
+ | HEJ70 shoulders | `shoulder_flexion`, `shoulder_abduction` | 75 | 1.8 |
37
+ | HEJ50 arms/neck | `shoulder_rotation`, `elbow_flexion`, … | 37.5 | 0.9 |
38
+
39
+ Legs/waist stay P200 (200/10) and are PD-held; they do not move in this task.
40
+
41
+ ## Training
42
+ Isaac Lab + RSL-RL PPO, 4096 envs, **500 iterations (early-stop)**. The base is driven by a
43
+ continuous, per-env randomized tilt sweep (±45°) so the policy learns to track a *moving* felt-gravity
44
+ vector. Reward: both arm segments (shoulder→elbow and elbow→wrist) aligned with world-down — a dense
45
+ `(cos+1)/2` term plus a tight `exp(-(1-cos)/0.10)` term — averaged over both arms, plus light
46
+ smoothness penalties (`action_rate −0.1`, `joint_acc −2.5e-6`, `energy −2e-5`).
47
+
48
+ Stability settings (vs the earlier run that eroded): `entropy_coef 0.005`, `desired_kl 0.005` (the
49
+ adaptive LR settles instead of wandering). Result: `point_gravity` peaks **9.63 @ iter 201** and holds
50
+ in a tight 9.4–9.6 band (drop-from-peak 0.24, vs 1.29 before); action std collapses 0.99 → 0.15.
51
+
52
+ ## Files
53
+ - `policy.pt` — TorchScript actor (obs 30 → 6 actions); the obs normalizer is folded in.
54
+ - `policy.onnx` — same network in ONNX.
55
+ - `deploy.yaml` — full deployment config: obs layout (term order/scale/history), per-joint PD gains
56
+ and limits, action joint order/scale, control rate, and the source checkpoint.
57
+
58
+ ## Use
59
+ Drive it with the production deploy brain (`DeployConfig` + `PolicyController` + `ObservationBuilder`).
60
+ Validate in MuJoCo with `sim2sim_imu_tilt.py` (tilt the base, watch the arms aim down) or feed the
61
+ real Xsens MTi-610 IMU. base_link convention: +X = right, +Y = forward, +Z = up.
ETHRC-humanoidv1-suspended-imu-v0_u4_imu_arm_lg_es500/deploy.yaml ADDED
@@ -0,0 +1,261 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ experiment_name: ETHRC-humanoidv1-suspended-imu-v0_u4_imu_arm_lg_es500
2
+ task: ETHRC-humanoidv1-suspended-imu-v0
3
+ interface: lo
4
+ sim2sim: true
5
+ msg_type: go
6
+ imu_type: base_link
7
+ lowcmd_topic: rt/lowcmd
8
+ lowstate_topic: rt/lowstate
9
+ policy_path: policy.pt
10
+ num_actions: 6
11
+ action_scale: 0.5
12
+ use_default_offset: true
13
+ control_dt: 0.02
14
+ policy_rate_hz: 50.0
15
+ decimation: 4
16
+ sim_dt: 0.005
17
+ joint_names:
18
+ - shoulder_flexion_l
19
+ - shoulder_flexion_r
20
+ - shoulder_abduction_l
21
+ - shoulder_abduction_r
22
+ - elbow_flexion_l
23
+ - elbow_flexion_r
24
+ joint2motor_idx:
25
+ - 0
26
+ - 1
27
+ - 2
28
+ - 3
29
+ - 4
30
+ - 5
31
+ default_angles:
32
+ - 0.0
33
+ - 0.0
34
+ - 0.0
35
+ - 0.0
36
+ - 0.0
37
+ - 0.0
38
+ kps:
39
+ - 75.0
40
+ - 75.0
41
+ - 75.0
42
+ - 75.0
43
+ - 37.5
44
+ - 37.5
45
+ kds:
46
+ - 1.7999999523162842
47
+ - 1.7999999523162842
48
+ - 1.7999999523162842
49
+ - 1.7999999523162842
50
+ - 0.8999999761581421
51
+ - 0.8999999761581421
52
+ effort_limits:
53
+ - 80.0
54
+ - 80.0
55
+ - 80.0
56
+ - 80.0
57
+ - 40.0
58
+ - 40.0
59
+ articulation_joint_names:
60
+ - hip_central_rotation
61
+ - waist_upperbody_yaw
62
+ - leg_flexion_l
63
+ - leg_flexion_r
64
+ - neck_yaw
65
+ - shoulder_flexion_l
66
+ - shoulder_flexion_r
67
+ - leg_abduction_l
68
+ - leg_abduction_r
69
+ - neck_pitch
70
+ - shoulder_abduction_l
71
+ - shoulder_abduction_r
72
+ - hip_yaw_l
73
+ - hip_yaw_r
74
+ - shoulder_rotation_l
75
+ - shoulder_rotation_r
76
+ - knee_l
77
+ - knee_r
78
+ - elbow_flexion_l
79
+ - elbow_flexion_r
80
+ - foot_pitch_l
81
+ - foot_pitch_r
82
+ - wrist_rotation_l
83
+ - wrist_rotation_r
84
+ - foot_roll_l
85
+ - foot_roll_r
86
+ articulation_default_angles:
87
+ - 0.0
88
+ - 0.0
89
+ - 0.0
90
+ - 0.0
91
+ - 0.0
92
+ - 0.0
93
+ - 0.0
94
+ - 0.0
95
+ - 0.0
96
+ - 0.0
97
+ - 0.0
98
+ - 0.0
99
+ - 0.0
100
+ - 0.0
101
+ - 0.0
102
+ - 0.0
103
+ - 0.0
104
+ - 0.0
105
+ - 0.0
106
+ - 0.0
107
+ - 0.0
108
+ - 0.0
109
+ - 0.0
110
+ - 0.0
111
+ - 0.0
112
+ - 0.0
113
+ pd_per_joint:
114
+ hip_central_rotation:
115
+ kp: 200.0
116
+ kd: 10.0
117
+ effort_limit: 128.0
118
+ velocity_limit: 14.65999984741211
119
+ waist_upperbody_yaw:
120
+ kp: 200.0
121
+ kd: 10.0
122
+ effort_limit: 128.0
123
+ velocity_limit: 14.65999984741211
124
+ leg_flexion_l:
125
+ kp: 200.0
126
+ kd: 10.0
127
+ effort_limit: 128.0
128
+ velocity_limit: 14.65999984741211
129
+ leg_flexion_r:
130
+ kp: 200.0
131
+ kd: 10.0
132
+ effort_limit: 128.0
133
+ velocity_limit: 14.65999984741211
134
+ leg_abduction_l:
135
+ kp: 200.0
136
+ kd: 10.0
137
+ effort_limit: 128.0
138
+ velocity_limit: 14.65999984741211
139
+ leg_abduction_r:
140
+ kp: 200.0
141
+ kd: 10.0
142
+ effort_limit: 128.0
143
+ velocity_limit: 14.65999984741211
144
+ hip_yaw_l:
145
+ kp: 200.0
146
+ kd: 10.0
147
+ effort_limit: 128.0
148
+ velocity_limit: 14.65999984741211
149
+ hip_yaw_r:
150
+ kp: 200.0
151
+ kd: 10.0
152
+ effort_limit: 128.0
153
+ velocity_limit: 14.65999984741211
154
+ knee_l:
155
+ kp: 200.0
156
+ kd: 10.0
157
+ effort_limit: 128.0
158
+ velocity_limit: 14.65999984741211
159
+ knee_r:
160
+ kp: 200.0
161
+ kd: 10.0
162
+ effort_limit: 128.0
163
+ velocity_limit: 14.65999984741211
164
+ foot_pitch_l:
165
+ kp: 40.0
166
+ kd: 1.0
167
+ effort_limit: 40.0
168
+ velocity_limit: 15.0
169
+ foot_pitch_r:
170
+ kp: 40.0
171
+ kd: 1.0
172
+ effort_limit: 40.0
173
+ velocity_limit: 15.0
174
+ foot_roll_l:
175
+ kp: 40.0
176
+ kd: 1.0
177
+ effort_limit: 40.0
178
+ velocity_limit: 15.0
179
+ foot_roll_r:
180
+ kp: 40.0
181
+ kd: 1.0
182
+ effort_limit: 40.0
183
+ velocity_limit: 15.0
184
+ shoulder_flexion_l:
185
+ kp: 75.0
186
+ kd: 1.7999999523162842
187
+ effort_limit: 80.0
188
+ velocity_limit: 18.0
189
+ shoulder_flexion_r:
190
+ kp: 75.0
191
+ kd: 1.7999999523162842
192
+ effort_limit: 80.0
193
+ velocity_limit: 18.0
194
+ shoulder_abduction_l:
195
+ kp: 75.0
196
+ kd: 1.7999999523162842
197
+ effort_limit: 80.0
198
+ velocity_limit: 18.0
199
+ shoulder_abduction_r:
200
+ kp: 75.0
201
+ kd: 1.7999999523162842
202
+ effort_limit: 80.0
203
+ velocity_limit: 18.0
204
+ neck_yaw:
205
+ kp: 37.5
206
+ kd: 0.8999999761581421
207
+ effort_limit: 40.0
208
+ velocity_limit: 20.0
209
+ neck_pitch:
210
+ kp: 37.5
211
+ kd: 0.8999999761581421
212
+ effort_limit: 40.0
213
+ velocity_limit: 20.0
214
+ shoulder_rotation_l:
215
+ kp: 37.5
216
+ kd: 0.8999999761581421
217
+ effort_limit: 40.0
218
+ velocity_limit: 20.0
219
+ shoulder_rotation_r:
220
+ kp: 37.5
221
+ kd: 0.8999999761581421
222
+ effort_limit: 40.0
223
+ velocity_limit: 20.0
224
+ elbow_flexion_l:
225
+ kp: 37.5
226
+ kd: 0.8999999761581421
227
+ effort_limit: 40.0
228
+ velocity_limit: 20.0
229
+ elbow_flexion_r:
230
+ kp: 37.5
231
+ kd: 0.8999999761581421
232
+ effort_limit: 40.0
233
+ velocity_limit: 20.0
234
+ wrist_rotation_l:
235
+ kp: 37.5
236
+ kd: 0.8999999761581421
237
+ effort_limit: 40.0
238
+ velocity_limit: 20.0
239
+ wrist_rotation_r:
240
+ kp: 37.5
241
+ kd: 0.8999999761581421
242
+ effort_limit: 40.0
243
+ velocity_limit: 20.0
244
+ obs_history_length: 5
245
+ obs_single_frame_dim: 6
246
+ obs_total_dim: 30
247
+ obs_layout:
248
+ - name: base_ang_vel
249
+ start: 0
250
+ end: 15
251
+ base_dim: 3
252
+ block_dim: 15
253
+ scale: 0.2
254
+ - name: projected_gravity
255
+ start: 15
256
+ end: 30
257
+ base_dim: 3
258
+ block_dim: 15
259
+ scale: 1.0
260
+ command_ranges: null
261
+ checkpoint: /home/sjohn/Documents/project/rc_humanoid_rl_lab/logs/rsl_rl/humanoid_v1_locomotion/ETHRC-humanoidv1-suspended-imu-v0_u4_imu_arm_lg_es500/model_175.pt
ETHRC-humanoidv1-suspended-imu-v0_u4_imu_arm_lg_es500/policy.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:13c5ec50d2344b8f817a3bddfed723cdef61c7f929003dd2af1e99c2bb5164d5
3
+ size 724625
ETHRC-humanoidv1-suspended-imu-v0_u4_imu_arm_lg_es500/policy.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:562ac44afc4bbd4c350d70273729ad6da1ebfd500b19f75673dac429e076005a
3
+ size 736662
deploy.md ADDED
@@ -0,0 +1,55 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # policy_checkpoints — deploy workflow
2
+
3
+ Deployment checkpoints for the ETHRC humanoid live here, synced with the **private**
4
+ HuggingFace bucket `ETHRC-humanoid/rc-humanoid-locomotion`. Each run is one subfolder
5
+ holding the deploy contract + weights (`deploy.yaml`, `policy.pt`, `policy.onnx[.data]`).
6
+
7
+ **Nothing in this folder is git-tracked** except the sync scripts and this doc (see
8
+ `.gitignore`) — the weights and configs live on HF and are pulled on demand, so the repo
9
+ never bloats and never drifts from the bucket.
10
+
11
+ ## 0. One-time access (per machine)
12
+
13
+ You need to be in the `ETHRC-humanoid` HF org. In the training repo, run:
14
+
15
+ ```bash
16
+ cd <rc_humanoid_rl_lab> && ./setup_hf.sh
17
+ ```
18
+
19
+ It tells you to **ask Darius on Discord** for org access, then stores your token
20
+ (`~/.hf_token` + `HF_TOKEN` in your shell profile). On a cluster, just set `HF_TOKEN` as a
21
+ job secret instead. The pull/push scripts here read it automatically.
22
+
23
+ ## 1. Pull the latest checkpoints
24
+
25
+ ```bash
26
+ ./pull_from_hf.sh
27
+ ```
28
+
29
+ Downloads every run folder from the bucket into this directory.
30
+
31
+ ## 2. Deploy in sim2sim
32
+
33
+ The `sim2sim` package drives any pulled checkpoint via `--yaml <RUN_FOLDER>/deploy.yaml`. Use a python
34
+ with torch + mujoco (the `ethrc-deploy` / `env_isaaclab` env). After
35
+ `colcon build --packages-select sim2sim deployment`:
36
+
37
+ ```bash
38
+ # suspended (hang) — bring-up policies (knee/arm/ankle cycle, joint tracking)
39
+ ros2 run sim2sim sim2sim --yaml "$(pwd)/<RUN_FOLDER>/deploy.yaml" --base hang --view
40
+
41
+ # ground / soft-release (G1 elastic band) — standing / locomotion policies
42
+ ros2 run sim2sim sim2sim --yaml "$(pwd)/<RUN_FOLDER>/deploy.yaml" --base release --view
43
+ ```
44
+
45
+ Without a built workspace, run the module directly:
46
+ `PYTHONPATH=<repo>/ros2_ws/src/sim2sim:<repo>/ros2_ws/src/deployment python -m sim2sim.sim2sim --yaml <…>/deploy.yaml --base hang --view`
47
+
48
+ ## 3. Push new checkpoints
49
+
50
+ The primary push is from the training repo right after `play_export.py`
51
+ (`<rc_humanoid_rl_lab>/upload_deploy_to_hf.sh`). To push checkpoints curated here:
52
+
53
+ ```bash
54
+ ./push_to_hf.sh
55
+ ```