--- license: other library_name: rsl-rl tags: - reinforcement-learning - robotics - humanoid - locomotion - isaac-lab --- # Humanoid V1 — Velocity Locomotion Policy A velocity-tracking locomotion policy for the ETHRC Humanoid V1 robot, trained in [Isaac Lab](https://github.com/isaac-sim/IsaacLab) with [rsl-rl](https://github.com/leggedrobotics/rsl_rl) (PPO) on the `ETHRC-humanoidv1-v0` task. The policy maps proprioceptive observations plus a target base-velocity command to joint position targets, and tracks that command while keeping the robot balanced. ## Velocity command The policy is driven by a 3-D `base_velocity` command `[lin_vel_x, lin_vel_y, ang_vel_z]`: | Component | Controls | Sign | | --- | --- | --- | | `lin_vel_x` | lateral (strafe) | `+` = right, `−` = left | | `lin_vel_y` | forward / backward | `+` = forward, `−` = backward | | `ang_vel_z` | yaw (turn) | `+` = turn left | Linear components are in m/s, angular in rad/s. ## Capabilities Measured as the mean over 300 steps under fixed commands: | Command | Achieved | Tracking | | --- | --- | --- | | Forward (`lin_vel_y = 0.5`) | ~0.36 m/s | ~71% | | Sidestep (`lin_vel_x = ±0.5`) | ~0.36 m/s | ~71% | | Turn (`ang_vel_z = 0.5`) | 0.34 rad/s | 68% | | Diagonal (`lin_vel_y = 0.4`, `lin_vel_x = 0.3`) | 0.33 / 0.21 m/s | 70–82% | | Top speed (`lin_vel_y = 1.0`) | 0.71 m/s | 71% | The policy is omnidirectional — forward, backward, lateral, turning, and combined commands (translation + yaw together) all work, and it stays upright under disturbances. Tracking undershoots roughly uniformly (~70%), so a practical rule is to **command ≈ 1.4× the target speed** to reach it. Sustained top forward speed is about **0.7 m/s**. Example rollouts for each command are in [`videos/`](./videos). ## Files | Path | Description | | --- | --- | | `model_3999.pt` | rsl-rl checkpoint (final, iteration 3999) | | `exported/policy.pt` | TorchScript policy (deployment) | | `exported/policy.onnx` (+ `.data`) | ONNX policy (deployment) | | `params/agent.yaml`, `params/env.yaml` | training / environment configs | | `videos/` | evaluation clips per command | ## Usage Evaluate or record a rollout with Isaac Lab (requires rsl-rl ≥ 5.0): ```bash cd isaaclab.sh -p scripts/rsl_rl/play.py \ --task=ETHRC-humanoidv1-v0 --num_envs=4 --headless --video --video_length=300 \ --checkpoint model_3999.pt \ env.commands.base_velocity.ranges.lin_vel_y="[0.5, 0.5]" \ env.commands.base_velocity.ranges.lin_vel_x="[0.0, 0.0]" \ env.commands.base_velocity.ranges.ang_vel_z="[0.0, 0.0]" ``` Set `lin_vel_y` for forward/backward, `lin_vel_x` for strafing, and `ang_vel_z` for turning. For deployment, load `exported/policy.pt` (TorchScript) or `exported/policy.onnx` (ONNX) — both take the observation vector and return joint position targets. ## Training - **Framework:** Isaac Lab + rsl-rl (PPO) - **Task:** `ETHRC-humanoidv1-v0` — velocity-tracking locomotion - **Observations:** proprioception + base-velocity command - **Actions:** joint position targets - **Curriculum:** `lin_vel_cmd_levels` progressively widens the commanded velocity range as the `track_lin_vel_xy` reward improves, so the policy is trained on harder commands only once it tracks the easier ones. Exact hyperparameters and reward terms are in `params/agent.yaml` and `params/env.yaml`.