Humanoid V1 β Velocity Locomotion Policy
A velocity-tracking locomotion policy for the ETHRC Humanoid V1 robot, trained in
Isaac Lab with rsl-rl
(PPO) on the ETHRC-humanoidv1-v0 task. The policy maps proprioceptive observations plus a target
base-velocity command to joint position targets, and tracks that command while keeping the robot balanced.
Velocity command
The policy is driven by a 3-D base_velocity command [lin_vel_x, lin_vel_y, ang_vel_z]:
| Component | Controls | Sign |
|---|---|---|
lin_vel_x |
lateral (strafe) | + = right, β = left |
lin_vel_y |
forward / backward | + = forward, β = backward |
ang_vel_z |
yaw (turn) | + = turn left |
Linear components are in m/s, angular in rad/s.
Capabilities
Measured as the mean over 300 steps under fixed commands:
| Command | Achieved | Tracking |
|---|---|---|
Forward (lin_vel_y = 0.5) |
~0.36 m/s | ~71% |
Sidestep (lin_vel_x = Β±0.5) |
~0.36 m/s | ~71% |
Turn (ang_vel_z = 0.5) |
0.34 rad/s | 68% |
Diagonal (lin_vel_y = 0.4, lin_vel_x = 0.3) |
0.33 / 0.21 m/s | 70β82% |
Top speed (lin_vel_y = 1.0) |
0.71 m/s | 71% |
The policy is omnidirectional β forward, backward, lateral, turning, and combined commands (translation + yaw together) all work, and it stays upright under disturbances. Tracking undershoots roughly uniformly (~70%), so a practical rule is to command β 1.4Γ the target speed to reach it. Sustained top forward speed is about 0.7 m/s.
Example rollouts for each command are in videos/.
Files
| Path | Description |
|---|---|
model_3999.pt |
rsl-rl checkpoint (final, iteration 3999) |
exported/policy.pt |
TorchScript policy (deployment) |
exported/policy.onnx (+ .data) |
ONNX policy (deployment) |
params/agent.yaml, params/env.yaml |
training / environment configs |
videos/ |
evaluation clips per command |
Usage
Evaluate or record a rollout with Isaac Lab (requires rsl-rl β₯ 5.0):
cd <rc_humanoid_rl_lab>
isaaclab.sh -p scripts/rsl_rl/play.py \
--task=ETHRC-humanoidv1-v0 --num_envs=4 --headless --video --video_length=300 \
--checkpoint model_3999.pt \
env.commands.base_velocity.ranges.lin_vel_y="[0.5, 0.5]" \
env.commands.base_velocity.ranges.lin_vel_x="[0.0, 0.0]" \
env.commands.base_velocity.ranges.ang_vel_z="[0.0, 0.0]"
Set lin_vel_y for forward/backward, lin_vel_x for strafing, and ang_vel_z for turning.
For deployment, load exported/policy.pt (TorchScript) or exported/policy.onnx (ONNX) β both take
the observation vector and return joint position targets.
Training
- Framework: Isaac Lab + rsl-rl (PPO)
- Task:
ETHRC-humanoidv1-v0β velocity-tracking locomotion - Observations: proprioception + base-velocity command
- Actions: joint position targets
- Curriculum:
lin_vel_cmd_levelsprogressively widens the commanded velocity range as thetrack_lin_vel_xyreward improves, so the policy is trained on harder commands only once it tracks the easier ones.
Exact hyperparameters and reward terms are in params/agent.yaml and params/env.yaml.