userdarius's picture
Rewrite model card to focus on how the policy works
95ce045 verified
|
Raw History Blame Contribute Delete
3.39 kB
---
license: other
library_name: rsl-rl
tags:
- reinforcement-learning
- robotics
- humanoid
- locomotion
- isaac-lab
---
# Humanoid V1 β€” Velocity Locomotion Policy
A velocity-tracking locomotion policy for the ETHRC Humanoid V1 robot, trained in
[Isaac Lab](https://github.com/isaac-sim/IsaacLab) with [rsl-rl](https://github.com/leggedrobotics/rsl_rl)
(PPO) on the `ETHRC-humanoidv1-v0` task. The policy maps proprioceptive observations plus a target
base-velocity command to joint position targets, and tracks that command while keeping the robot balanced.
## Velocity command
The policy is driven by a 3-D `base_velocity` command `[lin_vel_x, lin_vel_y, ang_vel_z]`:
| Component | Controls | Sign |
| --- | --- | --- |
| `lin_vel_x` | lateral (strafe) | `+` = right, `βˆ’` = left |
| `lin_vel_y` | forward / backward | `+` = forward, `βˆ’` = backward |
| `ang_vel_z` | yaw (turn) | `+` = turn left |
Linear components are in m/s, angular in rad/s.
## Capabilities
Measured as the mean over 300 steps under fixed commands:
| Command | Achieved | Tracking |
| --- | --- | --- |
| Forward (`lin_vel_y = 0.5`) | ~0.36 m/s | ~71% |
| Sidestep (`lin_vel_x = Β±0.5`) | ~0.36 m/s | ~71% |
| Turn (`ang_vel_z = 0.5`) | 0.34 rad/s | 68% |
| Diagonal (`lin_vel_y = 0.4`, `lin_vel_x = 0.3`) | 0.33 / 0.21 m/s | 70–82% |
| Top speed (`lin_vel_y = 1.0`) | 0.71 m/s | 71% |
The policy is omnidirectional β€” forward, backward, lateral, turning, and combined commands
(translation + yaw together) all work, and it stays upright under disturbances. Tracking undershoots
roughly uniformly (~70%), so a practical rule is to **command β‰ˆ 1.4Γ— the target speed** to reach it.
Sustained top forward speed is about **0.7 m/s**.
Example rollouts for each command are in [`videos/`](./videos).
## Files
| Path | Description |
| --- | --- |
| `model_3999.pt` | rsl-rl checkpoint (final, iteration 3999) |
| `exported/policy.pt` | TorchScript policy (deployment) |
| `exported/policy.onnx` (+ `.data`) | ONNX policy (deployment) |
| `params/agent.yaml`, `params/env.yaml` | training / environment configs |
| `videos/` | evaluation clips per command |
## Usage
Evaluate or record a rollout with Isaac Lab (requires rsl-rl β‰₯ 5.0):
```bash
cd <rc_humanoid_rl_lab>
isaaclab.sh -p scripts/rsl_rl/play.py \
--task=ETHRC-humanoidv1-v0 --num_envs=4 --headless --video --video_length=300 \
--checkpoint model_3999.pt \
env.commands.base_velocity.ranges.lin_vel_y="[0.5, 0.5]" \
env.commands.base_velocity.ranges.lin_vel_x="[0.0, 0.0]" \
env.commands.base_velocity.ranges.ang_vel_z="[0.0, 0.0]"
```
Set `lin_vel_y` for forward/backward, `lin_vel_x` for strafing, and `ang_vel_z` for turning.
For deployment, load `exported/policy.pt` (TorchScript) or `exported/policy.onnx` (ONNX) β€” both take
the observation vector and return joint position targets.
## Training
- **Framework:** Isaac Lab + rsl-rl (PPO)
- **Task:** `ETHRC-humanoidv1-v0` β€” velocity-tracking locomotion
- **Observations:** proprioception + base-velocity command
- **Actions:** joint position targets
- **Curriculum:** `lin_vel_cmd_levels` progressively widens the commanded velocity range as the
`track_lin_vel_xy` reward improves, so the policy is trained on harder commands only once it tracks
the easier ones.
Exact hyperparameters and reward terms are in `params/agent.yaml` and `params/env.yaml`.