strands-isaaclab-h1-rough-policy

rsl_rl PPO policy for Unitree H1 humanoid rough-terrain locomotion (Isaac-Velocity-Rough-H1), trained with strands-robots' isaaclab train_policy provider (PR #4227). Rollouts recorded through strands: cagataydev/strands-isaaclab-h1-rough.

playback: 4 parallel envs

envs 0–3 (2×2) from the strands camera — mp4.

Results

parallel envs × iterations 4096 × 1500 (147 M env steps), PhysX (physics=isaacsim_physx)
wall time 63 min 38 s on 1× NVIDIA L40S
env-steps/s median 48 k, max 121 k (GPU shared with other runs most of the time)
mean reward -0.16 → 26.1 best → 25.4 last
velocity-tracking success 1.00 last iteration; mean episode length 989 of 1000; terrain level 5.9
recorded rollout 12 / 12 envs walked the full 10 s window; mean return 17.7
PPO iteration mean reward velocity-tracking success mean ep. length (of 1000) terrain curriculum level env-steps/s
0 -0.16 0.042 12 3.51 25,864
150 4.74 0.000 1000 0.80 48,995
375 12.47 0.954 951 4.27 48,268
750 15.59 1.000 977 5.97 48,987
1125 23.62 1.000 989 5.77 49,684
1499 25.43 1.000 989 5.86 117,359

Full per-iteration metrics: train_curve.json.

Files

file what
model_1499.pt final rsl_rl checkpoint (OnPolicyRunner.load) — actor + critic + optimizer
exported/policy.pt TorchScript actor incl. observation normalizer (deterministic mean action); input (N, 256) → (N, 19)
exported/policy.onnx (+ .onnx.data) the same actor as ONNX
params/env.yaml, params/agent.yaml exact Isaac Lab env + rsl_rl agent configs of the run (seed 1)
train_curve.json · record.json · verify.json training curve · recording metadata (obs/action names) · dataset verification
examples/record_trained_policy.py the strands recording script used for the dataset
playback.mp4/.gif, frame.png media

How it was made with strands-robots

Setup

# Isaac Lab in its OWN venv (its pins clash with strands; strands never imports it)
uv venv --python 3.12 ~/il && uv pip install --python ~/il/bin/python --prerelease=allow \
  --index https://pypi.nvidia.com --index-strategy unsafe-best-match "isaaclab[rsl-rl,isaacsim]==3.0.0rc1"
export ISAACLAB_PYTHON=~/il/bin/python
export OMNI_KIT_ACCEPT_EULA=YES          # you accept the NVIDIA Omniverse / Isaac Sim EULA yourself
pip install "git+https://github.com/cagataycali/robots@feat/isaaclab-trainer"   # strands-robots with PR #4227

Train

As an agent tool call (the train_policy tool is a Strands @tool):

from strands import Agent
from strands_robots.tools.train_policy import train_policy

agent = Agent(tools=[train_policy])
agent("Train the Unitree H1 to walk on rough terrain with the isaaclab provider: task Isaac-Velocity-Rough-H1, 4096 envs, PhysX, 1500 iterations, seed 1.")
# -> train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c4_h1_rough",
#                 extra={"task": "Isaac-Velocity-Rough-H1", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800})

As plain Python (exactly what produced this run):

from strands_robots.tools.train_policy import train_policy

job = train_policy(action="train", provider="isaaclab", steps=1500, seed=1,
                   output_dir="runs/c4_h1_rough",
                   extra={"task": "Isaac-Velocity-Rough-H1", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800})
# poll: iteration, rewards, learning verdict, steps_per_s, checkpoint_dir
train_policy(action="status", provider="isaaclab", job_id="<job_id from the result>")

Under the hood the provider runs python -m isaaclab train --rl_library rsl_rl --task Isaac-Velocity-Rough-H1 --max_iterations 1500 --num_envs 4096 --seed 1 physics=isaacsim_physx in $ISAACLAB_PYTHON and parses its log. Job id of this run: isaaclab-20260929-054143-8d5e0dcdd86d. Docs: docs/learn/training/isaaclab.md · PR: strands-labs/robots#4227.

Record (dataset repo)

The final checkpoint was rolled out and recorded with examples/record_trained_policy.py (included in this repo). It runs in the Isaac Lab venv with strands on PYTHONPATH and:

  1. rebuilds the task env in play mode and adds an RTX camera per env;
  2. loads model_1499.pt with rsl_rl's OnPolicyRunner and exports TorchScript/ONNX with Isaac Lab's exporter;
  3. wraps the exported actor as a strands Policy (RslRlJitPolicy, max |Δa| vs rsl_rl inference = 1.8e-07);
  4. steps the env with policy.get_actions_sync(...) and writes every frame through strands DatasetRecorder (strands_robots.dataset_recorder.DatasetRecorder.create(...) → add_frame → save_episode, LeRobot v3);
  5. verifies the result with strands verify_dataset + LeRobotDataset load + video decode + NaN scan.
OMNI_KIT_ACCEPT_EULA=YES PYTHONPATH=/path/to/strands-robots $ISAACLAB_PYTHON examples/record_trained_policy.py \
  --task Isaac-Velocity-Rough-H1 --checkpoint model_1499.pt --episodes 12 --frames 500 \
  --cam chase --override physics=isaacsim_physx \
  --task_str "walk over rough terrain following the commanded base velocity" --robot_type unitree_h1 \
  --root out/ds --repo_id cagataydev/strands-isaaclab-h1-rough --attach Robot/pelvis
$ISAACLAB_PYTHON examples/record_trained_policy.py --verify out/ds --repo_id cagataydev/strands-isaaclab-h1-rough

Use it

Play it in Isaac Lab (in the Isaac Lab venv):

huggingface-cli download cagataydev/strands-isaaclab-h1-rough-policy --local-dir h1_policy
OMNI_KIT_ACCEPT_EULA=YES $ISAACLAB_PYTHON -m isaaclab play --rl_library rsl_rl --task Isaac-Velocity-Rough-H1 \
  --num_envs 16 --checkpoint h1_policy/model_1499.pt physics=isaacsim_physx   # PhysX! (IL-X-011); add --video --video_length 500

Raw TorchScript actor:

import torch
pi = torch.jit.load("h1_policy/exported/policy.pt").eval()
actions = pi(obs)          # obs: (N, 256) concatenated Isaac Lab 'policy' observation group -> (N, 19)

As a strands Policy. create_policy("rl") loads this repo directly (strands-robots main after strands-labs/robots#4440; before it, wrap exported/policy.pt as the adapter in examples/record_trained_policy.py does). The rsl_rl actor (model_1499.pt, ELU layers, activation read from params/agent.yaml) is rebuilt once into a strands checkpoint beside the snapshot, and the 19 action_names in record.json bind the outputs:

from strands_robots.policies import create_policy

policy = create_policy("rl", checkpoint_dir="cagataydev/strands-isaaclab-h1-rough-policy")          # or a local run dir / model_1499.pt path
chunk = policy.get_actions_sync({"policy_obs": obs_256}, "walk forward")   # [{"joint_pos.left_hip_yaw": ..., 19 keys in record.json order}]

Measured against torch.jit.load("exported/policy.pt") on 300 random observations of width 256 (three scales) plus one get_actions_sync call: max |delta| = 0.0 (same float32 weights, same Linear/ELU chain; this run trained with obs_normalization: false, so the exported wrapper's normalizer is an identity).

The observation has to come from the Isaac Lab task (base velocities, gravity, velocity command, joint states, last action, height scan …), so the policy runs inside Isaac Lab; see examples/record_trained_policy.py for the full env + camera + DatasetRecorder loop.

Provenance

  • strands-robots: feat/isaaclab-trainer @ fa66fc68 — strands-labs/robots#4227 (isaaclab train_policy provider, IsaacLabTrainer; DatasetRecorder; verify_dataset)
  • Isaac Lab 3.0.0rc1 · Isaac Sim 6.1.0.0 · PhysX (physics=isaacsim_physx) · rsl-rl-lib 5.4.1 (PPO) · lerobot 0.6.1 · torch on CUDA
  • GPU: 1× NVIDIA L40S (46 GB), shared with the Go2 rough-terrain run for all of training
  • Seeds: training seed 1 (params/agent.yaml, params/env.yaml); recording seed 7
  • Training job: isaaclab-20260929-054143-8d5e0dcdd86d, 2026-09-29

Limitations

  • Simulation only. Nothing here was run on a real Unitree H1; no sim-to-real claims (the policy was not trained with sim-to-real hardening beyond Isaac Lab's default randomization).
  • Release candidates: Isaac Lab 3.0.0rc1 on Isaac Sim 6.1.0.0; APIs and physics may change. PhysX and Newton results differ.
  • Physics preset matters (IL-X-011): trained and recorded on PhysX only; always pass physics=isaacsim_physx when playing it (the provider does not remember the preset; the G1 PhysX policy falls in < 1.1 s when replayed on Newton).
  • Recording: 12 parallel envs from a common reset, 500 frames (10 s) each with a chase camera; all 12 ran the full window. observation.state mixes joint positions with the policy observation (incl. height scan), so it is wide (282-D) and not a standard LeRobot "robot state".
  • create_policy("rl") in strands cannot load this rsl_rl checkpoint yet (finding IL-X-006: strands' RL actor is Tanh, rsl_rl is ELU
    • obs-normalizer); use the exported TorchScript + the small wrapper shown above.
  • Known provider findings tracked with the PR: IL-X-001 (NaN reward not surfaced), IL-X-004/005 (no stop / play action), IL-X-007 (status text hid a crash traceback).

License

Card choice: license: other — our generated data / weights under CC-BY-4.0, plus NVIDIA notices. Why:

  • The recorded trajectories, rendered camera video, playback clips and the trained policy weights are user-generated content produced with NVIDIA Isaac Sim / Isaac Lab. The NVIDIA Omniverse License Agreement (which governs Isaac Sim 6.1, shipped as isaacsim/LICENSE.txt) §2.1 explicitly allows you to "distribute user generated content that you develop using Omniverse, such as video, audio, stills, models, 3D assets and screen captures". We release that content under CC-BY-4.0.
  • No NVIDIA Content is redistributed: the Unitree H1 USD (IsaacLab/Robots/Unitree/H1/h1_minimal.usd) and scene assets come from the Isaac Lab / Isaac Sim asset packs on NVIDIA's asset server and are not in this repo; params/env.yaml only references their paths. To reproduce you download them under your own NVIDIA EULA acceptance. The rough terrain is procedurally generated by Isaac Lab. "Unitree H1" is a product of Unitree Robotics; no endorsement by Unitree or NVIDIA is implied.
  • params/*.yaml are Isaac Lab task / agent configurations (Isaac Lab is BSD-3-Clause); the example script is Apache-2.0 like strands-robots. Running Isaac Sim itself requires accepting the NVIDIA Isaac Sim / Omniverse EULA.
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Dataset used to train cagataydev/strands-isaaclab-h1-rough-policy

Collection including cagataydev/strands-isaaclab-h1-rough-policy

Evaluation results

  • velocity-tracking success rate (last training iteration) on Isaac Lab Isaac-Velocity-Rough-H1 (4096 envs, PhysX)
    self-reported
    1.000
  • mean episode reward (last iteration) on Isaac Lab Isaac-Velocity-Rough-H1 (4096 envs, PhysX)
    self-reported
    25.430