strands-isaaclab-cassie-rough-policy

rsl_rl PPO policy for Agility Cassie biped rough-terrain locomotion (Isaac-Velocity-Rough-Cassie), trained with strands-robots' isaaclab train_policy provider (PR #4227). Rollouts recorded through strands: cagataydev/strands-isaaclab-cassie-rough.

playback: 4 parallel envs

envs 0–3 (2×2) from the strands camera — mp4.

Results

parallel envs × iterations 4096 × 1500 (147 M env steps), PhysX (physics=isaacsim_physx)
wall time 63 min 38 s on 1× NVIDIA L40S
env-steps/s median 48 k, max 95 k (GPU shared with other runs most of the time)
mean reward -6.71 → 33.7 best → 30.4 last
velocity-tracking success 1.00 last iteration; mean episode length 957 of 1000; terrain level 5.8
recorded rollout 12 / 12 envs walked the full 10 s window; mean return 23.4
PPO iteration mean reward velocity-tracking success mean ep. length (of 1000) terrain curriculum level env-steps/s
0 -6.71 0.050 24 3.50 39,565
150 -0.27 0.126 771 0.04 43,755
375 15.96 0.867 932 2.92 47,451
750 24.61 1.000 965 6.12 50,011
1125 30.66 1.000 998 5.87 49,517
1499 30.37 1.000 957 5.79 48,889

Full per-iteration metrics: train_curve.json.

Files

file what
model_1499.pt final rsl_rl checkpoint (OnPolicyRunner.load) — actor + critic + optimizer
exported/policy.pt TorchScript actor incl. observation normalizer (deterministic mean action); input (N, 235) → (N, 12)
exported/policy.onnx (+ .onnx.data) the same actor as ONNX
params/env.yaml, params/agent.yaml exact Isaac Lab env + rsl_rl agent configs of the run (seed 1)
train_curve.json · record.json · verify.json training curve · recording metadata (obs/action names) · dataset verification
examples/record_trained_policy.py the strands recording script used for the dataset
playback.mp4/.gif, frame.png media

How it was made with strands-robots

Setup

# Isaac Lab in its OWN venv (its pins clash with strands; strands never imports it)
uv venv --python 3.12 ~/il && uv pip install --python ~/il/bin/python --prerelease=allow \
  --index https://pypi.nvidia.com --index-strategy unsafe-best-match "isaaclab[rsl-rl,isaacsim]==3.0.0rc1"
export ISAACLAB_PYTHON=~/il/bin/python
export OMNI_KIT_ACCEPT_EULA=YES          # you accept the NVIDIA Omniverse / Isaac Sim EULA yourself
pip install "git+https://github.com/cagataycali/robots@feat/isaaclab-trainer"   # strands-robots with PR #4227

Train

As an agent tool call (the train_policy tool is a Strands @tool):

from strands import Agent
from strands_robots.tools.train_policy import train_policy

agent = Agent(tools=[train_policy])
agent("Train the Agility Cassie to walk on rough terrain with the isaaclab provider: task Isaac-Velocity-Rough-Cassie, 4096 envs, PhysX, 1500 iterations, seed 1.")
# -> train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c4_cassie_rough",
#                 extra={"task": "Isaac-Velocity-Rough-Cassie", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800})

As plain Python (exactly what produced this run):

from strands_robots.tools.train_policy import train_policy

job = train_policy(action="train", provider="isaaclab", steps=1500, seed=1,
                   output_dir="runs/c4_cassie_rough",
                   extra={"task": "Isaac-Velocity-Rough-Cassie", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800})
# poll: iteration, rewards, learning verdict, steps_per_s, checkpoint_dir
train_policy(action="status", provider="isaaclab", job_id="<job_id from the result>")

Under the hood the provider runs python -m isaaclab train --rl_library rsl_rl --task Isaac-Velocity-Rough-Cassie --max_iterations 1500 --num_envs 4096 --seed 1 physics=isaacsim_physx in $ISAACLAB_PYTHON and parses its log. Job id of this run: isaaclab-20260929-064750-053b21124489. Docs: docs/learn/training/isaaclab.md · PR: strands-labs/robots#4227.

Record (dataset repo)

The final checkpoint was rolled out and recorded with examples/record_trained_policy.py (included in this repo). It runs in the Isaac Lab venv with strands on PYTHONPATH and:

  1. rebuilds the task env in play mode and adds an RTX camera per env;
  2. loads model_1499.pt with rsl_rl's OnPolicyRunner and exports TorchScript/ONNX with Isaac Lab's exporter;
  3. wraps the exported actor as a strands Policy (RslRlJitPolicy, max |Δa| vs rsl_rl inference = 1.5e-07);
  4. steps the env with policy.get_actions_sync(...) and writes every frame through strands DatasetRecorder (strands_robots.dataset_recorder.DatasetRecorder.create(...) → add_frame → save_episode, LeRobot v3);
  5. verifies the result with strands verify_dataset + LeRobotDataset load + video decode + NaN scan.
OMNI_KIT_ACCEPT_EULA=YES PYTHONPATH=/path/to/strands-robots $ISAACLAB_PYTHON examples/record_trained_policy.py \
  --task Isaac-Velocity-Rough-Cassie --checkpoint model_1499.pt --episodes 12 --frames 500 \
  --cam chase --override physics=isaacsim_physx \
  --task_str "walk over rough terrain following the commanded base velocity" --robot_type cassie \
  --root out/ds --repo_id cagataydev/strands-isaaclab-cassie-rough --attach Robot/pelvis
$ISAACLAB_PYTHON examples/record_trained_policy.py --verify out/ds --repo_id cagataydev/strands-isaaclab-cassie-rough

Use it

Play it in Isaac Lab (in the Isaac Lab venv):

huggingface-cli download cagataydev/strands-isaaclab-cassie-rough-policy --local-dir cassie_policy
OMNI_KIT_ACCEPT_EULA=YES $ISAACLAB_PYTHON -m isaaclab play --rl_library rsl_rl --task Isaac-Velocity-Rough-Cassie \
  --num_envs 16 --checkpoint cassie_policy/model_1499.pt physics=isaacsim_physx   # PhysX! (IL-X-011); add --video --video_length 500

Raw TorchScript actor:

import torch
pi = torch.jit.load("cassie_policy/exported/policy.pt").eval()
actions = pi(obs)          # obs: (N, 235) concatenated Isaac Lab 'policy' observation group -> (N, 12)

As a strands Policy. create_policy("rl") cannot load rsl_rl checkpoints yet (IL-X-006: strands' RL actor is a Tanh MLP, rsl_rl's is ELU with a baked-in normalizer), so wrap the exported actor — this is the adapter the recording script uses:

import numpy as np, torch
from strands_robots.policies.base import Policy

class RslRlJitPolicy(Policy):
    def __init__(self, jit_path, action_names, device="cpu"):
        self.net = torch.jit.load(jit_path, map_location=device).eval()
        self.device, self.action_names, self.robot_state_keys = device, list(action_names), []
    @property
    def provider_name(self): return "isaaclab_rsl_rl_jit"
    def set_robot_state_keys(self, keys): self.robot_state_keys = list(keys)
    def reset(self, seed=None): self.net.reset()
    async def get_actions(self, observation_dict, instruction, **kw):
        x = torch.as_tensor(np.asarray(observation_dict["policy_obs"], np.float32), device=self.device)
        with torch.inference_mode():
            y = self.net(x.unsqueeze(0))[0].cpu().numpy()
        return [dict(zip(self.action_names, map(float, y)))]

import json
names = json.load(open("cassie_policy/record.json"))["action_names"]
policy = RslRlJitPolicy("cassie_policy/exported/policy.pt", names)
chunk = policy.get_actions_sync({"policy_obs": obs_235}, "walk forward")   # [{"joint_pos.hip_abduction_left": ..., ...}]

The observation has to come from the Isaac Lab task (base velocities, gravity, velocity command, joint states, last action, height scan …), so the policy runs inside Isaac Lab; see examples/record_trained_policy.py for the full env + camera + DatasetRecorder loop.

Provenance

  • strands-robots: feat/isaaclab-trainer @ fa66fc68 — strands-labs/robots#4227 (isaaclab train_policy provider, IsaacLabTrainer; DatasetRecorder; verify_dataset)
  • Isaac Lab 3.0.0rc1 · Isaac Sim 6.1.0.0 · PhysX (physics=isaacsim_physx) · rsl-rl-lib 5.4.1 (PPO) · lerobot 0.6.1 · torch on CUDA
  • GPU: 1× NVIDIA L40S (46 GB), shared with the ANYmal-D rough-terrain run for all of training
  • Seeds: training seed 1 (params/agent.yaml, params/env.yaml); recording seed 7
  • Training job: isaaclab-20260929-064750-053b21124489, 2026-09-29

Limitations

  • Simulation only. Nothing here was run on a real Agility Cassie; no sim-to-real claims (the policy was not trained with sim-to-real hardening beyond Isaac Lab's default randomization).
  • Release candidates: Isaac Lab 3.0.0rc1 on Isaac Sim 6.1.0.0; APIs and physics may change. PhysX and Newton results differ.
  • Physics preset matters (IL-X-011): trained and recorded on PhysX only; always pass physics=isaacsim_physx when playing it (the provider does not remember the preset; the G1 PhysX policy falls in < 1.1 s when replayed on Newton).
  • Recording: 12 parallel envs from a common reset, 500 frames (10 s) each with a chase camera; all 12 ran the full window. observation.state mixes joint positions with the policy observation (incl. height scan), so it is wide (254-D) and not a standard LeRobot "robot state".
  • create_policy("rl") in strands cannot load this rsl_rl checkpoint yet (finding IL-X-006: strands' RL actor is Tanh, rsl_rl is ELU
    • obs-normalizer); use the exported TorchScript + the small wrapper shown above.
  • Known provider findings tracked with the PR: IL-X-001 (NaN reward not surfaced), IL-X-004/005 (no stop / play action), IL-X-007 (status text hid a crash traceback).

License

Card choice: license: other — our generated data / weights under CC-BY-4.0, plus NVIDIA notices. Why:

  • The recorded trajectories, rendered camera video, playback clips and the trained policy weights are user-generated content produced with NVIDIA Isaac Sim / Isaac Lab. The NVIDIA Omniverse License Agreement (which governs Isaac Sim 6.1, shipped as isaacsim/LICENSE.txt) §2.1 explicitly allows you to "distribute user generated content that you develop using Omniverse, such as video, audio, stills, models, 3D assets and screen captures". We release that content under CC-BY-4.0.
  • No NVIDIA Content is redistributed: the Agility Cassie USD (IsaacLab/Robots/Agility/Cassie/cassie.usd) and scene assets come from the Isaac Lab / Isaac Sim asset packs on NVIDIA's asset server and are not in this repo; params/env.yaml only references their paths. To reproduce you download them under your own NVIDIA EULA acceptance. The rough terrain is procedurally generated by Isaac Lab. "Agility Cassie" is a product of Agility Robotics; no endorsement by Agility or NVIDIA is implied.
  • params/*.yaml are Isaac Lab task / agent configurations (Isaac Lab is BSD-3-Clause); the example script is Apache-2.0 like strands-robots. Running Isaac Sim itself requires accepting the NVIDIA Isaac Sim / Omniverse EULA.
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Dataset used to train cagataydev/strands-isaaclab-cassie-rough-policy

Collection including cagataydev/strands-isaaclab-cassie-rough-policy

Evaluation results

  • velocity-tracking success rate (last training iteration) on Isaac Lab Isaac-Velocity-Rough-Cassie (4096 envs, PhysX)
    self-reported
    1.000
  • mean episode reward (last iteration) on Isaac Lab Isaac-Velocity-Rough-Cassie (4096 envs, PhysX)
    self-reported
    30.370