--- license: other license_name: cc-by-4.0-generated-with-nvidia-isaac license_link: LICENSE.md library_name: rsl_rl pipeline_tag: robotics tags: - isaaclab - isaac-sim - strands-robots - lerobot - robotics - reinforcement-learning - rsl_rl - ppo - cassie - biped - locomotion - rough-terrain datasets: - cagataydev/strands-isaaclab-cassie-rough model-index: - name: strands-isaaclab-cassie-rough-policy results: - task: type: reinforcement-learning name: Rough-terrain velocity-tracking locomotion dataset: type: Isaac-Velocity-Rough-Cassie name: Isaac Lab Isaac-Velocity-Rough-Cassie (4096 envs, PhysX) metrics: - type: success_rate value: 1.000 name: velocity-tracking success rate (last training iteration) - type: mean_reward value: 30.37 name: mean episode reward (last iteration) --- # strands-isaaclab-cassie-rough-policy **rsl_rl PPO policy for Agility Cassie biped rough-terrain locomotion (`Isaac-Velocity-Rough-Cassie`), trained with strands-robots' `isaaclab` train_policy provider ([PR #4227](https://github.com/strands-labs/robots/pull/4227)).** Rollouts recorded through strands: [cagataydev/strands-isaaclab-cassie-rough](https://huggingface.co/datasets/cagataydev/strands-isaaclab-cassie-rough). ![playback: 4 parallel envs](playback.gif) *envs 0–3 (2×2) from the strands camera — [mp4](playback.mp4).* ## Results | | | |---|---| | parallel envs × iterations | **4096 × 1500** (147 M env steps), PhysX (`physics=isaacsim_physx`) | | wall time | **63 min 38 s** on 1× NVIDIA L40S | | env-steps/s | median **48 k**, max 95 k (GPU shared with other runs most of the time) | | mean reward | -6.71 → **33.7** best → 30.4 last | | velocity-tracking success | **1.00** last iteration; mean episode length 957 of 1000; terrain level 5.8 | | recorded rollout | 12 / 12 envs walked the full 10 s window; mean return 23.4 | | PPO iteration | mean reward | velocity-tracking success | mean ep. length (of 1000) | terrain curriculum level | env-steps/s | |---|---|---|---|---|---| | 0 | -6.71 | 0.050 | 24 | 3.50 | 39,565 | | 150 | -0.27 | 0.126 | 771 | 0.04 | 43,755 | | 375 | 15.96 | 0.867 | 932 | 2.92 | 47,451 | | 750 | 24.61 | 1.000 | 965 | 6.12 | 50,011 | | 1125 | 30.66 | 1.000 | 998 | 5.87 | 49,517 | | 1499 | 30.37 | 1.000 | 957 | 5.79 | 48,889 | Full per-iteration metrics: [`train_curve.json`](train_curve.json). ## Files | file | what | |---|---| | `model_1499.pt` | final rsl_rl checkpoint (`OnPolicyRunner.load`) — actor + critic + optimizer | | `exported/policy.pt` | TorchScript actor **incl. observation normalizer** (deterministic mean action); input `(N, 235)` → `(N, 12)` | | `exported/policy.onnx` (+ `.onnx.data`) | the same actor as ONNX | | `params/env.yaml`, `params/agent.yaml` | exact Isaac Lab env + rsl_rl agent configs of the run (seed 1) | | `train_curve.json` · `record.json` · `verify.json` | training curve · recording metadata (obs/action names) · dataset verification | | `examples/record_trained_policy.py` | the strands recording script used for the dataset | | `playback.mp4/.gif`, `frame.png` | media | ## How it was made with strands-robots **Setup** ```bash # Isaac Lab in its OWN venv (its pins clash with strands; strands never imports it) uv venv --python 3.12 ~/il && uv pip install --python ~/il/bin/python --prerelease=allow \ --index https://pypi.nvidia.com --index-strategy unsafe-best-match "isaaclab[rsl-rl,isaacsim]==3.0.0rc1" export ISAACLAB_PYTHON=~/il/bin/python export OMNI_KIT_ACCEPT_EULA=YES # you accept the NVIDIA Omniverse / Isaac Sim EULA yourself pip install "git+https://github.com/cagataycali/robots@feat/isaaclab-trainer" # strands-robots with PR #4227 ``` **Train** **As an agent tool call** (the `train_policy` tool is a Strands `@tool`): ```python from strands import Agent from strands_robots.tools.train_policy import train_policy agent = Agent(tools=[train_policy]) agent("Train the Agility Cassie to walk on rough terrain with the isaaclab provider: task Isaac-Velocity-Rough-Cassie, 4096 envs, PhysX, 1500 iterations, seed 1.") # -> train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c4_cassie_rough", # extra={"task": "Isaac-Velocity-Rough-Cassie", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800}) ``` **As plain Python** (exactly what produced this run): ```python from strands_robots.tools.train_policy import train_policy job = train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c4_cassie_rough", extra={"task": "Isaac-Velocity-Rough-Cassie", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800}) # poll: iteration, rewards, learning verdict, steps_per_s, checkpoint_dir train_policy(action="status", provider="isaaclab", job_id="") ``` Under the hood the provider runs `python -m isaaclab train --rl_library rsl_rl --task Isaac-Velocity-Rough-Cassie --max_iterations 1500 --num_envs 4096 --seed 1 physics=isaacsim_physx` in `$ISAACLAB_PYTHON` and parses its log. Job id of this run: `isaaclab-20260929-064750-053b21124489`. Docs: [`docs/learn/training/isaaclab.md`](https://github.com/cagataycali/robots/blob/feat/isaaclab-trainer/docs/learn/training/isaaclab.md) · PR: [strands-labs/robots#4227](https://github.com/strands-labs/robots/pull/4227). **Record** (dataset repo) The final checkpoint was rolled out and recorded with [`examples/record_trained_policy.py`](examples/record_trained_policy.py) (included in this repo). It runs in the Isaac Lab venv with strands on `PYTHONPATH` and: 1. rebuilds the task env in play mode and adds an RTX camera per env; 2. loads `model_1499.pt` with rsl_rl's `OnPolicyRunner` and exports TorchScript/ONNX with Isaac Lab's exporter; 3. wraps the exported actor as a **strands `Policy`** (`RslRlJitPolicy`, max |Δa| vs rsl_rl inference = 1.5e-07); 4. steps the env with `policy.get_actions_sync(...)` and writes every frame through **strands `DatasetRecorder`** (`strands_robots.dataset_recorder.DatasetRecorder.create(...)` → `add_frame` → `save_episode`, LeRobot v3); 5. verifies the result with strands `verify_dataset` + `LeRobotDataset` load + video decode + NaN scan. ```bash OMNI_KIT_ACCEPT_EULA=YES PYTHONPATH=/path/to/strands-robots $ISAACLAB_PYTHON examples/record_trained_policy.py \ --task Isaac-Velocity-Rough-Cassie --checkpoint model_1499.pt --episodes 12 --frames 500 \ --cam chase --override physics=isaacsim_physx \ --task_str "walk over rough terrain following the commanded base velocity" --robot_type cassie \ --root out/ds --repo_id cagataydev/strands-isaaclab-cassie-rough --attach Robot/pelvis $ISAACLAB_PYTHON examples/record_trained_policy.py --verify out/ds --repo_id cagataydev/strands-isaaclab-cassie-rough ``` ## Use it **Play it in Isaac Lab** (in the Isaac Lab venv): ```bash huggingface-cli download cagataydev/strands-isaaclab-cassie-rough-policy --local-dir cassie_policy OMNI_KIT_ACCEPT_EULA=YES $ISAACLAB_PYTHON -m isaaclab play --rl_library rsl_rl --task Isaac-Velocity-Rough-Cassie \ --num_envs 16 --checkpoint cassie_policy/model_1499.pt physics=isaacsim_physx # PhysX! (IL-X-011); add --video --video_length 500 ``` **Raw TorchScript actor**: ```python import torch pi = torch.jit.load("cassie_policy/exported/policy.pt").eval() actions = pi(obs) # obs: (N, 235) concatenated Isaac Lab 'policy' observation group -> (N, 12) ``` **As a strands `Policy`.** `create_policy("rl")` **cannot load rsl_rl checkpoints yet** (IL-X-006: strands' RL actor is a Tanh MLP, rsl_rl's is ELU with a baked-in normalizer), so wrap the exported actor — this is the adapter the recording script uses: ```python import numpy as np, torch from strands_robots.policies.base import Policy class RslRlJitPolicy(Policy): def __init__(self, jit_path, action_names, device="cpu"): self.net = torch.jit.load(jit_path, map_location=device).eval() self.device, self.action_names, self.robot_state_keys = device, list(action_names), [] @property def provider_name(self): return "isaaclab_rsl_rl_jit" def set_robot_state_keys(self, keys): self.robot_state_keys = list(keys) def reset(self, seed=None): self.net.reset() async def get_actions(self, observation_dict, instruction, **kw): x = torch.as_tensor(np.asarray(observation_dict["policy_obs"], np.float32), device=self.device) with torch.inference_mode(): y = self.net(x.unsqueeze(0))[0].cpu().numpy() return [dict(zip(self.action_names, map(float, y)))] import json names = json.load(open("cassie_policy/record.json"))["action_names"] policy = RslRlJitPolicy("cassie_policy/exported/policy.pt", names) chunk = policy.get_actions_sync({"policy_obs": obs_235}, "walk forward") # [{"joint_pos.hip_abduction_left": ..., ...}] ``` The observation has to come from the Isaac Lab task (base velocities, gravity, velocity command, joint states, last action, height scan …), so the policy runs inside Isaac Lab; see [`examples/record_trained_policy.py`](examples/record_trained_policy.py) for the full env + camera + `DatasetRecorder` loop. ## Provenance - **strands-robots**: `feat/isaaclab-trainer` @ [`fa66fc68`](https://github.com/cagataycali/robots/commit/fa66fc682047f3cb70e2e2bc60f0b2dde2fbebae) — [strands-labs/robots#4227](https://github.com/strands-labs/robots/pull/4227) (`isaaclab` train_policy provider, `IsaacLabTrainer`; `DatasetRecorder`; `verify_dataset`) - **Isaac Lab** 3.0.0rc1 · **Isaac Sim** 6.1.0.0 · PhysX (`physics=isaacsim_physx`) · rsl-rl-lib 5.4.1 (PPO) · lerobot 0.6.1 · torch on CUDA - **GPU**: 1× NVIDIA L40S (46 GB), shared with the ANYmal-D rough-terrain run for all of training - **Seeds**: training seed 1 (`params/agent.yaml`, `params/env.yaml`); recording seed 7 - **Training job**: `isaaclab-20260929-064750-053b21124489`, 2026-09-29 ## Limitations - **Simulation only.** Nothing here was run on a real Agility Cassie; no sim-to-real claims (the policy was not trained with sim-to-real hardening beyond Isaac Lab's default randomization). - **Release candidates**: Isaac Lab 3.0.0rc1 on Isaac Sim 6.1.0.0; APIs and physics may change. PhysX and Newton results differ. - **Physics preset matters (IL-X-011):** trained and recorded on PhysX only; always pass `physics=isaacsim_physx` when playing it (the provider does not remember the preset; the G1 PhysX policy falls in < 1.1 s when replayed on Newton). - Recording: 12 parallel envs from a common reset, 500 frames (10 s) each with a chase camera; all 12 ran the full window. `observation.state` mixes joint positions with the policy observation (incl. height scan), so it is wide (254-D) and not a standard LeRobot "robot state". - `create_policy("rl")` in strands **cannot load this rsl_rl checkpoint yet** (finding IL-X-006: strands' RL actor is Tanh, rsl_rl is ELU + obs-normalizer); use the exported TorchScript + the small wrapper shown above. - Known provider findings tracked with the PR: IL-X-001 (NaN reward not surfaced), IL-X-004/005 (no stop / play action), IL-X-007 (status text hid a crash traceback). ## License **Card choice: `license: other` — our generated data / weights under CC-BY-4.0, plus NVIDIA notices.** Why: - The recorded trajectories, rendered camera video, playback clips and the trained policy weights are *user-generated content* produced with NVIDIA Isaac Sim / Isaac Lab. The NVIDIA Omniverse License Agreement (which governs Isaac Sim 6.1, shipped as `isaacsim/LICENSE.txt`) §2.1 explicitly allows you to *"distribute user generated content that you develop using Omniverse, such as video, audio, stills, models, 3D assets and screen captures"*. We release that content under **CC-BY-4.0**. - **No NVIDIA Content is redistributed**: the Agility Cassie USD (`IsaacLab/Robots/Agility/Cassie/cassie.usd`) and scene assets come from the Isaac Lab / Isaac Sim asset packs on NVIDIA's asset server and are **not** in this repo; `params/env.yaml` only references their paths. To reproduce you download them under your own NVIDIA EULA acceptance. The rough terrain is procedurally generated by Isaac Lab. "Agility Cassie" is a product of Agility Robotics; no endorsement by Agility or NVIDIA is implied. - `params/*.yaml` are Isaac Lab task / agent configurations (Isaac Lab is **BSD-3-Clause**); the example script is Apache-2.0 like strands-robots. Running Isaac Sim itself requires accepting the NVIDIA Isaac Sim / Omniverse EULA.