Instructions to use cagataydev/strands-isaaclab-cassie-rough-policy with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use cagataydev/strands-isaaclab-cassie-rough-policy with LeRobot:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
strands-isaaclab-cassie-rough-policy
rsl_rl PPO policy for Agility Cassie biped rough-terrain locomotion (Isaac-Velocity-Rough-Cassie), trained with strands-robots' isaaclab
train_policy provider (PR #4227). Rollouts recorded through strands: cagataydev/strands-isaaclab-cassie-rough.
envs 0–3 (2×2) from the strands camera — mp4.
Results
| parallel envs × iterations | 4096 × 1500 (147 M env steps), PhysX (physics=isaacsim_physx) |
| wall time | 63 min 38 s on 1× NVIDIA L40S |
| env-steps/s | median 48 k, max 95 k (GPU shared with other runs most of the time) |
| mean reward | -6.71 → 33.7 best → 30.4 last |
| velocity-tracking success | 1.00 last iteration; mean episode length 957 of 1000; terrain level 5.8 |
| recorded rollout | 12 / 12 envs walked the full 10 s window; mean return 23.4 |
| PPO iteration | mean reward | velocity-tracking success | mean ep. length (of 1000) | terrain curriculum level | env-steps/s |
|---|---|---|---|---|---|
| 0 | -6.71 | 0.050 | 24 | 3.50 | 39,565 |
| 150 | -0.27 | 0.126 | 771 | 0.04 | 43,755 |
| 375 | 15.96 | 0.867 | 932 | 2.92 | 47,451 |
| 750 | 24.61 | 1.000 | 965 | 6.12 | 50,011 |
| 1125 | 30.66 | 1.000 | 998 | 5.87 | 49,517 |
| 1499 | 30.37 | 1.000 | 957 | 5.79 | 48,889 |
Full per-iteration metrics: train_curve.json.
Files
| file | what |
|---|---|
model_1499.pt |
final rsl_rl checkpoint (OnPolicyRunner.load) — actor + critic + optimizer |
exported/policy.pt |
TorchScript actor incl. observation normalizer (deterministic mean action); input (N, 235) → (N, 12) |
exported/policy.onnx (+ .onnx.data) |
the same actor as ONNX |
params/env.yaml, params/agent.yaml |
exact Isaac Lab env + rsl_rl agent configs of the run (seed 1) |
train_curve.json · record.json · verify.json |
training curve · recording metadata (obs/action names) · dataset verification |
examples/record_trained_policy.py |
the strands recording script used for the dataset |
playback.mp4/.gif, frame.png |
media |
How it was made with strands-robots
Setup
# Isaac Lab in its OWN venv (its pins clash with strands; strands never imports it)
uv venv --python 3.12 ~/il && uv pip install --python ~/il/bin/python --prerelease=allow \
--index https://pypi.nvidia.com --index-strategy unsafe-best-match "isaaclab[rsl-rl,isaacsim]==3.0.0rc1"
export ISAACLAB_PYTHON=~/il/bin/python
export OMNI_KIT_ACCEPT_EULA=YES # you accept the NVIDIA Omniverse / Isaac Sim EULA yourself
pip install "git+https://github.com/cagataycali/robots@feat/isaaclab-trainer" # strands-robots with PR #4227
Train
As an agent tool call (the train_policy tool is a Strands @tool):
from strands import Agent
from strands_robots.tools.train_policy import train_policy
agent = Agent(tools=[train_policy])
agent("Train the Agility Cassie to walk on rough terrain with the isaaclab provider: task Isaac-Velocity-Rough-Cassie, 4096 envs, PhysX, 1500 iterations, seed 1.")
# -> train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c4_cassie_rough",
# extra={"task": "Isaac-Velocity-Rough-Cassie", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800})
As plain Python (exactly what produced this run):
from strands_robots.tools.train_policy import train_policy
job = train_policy(action="train", provider="isaaclab", steps=1500, seed=1,
output_dir="runs/c4_cassie_rough",
extra={"task": "Isaac-Velocity-Rough-Cassie", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800})
# poll: iteration, rewards, learning verdict, steps_per_s, checkpoint_dir
train_policy(action="status", provider="isaaclab", job_id="<job_id from the result>")
Under the hood the provider runs python -m isaaclab train --rl_library rsl_rl --task Isaac-Velocity-Rough-Cassie --max_iterations 1500 --num_envs 4096 --seed 1 physics=isaacsim_physx
in $ISAACLAB_PYTHON and parses its log. Job id of this run: isaaclab-20260929-064750-053b21124489.
Docs: docs/learn/training/isaaclab.md · PR: strands-labs/robots#4227.
Record (dataset repo)
The final checkpoint was rolled out and recorded with examples/record_trained_policy.py
(included in this repo). It runs in the Isaac Lab venv with strands on PYTHONPATH and:
- rebuilds the task env in play mode and adds an RTX camera per env;
- loads
model_1499.ptwith rsl_rl'sOnPolicyRunnerand exports TorchScript/ONNX with Isaac Lab's exporter; - wraps the exported actor as a strands
Policy(RslRlJitPolicy, max |Δa| vs rsl_rl inference = 1.5e-07); - steps the env with
policy.get_actions_sync(...)and writes every frame through strandsDatasetRecorder(strands_robots.dataset_recorder.DatasetRecorder.create(...)→add_frame→save_episode, LeRobot v3); - verifies the result with strands
verify_dataset+LeRobotDatasetload + video decode + NaN scan.
OMNI_KIT_ACCEPT_EULA=YES PYTHONPATH=/path/to/strands-robots $ISAACLAB_PYTHON examples/record_trained_policy.py \
--task Isaac-Velocity-Rough-Cassie --checkpoint model_1499.pt --episodes 12 --frames 500 \
--cam chase --override physics=isaacsim_physx \
--task_str "walk over rough terrain following the commanded base velocity" --robot_type cassie \
--root out/ds --repo_id cagataydev/strands-isaaclab-cassie-rough --attach Robot/pelvis
$ISAACLAB_PYTHON examples/record_trained_policy.py --verify out/ds --repo_id cagataydev/strands-isaaclab-cassie-rough
Use it
Play it in Isaac Lab (in the Isaac Lab venv):
huggingface-cli download cagataydev/strands-isaaclab-cassie-rough-policy --local-dir cassie_policy
OMNI_KIT_ACCEPT_EULA=YES $ISAACLAB_PYTHON -m isaaclab play --rl_library rsl_rl --task Isaac-Velocity-Rough-Cassie \
--num_envs 16 --checkpoint cassie_policy/model_1499.pt physics=isaacsim_physx # PhysX! (IL-X-011); add --video --video_length 500
Raw TorchScript actor:
import torch
pi = torch.jit.load("cassie_policy/exported/policy.pt").eval()
actions = pi(obs) # obs: (N, 235) concatenated Isaac Lab 'policy' observation group -> (N, 12)
As a strands Policy. create_policy("rl") cannot load rsl_rl checkpoints yet (IL-X-006: strands' RL actor is a Tanh MLP,
rsl_rl's is ELU with a baked-in normalizer), so wrap the exported actor — this is the adapter the recording script uses:
import numpy as np, torch
from strands_robots.policies.base import Policy
class RslRlJitPolicy(Policy):
def __init__(self, jit_path, action_names, device="cpu"):
self.net = torch.jit.load(jit_path, map_location=device).eval()
self.device, self.action_names, self.robot_state_keys = device, list(action_names), []
@property
def provider_name(self): return "isaaclab_rsl_rl_jit"
def set_robot_state_keys(self, keys): self.robot_state_keys = list(keys)
def reset(self, seed=None): self.net.reset()
async def get_actions(self, observation_dict, instruction, **kw):
x = torch.as_tensor(np.asarray(observation_dict["policy_obs"], np.float32), device=self.device)
with torch.inference_mode():
y = self.net(x.unsqueeze(0))[0].cpu().numpy()
return [dict(zip(self.action_names, map(float, y)))]
import json
names = json.load(open("cassie_policy/record.json"))["action_names"]
policy = RslRlJitPolicy("cassie_policy/exported/policy.pt", names)
chunk = policy.get_actions_sync({"policy_obs": obs_235}, "walk forward") # [{"joint_pos.hip_abduction_left": ..., ...}]
The observation has to come from the Isaac Lab task (base velocities, gravity, velocity command, joint states, last action, height scan …), so the
policy runs inside Isaac Lab; see examples/record_trained_policy.py for the full
env + camera + DatasetRecorder loop.
Provenance
- strands-robots:
feat/isaaclab-trainer@fa66fc68— strands-labs/robots#4227 (isaaclabtrain_policy provider,IsaacLabTrainer;DatasetRecorder;verify_dataset) - Isaac Lab 3.0.0rc1 · Isaac Sim 6.1.0.0 · PhysX (
physics=isaacsim_physx) · rsl-rl-lib 5.4.1 (PPO) · lerobot 0.6.1 · torch on CUDA - GPU: 1× NVIDIA L40S (46 GB), shared with the ANYmal-D rough-terrain run for all of training
- Seeds: training seed 1 (
params/agent.yaml,params/env.yaml); recording seed 7 - Training job:
isaaclab-20260929-064750-053b21124489, 2026-09-29
Limitations
- Simulation only. Nothing here was run on a real Agility Cassie; no sim-to-real claims (the policy was not trained with sim-to-real hardening beyond Isaac Lab's default randomization).
- Release candidates: Isaac Lab 3.0.0rc1 on Isaac Sim 6.1.0.0; APIs and physics may change. PhysX and Newton results differ.
- Physics preset matters (IL-X-011): trained and recorded on PhysX only; always pass
physics=isaacsim_physxwhen playing it (the provider does not remember the preset; the G1 PhysX policy falls in < 1.1 s when replayed on Newton). - Recording: 12 parallel envs from a common reset, 500 frames (10 s) each with a chase camera; all 12 ran the full window.
observation.statemixes joint positions with the policy observation (incl. height scan), so it is wide (254-D) and not a standard LeRobot "robot state". create_policy("rl")in strands cannot load this rsl_rl checkpoint yet (finding IL-X-006: strands' RL actor is Tanh, rsl_rl is ELU- obs-normalizer); use the exported TorchScript + the small wrapper shown above.
- Known provider findings tracked with the PR: IL-X-001 (NaN reward not surfaced), IL-X-004/005 (no stop / play action), IL-X-007 (status text hid a crash traceback).
License
Card choice: license: other — our generated data / weights under CC-BY-4.0, plus NVIDIA notices. Why:
- The recorded trajectories, rendered camera video, playback clips and the trained policy weights are user-generated content
produced with NVIDIA Isaac Sim / Isaac Lab. The NVIDIA Omniverse License Agreement (which governs Isaac Sim 6.1, shipped as
isaacsim/LICENSE.txt) §2.1 explicitly allows you to "distribute user generated content that you develop using Omniverse, such as video, audio, stills, models, 3D assets and screen captures". We release that content under CC-BY-4.0. - No NVIDIA Content is redistributed: the Agility Cassie USD (
IsaacLab/Robots/Agility/Cassie/cassie.usd) and scene assets come from the Isaac Lab / Isaac Sim asset packs on NVIDIA's asset server and are not in this repo;params/env.yamlonly references their paths. To reproduce you download them under your own NVIDIA EULA acceptance. The rough terrain is procedurally generated by Isaac Lab. "Agility Cassie" is a product of Agility Robotics; no endorsement by Agility or NVIDIA is implied. params/*.yamlare Isaac Lab task / agent configurations (Isaac Lab is BSD-3-Clause); the example script is Apache-2.0 like strands-robots. Running Isaac Sim itself requires accepting the NVIDIA Isaac Sim / Omniverse EULA.
Dataset used to train cagataydev/strands-isaaclab-cassie-rough-policy
Collection including cagataydev/strands-isaaclab-cassie-rough-policy
Evaluation results
- velocity-tracking success rate (last training iteration) on Isaac Lab Isaac-Velocity-Rough-Cassie (4096 envs, PhysX)self-reported1.000
- mean episode reward (last iteration) on Isaac Lab Isaac-Velocity-Rough-Cassie (4096 envs, PhysX)self-reported30.370
