--- license: other license_name: cc-by-4.0-generated-with-nvidia-isaac license_link: LICENSE.md library_name: rsl_rl pipeline_tag: robotics tags: - isaaclab - isaac-sim - strands-robots - lerobot - robotics - reinforcement-learning - rsl_rl - ppo - unitree-g1 - humanoid - locomotion - rough-terrain datasets: - cagataydev/strands-isaaclab-g1-rough model-index: - name: strands-isaaclab-g1-rough-policy results: - task: type: reinforcement-learning name: Rough-terrain velocity-tracking locomotion dataset: type: Isaac-Velocity-Rough-G1 name: Isaac Lab Isaac-Velocity-Rough-G1 (4096 envs, PhysX) metrics: - type: success_rate value: 0.990 name: velocity-tracking success rate (last training iteration) - type: mean_reward value: 18.52 name: mean episode reward (last iteration) --- # strands-isaaclab-g1-rough-policy **rsl_rl PPO policy for Unitree G1 humanoid rough-terrain locomotion (`Isaac-Velocity-Rough-G1`), trained with strands-robots' `isaaclab` train_policy provider ([PR #4227](https://github.com/strands-labs/robots/pull/4227)).** Rollouts recorded through strands: [cagataydev/strands-isaaclab-g1-rough](https://huggingface.co/datasets/cagataydev/strands-isaaclab-g1-rough). ![playback: 4 parallel envs](playback.gif) *envs 0–3 (2×2) from the strands camera — [mp4](playback.mp4).* ## Results | | | |---|---| | parallel envs × iterations | **4096 × 1500** (147 M env steps), PhysX (`physics=isaacsim_physx`) | | wall time | **63 min 38 s** on 1× NVIDIA L40S | | env-steps/s | median **46 k**, max 77 k (GPU shared with other runs most of the time) | | mean reward | -0.90 → **19.5** best → 18.5 last | | velocity-tracking success | **0.99** last iteration; mean episode length 966 of 1000; terrain level 5.8 | | recorded rollout | 12 / 12 envs walked the full 10 s window; mean return 22.9 | | PPO iteration | mean reward | velocity-tracking success | mean ep. length (of 1000) | terrain curriculum level | env-steps/s | |---|---|---|---|---|---| | 0 | -0.90 | 0.050 | 14 | 3.51 | 18,052 | | 150 | -8.77 | 0.000 | 652 | 0.33 | 28,265 | | 375 | 4.24 | 0.000 | 941 | 3.15 | 28,978 | | 750 | 5.48 | 0.014 | 943 | 6.12 | 71,006 | | 1125 | 11.92 | 0.928 | 946 | 5.89 | 47,384 | | 1499 | 18.52 | 0.990 | 966 | 5.78 | 47,585 | Full per-iteration metrics: [`train_curve.json`](train_curve.json). ## Files | file | what | |---|---| | `model_1499.pt` | final rsl_rl checkpoint (`OnPolicyRunner.load`) — actor + critic + optimizer | | `exported/policy.pt` | TorchScript actor **incl. observation normalizer** (deterministic mean action); input `(N, 310)` → `(N, 37)` | | `exported/policy.onnx` (+ `.onnx.data`) | the same actor as ONNX | | `params/env.yaml`, `params/agent.yaml` | exact Isaac Lab env + rsl_rl agent configs of the run (seed 1) | | `train_curve.json` · `record.json` · `verify.json` | training curve · recording metadata (obs/action names) · dataset verification | | `examples/record_trained_policy.py` | the strands recording script used for the dataset | | `playback.mp4/.gif`, `frame.png` | media | ## How it was made with strands-robots **Setup** ```bash # Isaac Lab in its OWN venv (its pins clash with strands; strands never imports it) uv venv --python 3.12 ~/il && uv pip install --python ~/il/bin/python --prerelease=allow \ --index https://pypi.nvidia.com --index-strategy unsafe-best-match "isaaclab[rsl-rl,isaacsim]==3.0.0rc1" export ISAACLAB_PYTHON=~/il/bin/python export OMNI_KIT_ACCEPT_EULA=YES # you accept the NVIDIA Omniverse / Isaac Sim EULA yourself pip install "git+https://github.com/cagataycali/robots@feat/isaaclab-trainer" # strands-robots with PR #4227 ``` **Train** **As an agent tool call** (the `train_policy` tool is a Strands `@tool`): ```python from strands import Agent from strands_robots.tools.train_policy import train_policy agent = Agent(tools=[train_policy]) agent("Train the Unitree G1 to walk on rough terrain with the isaaclab provider: task Isaac-Velocity-Rough-G1, 4096 envs, PhysX, 1500 iterations, seed 1.") # -> train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c1_g1_rough_physx", # extra={"task": "Isaac-Velocity-Rough-G1", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800}) ``` **As plain Python** (exactly what produced this run): ```python from strands_robots.tools.train_policy import train_policy job = train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c1_g1_rough_physx", extra={"task": "Isaac-Velocity-Rough-G1", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800}) # poll: iteration, rewards, learning verdict, steps_per_s, checkpoint_dir train_policy(action="status", provider="isaaclab", job_id="") ``` Under the hood the provider runs `python -m isaaclab train --rl_library rsl_rl --task Isaac-Velocity-Rough-G1 --max_iterations 1500 --num_envs 4096 --seed 1 physics=isaacsim_physx` in `$ISAACLAB_PYTHON` and parses its log. Job id of this run: `isaaclab-20260929-040828-b56afb4da6e9`. Docs: [`docs/learn/training/isaaclab.md`](https://github.com/cagataycali/robots/blob/feat/isaaclab-trainer/docs/learn/training/isaaclab.md) · PR: [strands-labs/robots#4227](https://github.com/strands-labs/robots/pull/4227). **Record** (dataset repo) The final checkpoint was rolled out and recorded with [`examples/record_trained_policy.py`](examples/record_trained_policy.py) (included in this repo). It runs in the Isaac Lab venv with strands on `PYTHONPATH` and: 1. rebuilds the task env in play mode and adds an RTX camera per env; 2. loads `model_1499.pt` with rsl_rl's `OnPolicyRunner` and exports TorchScript/ONNX with Isaac Lab's exporter; 3. wraps the exported actor as a **strands `Policy`** (`RslRlJitPolicy`, max |Δa| vs rsl_rl inference = 4.8e-07); 4. steps the env with `policy.get_actions_sync(...)` and writes every frame through **strands `DatasetRecorder`** (`strands_robots.dataset_recorder.DatasetRecorder.create(...)` → `add_frame` → `save_episode`, LeRobot v3); 5. verifies the result with strands `verify_dataset` + `LeRobotDataset` load + video decode + NaN scan. ```bash OMNI_KIT_ACCEPT_EULA=YES PYTHONPATH=/path/to/strands-robots $ISAACLAB_PYTHON examples/record_trained_policy.py \ --task Isaac-Velocity-Rough-G1 --checkpoint model_1499.pt --episodes 12 --frames 500 \ --cam chase --override physics=isaacsim_physx \ --task_str "walk over rough terrain following the commanded base velocity" --robot_type unitree_g1 \ --root out/ds --repo_id cagataydev/strands-isaaclab-g1-rough $ISAACLAB_PYTHON examples/record_trained_policy.py --verify out/ds --repo_id cagataydev/strands-isaaclab-g1-rough ``` ## Use it **Play it in Isaac Lab** (in the Isaac Lab venv): ```bash huggingface-cli download cagataydev/strands-isaaclab-g1-rough-policy --local-dir g1_policy OMNI_KIT_ACCEPT_EULA=YES $ISAACLAB_PYTHON -m isaaclab play --rl_library rsl_rl --task Isaac-Velocity-Rough-G1 \ --num_envs 16 --checkpoint g1_policy/model_1499.pt physics=isaacsim_physx # PhysX! (IL-X-011); add --video --video_length 500 ``` **Raw TorchScript actor**: ```python import torch pi = torch.jit.load("g1_policy/exported/policy.pt").eval() actions = pi(obs) # obs: (N, 310) concatenated Isaac Lab 'policy' observation group -> (N, 37) ``` **As a strands `Policy`.** `create_policy("rl")` **cannot load rsl_rl checkpoints yet** (IL-X-006: strands' RL actor is a Tanh MLP, rsl_rl's is ELU with a baked-in normalizer), so wrap the exported actor — this is the adapter the recording script uses: ```python import numpy as np, torch from strands_robots.policies.base import Policy class RslRlJitPolicy(Policy): def __init__(self, jit_path, action_names, device="cpu"): self.net = torch.jit.load(jit_path, map_location=device).eval() self.device, self.action_names, self.robot_state_keys = device, list(action_names), [] @property def provider_name(self): return "isaaclab_rsl_rl_jit" def set_robot_state_keys(self, keys): self.robot_state_keys = list(keys) def reset(self, seed=None): self.net.reset() async def get_actions(self, observation_dict, instruction, **kw): x = torch.as_tensor(np.asarray(observation_dict["policy_obs"], np.float32), device=self.device) with torch.inference_mode(): y = self.net(x.unsqueeze(0))[0].cpu().numpy() return [dict(zip(self.action_names, map(float, y)))] import json names = json.load(open("g1_policy/record.json"))["action_names"] policy = RslRlJitPolicy("g1_policy/exported/policy.pt", names) chunk = policy.get_actions_sync({"policy_obs": obs_310}, "walk forward") # [{"joint_pos.left_hip_pitch_joint": ..., ...}] ``` The observation has to come from the Isaac Lab task (base velocities, gravity, velocity command, joint states, last action, height scan …), so the policy runs inside Isaac Lab; see [`examples/record_trained_policy.py`](examples/record_trained_policy.py) for the full env + camera + `DatasetRecorder` loop. ## Provenance - **strands-robots**: `feat/isaaclab-trainer` @ [`fa66fc68`](https://github.com/cagataycali/robots/commit/fa66fc682047f3cb70e2e2bc60f0b2dde2fbebae) — [strands-labs/robots#4227](https://github.com/strands-labs/robots/pull/4227) (`isaaclab` train_policy provider, `IsaacLabTrainer`; `DatasetRecorder`; `verify_dataset`) - **Isaac Lab** 3.0.0rc1 · **Isaac Sim** 6.1.0.0 · PhysX (`physics=isaacsim_physx`) · rsl-rl-lib 5.4.1 (PPO) · lerobot 0.6.1 · torch on CUDA - **GPU**: 1× NVIDIA L40S (46 GB), shared with the Shadow run / agent sweep for most of training - **Seeds**: training seed 1 (`params/agent.yaml`, `params/env.yaml`); recording seed 7 - **Training job**: `isaaclab-20260929-040828-b56afb4da6e9`, 2026-09-29 ## Limitations - **Simulation only.** Nothing here was run on a real Unitree G1; no sim-to-real claims (the policy was not trained with sim-to-real hardening beyond Isaac Lab's default randomization). - **Release candidates**: Isaac Lab 3.0.0rc1 on Isaac Sim 6.1.0.0; APIs and physics may change. PhysX and Newton results differ. - **Physics preset matters (IL-X-011):** trained on PhysX; replayed on Newton (`physics=newton_mjwarp`) it falls in < 1.1 s. Always pass `physics=isaacsim_physx` when playing it. An earlier Newton training run of the same task crashed at iteration 611 on NaN observations (IL-X-007), which is why this run uses PhysX. - Recording: 12 parallel envs from a common reset, 500 frames (10 s) each with a chase camera; all 12 walked the full window. `observation.state` mixes joint positions with the policy observation (incl. height scan), so it is wide (354-D) and not a standard LeRobot "robot state". - `create_policy("rl")` in strands **cannot load this rsl_rl checkpoint yet** (finding IL-X-006: strands' RL actor is Tanh, rsl_rl is ELU + obs-normalizer); use the exported TorchScript + the small wrapper shown above. - Known provider findings tracked with the PR: IL-X-001 (NaN reward not surfaced), IL-X-004/005 (no stop / play action), IL-X-007 (status text hid a crash traceback). ## License **Card choice: `license: other` — our generated data / weights under CC-BY-4.0, plus NVIDIA notices.** Why: - The recorded trajectories, rendered camera video, playback clips and the trained policy weights are *user-generated content* produced with NVIDIA Isaac Sim / Isaac Lab. The NVIDIA Omniverse License Agreement (which governs Isaac Sim 6.1, shipped as `isaacsim/LICENSE.txt`) §2.1 explicitly allows you to *"distribute user generated content that you develop using Omniverse, such as video, audio, stills, models, 3D assets and screen captures"*. We release that content under **CC-BY-4.0**. - **No NVIDIA Content is redistributed**: the Unitree G1 USD (`IsaacLab/Robots/Unitree/G1/g1_minimal.usd`) and scene assets come from the Isaac Lab / Isaac Sim asset packs on NVIDIA's asset server and are **not** in this repo; `params/env.yaml` only references their paths. To reproduce you download them under your own NVIDIA EULA acceptance. The rough terrain is procedurally generated by Isaac Lab. "Unitree G1" is a product of Unitree Robotics; no endorsement by Unitree or NVIDIA is implied. - `params/*.yaml` are Isaac Lab task / agent configurations (Isaac Lab is **BSD-3-Clause**); the example script is Apache-2.0 like strands-robots. Running Isaac Sim itself requires accepting the NVIDIA Isaac Sim / Omniverse EULA.