Instructions to use cagataydev/strands-isaaclab-g1-rough-policy with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use cagataydev/strands-isaaclab-g1-rough-policy with LeRobot:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
strands-isaaclab-g1-rough-policy
rsl_rl PPO policy for Unitree G1 humanoid rough-terrain locomotion (Isaac-Velocity-Rough-G1), trained with strands-robots' isaaclab
train_policy provider (PR #4227). Rollouts recorded through strands: cagataydev/strands-isaaclab-g1-rough.
envs 0–3 (2×2) from the strands camera — mp4.
Results
| parallel envs × iterations | 4096 × 1500 (147 M env steps), PhysX (physics=isaacsim_physx) |
| wall time | 63 min 38 s on 1× NVIDIA L40S |
| env-steps/s | median 46 k, max 77 k (GPU shared with other runs most of the time) |
| mean reward | -0.90 → 19.5 best → 18.5 last |
| velocity-tracking success | 0.99 last iteration; mean episode length 966 of 1000; terrain level 5.8 |
| recorded rollout | 12 / 12 envs walked the full 10 s window; mean return 22.9 |
| PPO iteration | mean reward | velocity-tracking success | mean ep. length (of 1000) | terrain curriculum level | env-steps/s |
|---|---|---|---|---|---|
| 0 | -0.90 | 0.050 | 14 | 3.51 | 18,052 |
| 150 | -8.77 | 0.000 | 652 | 0.33 | 28,265 |
| 375 | 4.24 | 0.000 | 941 | 3.15 | 28,978 |
| 750 | 5.48 | 0.014 | 943 | 6.12 | 71,006 |
| 1125 | 11.92 | 0.928 | 946 | 5.89 | 47,384 |
| 1499 | 18.52 | 0.990 | 966 | 5.78 | 47,585 |
Full per-iteration metrics: train_curve.json.
Files
| file | what |
|---|---|
model_1499.pt |
final rsl_rl checkpoint (OnPolicyRunner.load) — actor + critic + optimizer |
exported/policy.pt |
TorchScript actor incl. observation normalizer (deterministic mean action); input (N, 310) → (N, 37) |
exported/policy.onnx (+ .onnx.data) |
the same actor as ONNX |
params/env.yaml, params/agent.yaml |
exact Isaac Lab env + rsl_rl agent configs of the run (seed 1) |
train_curve.json · record.json · verify.json |
training curve · recording metadata (obs/action names) · dataset verification |
examples/record_trained_policy.py |
the strands recording script used for the dataset |
playback.mp4/.gif, frame.png |
media |
How it was made with strands-robots
Setup
# Isaac Lab in its OWN venv (its pins clash with strands; strands never imports it)
uv venv --python 3.12 ~/il && uv pip install --python ~/il/bin/python --prerelease=allow \
--index https://pypi.nvidia.com --index-strategy unsafe-best-match "isaaclab[rsl-rl,isaacsim]==3.0.0rc1"
export ISAACLAB_PYTHON=~/il/bin/python
export OMNI_KIT_ACCEPT_EULA=YES # you accept the NVIDIA Omniverse / Isaac Sim EULA yourself
pip install "git+https://github.com/cagataycali/robots@feat/isaaclab-trainer" # strands-robots with PR #4227
Train
As an agent tool call (the train_policy tool is a Strands @tool):
from strands import Agent
from strands_robots.tools.train_policy import train_policy
agent = Agent(tools=[train_policy])
agent("Train the Unitree G1 to walk on rough terrain with the isaaclab provider: task Isaac-Velocity-Rough-G1, 4096 envs, PhysX, 1500 iterations, seed 1.")
# -> train_policy(action="train", provider="isaaclab", steps=1500, seed=1, output_dir="runs/c1_g1_rough_physx",
# extra={"task": "Isaac-Velocity-Rough-G1", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800})
As plain Python (exactly what produced this run):
from strands_robots.tools.train_policy import train_policy
job = train_policy(action="train", provider="isaaclab", steps=1500, seed=1,
output_dir="runs/c1_g1_rough_physx",
extra={"task": "Isaac-Velocity-Rough-G1", "num_envs": 4096, "physics": "isaacsim_physx", "timeout_s": 10800})
# poll: iteration, rewards, learning verdict, steps_per_s, checkpoint_dir
train_policy(action="status", provider="isaaclab", job_id="<job_id from the result>")
Under the hood the provider runs python -m isaaclab train --rl_library rsl_rl --task Isaac-Velocity-Rough-G1 --max_iterations 1500 --num_envs 4096 --seed 1 physics=isaacsim_physx
in $ISAACLAB_PYTHON and parses its log. Job id of this run: isaaclab-20260929-040828-b56afb4da6e9.
Docs: docs/learn/training/isaaclab.md · PR: strands-labs/robots#4227.
Record (dataset repo)
The final checkpoint was rolled out and recorded with examples/record_trained_policy.py
(included in this repo). It runs in the Isaac Lab venv with strands on PYTHONPATH and:
- rebuilds the task env in play mode and adds an RTX camera per env;
- loads
model_1499.ptwith rsl_rl'sOnPolicyRunnerand exports TorchScript/ONNX with Isaac Lab's exporter; - wraps the exported actor as a strands
Policy(RslRlJitPolicy, max |Δa| vs rsl_rl inference = 4.8e-07); - steps the env with
policy.get_actions_sync(...)and writes every frame through strandsDatasetRecorder(strands_robots.dataset_recorder.DatasetRecorder.create(...)→add_frame→save_episode, LeRobot v3); - verifies the result with strands
verify_dataset+LeRobotDatasetload + video decode + NaN scan.
OMNI_KIT_ACCEPT_EULA=YES PYTHONPATH=/path/to/strands-robots $ISAACLAB_PYTHON examples/record_trained_policy.py \
--task Isaac-Velocity-Rough-G1 --checkpoint model_1499.pt --episodes 12 --frames 500 \
--cam chase --override physics=isaacsim_physx \
--task_str "walk over rough terrain following the commanded base velocity" --robot_type unitree_g1 \
--root out/ds --repo_id cagataydev/strands-isaaclab-g1-rough
$ISAACLAB_PYTHON examples/record_trained_policy.py --verify out/ds --repo_id cagataydev/strands-isaaclab-g1-rough
Use it
Play it in Isaac Lab (in the Isaac Lab venv):
huggingface-cli download cagataydev/strands-isaaclab-g1-rough-policy --local-dir g1_policy
OMNI_KIT_ACCEPT_EULA=YES $ISAACLAB_PYTHON -m isaaclab play --rl_library rsl_rl --task Isaac-Velocity-Rough-G1 \
--num_envs 16 --checkpoint g1_policy/model_1499.pt physics=isaacsim_physx # PhysX! (IL-X-011); add --video --video_length 500
Raw TorchScript actor:
import torch
pi = torch.jit.load("g1_policy/exported/policy.pt").eval()
actions = pi(obs) # obs: (N, 310) concatenated Isaac Lab 'policy' observation group -> (N, 37)
As a strands Policy. create_policy("rl") cannot load rsl_rl checkpoints yet (IL-X-006: strands' RL actor is a Tanh MLP,
rsl_rl's is ELU with a baked-in normalizer), so wrap the exported actor — this is the adapter the recording script uses:
import numpy as np, torch
from strands_robots.policies.base import Policy
class RslRlJitPolicy(Policy):
def __init__(self, jit_path, action_names, device="cpu"):
self.net = torch.jit.load(jit_path, map_location=device).eval()
self.device, self.action_names, self.robot_state_keys = device, list(action_names), []
@property
def provider_name(self): return "isaaclab_rsl_rl_jit"
def set_robot_state_keys(self, keys): self.robot_state_keys = list(keys)
def reset(self, seed=None): self.net.reset()
async def get_actions(self, observation_dict, instruction, **kw):
x = torch.as_tensor(np.asarray(observation_dict["policy_obs"], np.float32), device=self.device)
with torch.inference_mode():
y = self.net(x.unsqueeze(0))[0].cpu().numpy()
return [dict(zip(self.action_names, map(float, y)))]
import json
names = json.load(open("g1_policy/record.json"))["action_names"]
policy = RslRlJitPolicy("g1_policy/exported/policy.pt", names)
chunk = policy.get_actions_sync({"policy_obs": obs_310}, "walk forward") # [{"joint_pos.left_hip_pitch_joint": ..., ...}]
The observation has to come from the Isaac Lab task (base velocities, gravity, velocity command, joint states, last action, height scan …), so the
policy runs inside Isaac Lab; see examples/record_trained_policy.py for the full
env + camera + DatasetRecorder loop.
Provenance
- strands-robots:
feat/isaaclab-trainer@fa66fc68— strands-labs/robots#4227 (isaaclabtrain_policy provider,IsaacLabTrainer;DatasetRecorder;verify_dataset) - Isaac Lab 3.0.0rc1 · Isaac Sim 6.1.0.0 · PhysX (
physics=isaacsim_physx) · rsl-rl-lib 5.4.1 (PPO) · lerobot 0.6.1 · torch on CUDA - GPU: 1× NVIDIA L40S (46 GB), shared with the Shadow run / agent sweep for most of training
- Seeds: training seed 1 (
params/agent.yaml,params/env.yaml); recording seed 7 - Training job:
isaaclab-20260929-040828-b56afb4da6e9, 2026-09-29
Limitations
- Simulation only. Nothing here was run on a real Unitree G1; no sim-to-real claims (the policy was not trained with sim-to-real hardening beyond Isaac Lab's default randomization).
- Release candidates: Isaac Lab 3.0.0rc1 on Isaac Sim 6.1.0.0; APIs and physics may change. PhysX and Newton results differ.
- Physics preset matters (IL-X-011): trained on PhysX; replayed on Newton (
physics=newton_mjwarp) it falls in < 1.1 s. Always passphysics=isaacsim_physxwhen playing it. An earlier Newton training run of the same task crashed at iteration 611 on NaN observations (IL-X-007), which is why this run uses PhysX. - Recording: 12 parallel envs from a common reset, 500 frames (10 s) each with a chase camera; all 12 walked the full window.
observation.statemixes joint positions with the policy observation (incl. height scan), so it is wide (354-D) and not a standard LeRobot "robot state". create_policy("rl")in strands cannot load this rsl_rl checkpoint yet (finding IL-X-006: strands' RL actor is Tanh, rsl_rl is ELU- obs-normalizer); use the exported TorchScript + the small wrapper shown above.
- Known provider findings tracked with the PR: IL-X-001 (NaN reward not surfaced), IL-X-004/005 (no stop / play action), IL-X-007 (status text hid a crash traceback).
License
Card choice: license: other — our generated data / weights under CC-BY-4.0, plus NVIDIA notices. Why:
- The recorded trajectories, rendered camera video, playback clips and the trained policy weights are user-generated content
produced with NVIDIA Isaac Sim / Isaac Lab. The NVIDIA Omniverse License Agreement (which governs Isaac Sim 6.1, shipped as
isaacsim/LICENSE.txt) §2.1 explicitly allows you to "distribute user generated content that you develop using Omniverse, such as video, audio, stills, models, 3D assets and screen captures". We release that content under CC-BY-4.0. - No NVIDIA Content is redistributed: the Unitree G1 USD (
IsaacLab/Robots/Unitree/G1/g1_minimal.usd) and scene assets come from the Isaac Lab / Isaac Sim asset packs on NVIDIA's asset server and are not in this repo;params/env.yamlonly references their paths. To reproduce you download them under your own NVIDIA EULA acceptance. The rough terrain is procedurally generated by Isaac Lab. "Unitree G1" is a product of Unitree Robotics; no endorsement by Unitree or NVIDIA is implied. params/*.yamlare Isaac Lab task / agent configurations (Isaac Lab is BSD-3-Clause); the example script is Apache-2.0 like strands-robots. Running Isaac Sim itself requires accepting the NVIDIA Isaac Sim / Omniverse EULA.
Dataset used to train cagataydev/strands-isaaclab-g1-rough-policy
Collection including cagataydev/strands-isaaclab-g1-rough-policy
Evaluation results
- velocity-tracking success rate (last training iteration) on Isaac Lab Isaac-Velocity-Rough-G1 (4096 envs, PhysX)self-reported0.990
- mean episode reward (last iteration) on Isaac Lab Isaac-Velocity-Rough-G1 (4096 envs, PhysX)self-reported18.520
