strands-isaaclab-shadow-reorient-policy

rsl_rl PPO policy for Shadow Hand in-hand cube reorientation (Isaac-Reorient-Cube-Shadow), trained with strands-robots' isaaclab train_policy provider (PR #4227). Rollouts recorded through strands: cagataydev/strands-isaaclab-shadow-reorient.

playback: 4 parallel envs

envs 0–3 (2×2) from the strands camera — mp4 · one full recorded episode: episode1.mp4.

Results

parallel envs × iterations 8192 × 3000 (393 M env steps), PhysX
wall time 59 min 50 s on 1× NVIDIA L40S
env-steps/s median 110 k, max 166 k (GPU shared with another run for part of training)
mean reward 0.02 → 102.6 best → 91.9 last
success rate 0.93 last iteration (best 0.96); mean episode length 535 of 600
recorded rollout 11 / 12 envs held the cube for the full 5 s window; mean return 56.1
PPO iteration mean reward success rate mean ep. length (of 600) orientation error (rad) env-steps/s
0 0.02 0.000 13 2.041 71,679
300 8.84 0.297 470 1.405 95,302
750 56.23 0.895 548 1.334 98,068
1500 83.34 0.915 568 1.354 112,092
2250 92.01 0.939 563 1.407 112,528
2999 91.93 0.934 535 1.393 111,316

Full per-iteration metrics: train_curve.json.

Files

file what
model_2999.pt final rsl_rl checkpoint (OnPolicyRunner.load) — actor + critic + optimizer
exported/policy.pt TorchScript actor incl. observation normalizer (deterministic mean action); input (N, 157) → (N, 20)
exported/policy.onnx (+ .onnx.data) the same actor as ONNX
params/env.yaml, params/agent.yaml exact Isaac Lab env + rsl_rl agent configs of the run (seed 1)
train_curve.json · record.json · verify.json training curve · recording metadata (obs/action names) · dataset verification
examples/record_trained_policy.py the strands recording script used for the dataset
playback.mp4/.gif, episode1.mp4, frame.png media

How it was made with strands-robots

Setup

# Isaac Lab in its OWN venv (its pins clash with strands; strands never imports it)
uv venv --python 3.12 ~/il && uv pip install --python ~/il/bin/python --prerelease=allow \
  --index https://pypi.nvidia.com --index-strategy unsafe-best-match "isaaclab[rsl-rl,isaacsim]==3.0.0rc1"
export ISAACLAB_PYTHON=~/il/bin/python
export OMNI_KIT_ACCEPT_EULA=YES          # you accept the NVIDIA Omniverse / Isaac Sim EULA yourself
pip install "git+https://github.com/cagataycali/robots@feat/isaaclab-trainer"   # strands-robots with PR #4227

Train

As an agent tool call (the train_policy tool is a Strands @tool):

from strands import Agent
from strands_robots.tools.train_policy import train_policy

agent = Agent(tools=[train_policy])
agent("Train the Shadow hand to reorient a cube with the isaaclab provider: task Isaac-Reorient-Cube-Shadow, 3000 iterations, seed 1.")
# -> train_policy(action="train", provider="isaaclab", steps=3000, seed=1, output_dir="runs/c2_shadow_reorient",
#                 extra={"task": "Isaac-Reorient-Cube-Shadow", "num_envs": 8192, "timeout_s": 10800})

As plain Python (exactly what produced this run; num_envs 8192 is also the task default):

from strands_robots.tools.train_policy import train_policy

job = train_policy(action="train", provider="isaaclab", steps=3000, seed=1,
                   output_dir="runs/c2_shadow_reorient",
                   extra={"task": "Isaac-Reorient-Cube-Shadow", "num_envs": 8192, "timeout_s": 10800})
# poll: iteration, rewards, learning verdict, steps_per_s, checkpoint_dir
train_policy(action="status", provider="isaaclab", job_id="<job_id from the result>")

Under the hood the provider runs python -m isaaclab train --rl_library rsl_rl --task Isaac-Reorient-Cube-Shadow --max_iterations 3000 --seed 1 in $ISAACLAB_PYTHON and parses its log. Job id of this run: isaaclab-20260929-034557-204417977ae5. Docs: docs/learn/training/isaaclab.md · PR: strands-labs/robots#4227.

Record (dataset repo)

The final checkpoint was rolled out and recorded with examples/record_trained_policy.py (included in this repo). It runs in the Isaac Lab venv with strands on PYTHONPATH and:

  1. rebuilds the task env in play mode and adds an RTX camera per env;
  2. loads model_2999.pt with rsl_rl's OnPolicyRunner and exports TorchScript/ONNX with Isaac Lab's exporter;
  3. wraps the exported actor as a strands Policy (RslRlJitPolicy, max |Δa| vs rsl_rl inference = 1.2e-05);
  4. steps the env with policy.get_actions_sync(...) and writes every frame through strands DatasetRecorder (strands_robots.dataset_recorder.DatasetRecorder.create(...) → add_frame → save_episode, LeRobot v3);
  5. verifies the result with strands verify_dataset + LeRobotDataset load + video decode + NaN scan.
OMNI_KIT_ACCEPT_EULA=YES PYTHONPATH=/path/to/strands-robots $ISAACLAB_PYTHON examples/record_trained_policy.py \
  --task Isaac-Reorient-Cube-Shadow --checkpoint model_2999.pt --episodes 12 --frames 300 \
  --cam fixed --cam_name front --eye 0.42,-0.72,0.85 --target 0,-0.39,0.55 \
  --task_str "reorient the cube in hand to match the goal orientation" --robot_type shadow_hand \
  --root out/ds --repo_id cagataydev/strands-isaaclab-shadow-reorient
$ISAACLAB_PYTHON examples/record_trained_policy.py --verify out/ds --repo_id cagataydev/strands-isaaclab-shadow-reorient

Use it

Play it in Isaac Lab (in the Isaac Lab venv):

huggingface-cli download cagataydev/strands-isaaclab-shadow-reorient-policy --local-dir shadow_policy
OMNI_KIT_ACCEPT_EULA=YES $ISAACLAB_PYTHON -m isaaclab play --rl_library rsl_rl --task Isaac-Reorient-Cube-Shadow \
  --num_envs 16 --checkpoint shadow_policy/model_2999.pt            # add --video --video_length 400 for an mp4

Raw TorchScript actor:

import torch
pi = torch.jit.load("shadow_policy/exported/policy.pt").eval()
actions = pi(obs)          # obs: (N, 157) concatenated Isaac Lab 'policy' observation group -> (N, 20)

As a strands Policy. create_policy("rl") cannot load rsl_rl checkpoints yet (IL-X-006: strands' RL actor is a Tanh MLP, rsl_rl's is ELU with a baked-in normalizer), so wrap the exported actor — this is the adapter the recording script uses:

import numpy as np, torch
from strands_robots.policies.base import Policy

class RslRlJitPolicy(Policy):
    def __init__(self, jit_path, action_names, device="cpu"):
        self.net = torch.jit.load(jit_path, map_location=device).eval()
        self.device, self.action_names, self.robot_state_keys = device, list(action_names), []
    @property
    def provider_name(self): return "isaaclab_rsl_rl_jit"
    def set_robot_state_keys(self, keys): self.robot_state_keys = list(keys)
    def reset(self, seed=None): self.net.reset()
    async def get_actions(self, observation_dict, instruction, **kw):
        x = torch.as_tensor(np.asarray(observation_dict["policy_obs"], np.float32), device=self.device)
        with torch.inference_mode():
            y = self.net(x.unsqueeze(0))[0].cpu().numpy()
        return [dict(zip(self.action_names, map(float, y)))]

import json
names = json.load(open("shadow_policy/record.json"))["action_names"]
policy = RslRlJitPolicy("shadow_policy/exported/policy.pt", names)
chunk = policy.get_actions_sync({"policy_obs": obs_157}, "reorient the cube")   # [{"joint_pos.rh_WRJ2": ..., ...}]

The observation has to come from the Isaac Lab task (object pose, goal quaternion, fingertip states, last action …), so the policy runs inside Isaac Lab; see examples/record_trained_policy.py for the full env + camera + DatasetRecorder loop.

Provenance

  • strands-robots: feat/isaaclab-trainer @ fa66fc68 — strands-labs/robots#4227 (isaaclab train_policy provider, IsaacLabTrainer; DatasetRecorder; verify_dataset)
  • Isaac Lab 3.0.0rc1 · Isaac Sim 6.1.0.0 · PhysX (task default physics) · rsl-rl-lib 5.4.1 (PPO) · lerobot 0.6.1 · torch on CUDA
  • GPU: 1× NVIDIA L40S (46 GB), shared with other runs for part of training
  • Seeds: training seed 1 (params/agent.yaml, params/env.yaml); recording seed 7
  • Training job: isaaclab-20260929-034557-204417977ae5, 2026-09-29

Limitations

  • Simulation only. Nothing here was run on a real Shadow Hand; no sim-to-real claims.
  • Release candidates: Isaac Lab 3.0.0rc1 on Isaac Sim 6.1.0.0; APIs and physics may change. PhysX and Newton results differ.
  • Recording: 12 parallel envs from a common reset, 300 frames (5 s) each — episode 0 dropped the cube at frame 55 (terminated); the other 11 held it for the full window. observation.state mixes joint positions with the policy observation, so it is wide (188-D) and not a standard LeRobot "robot state".
  • create_policy("rl") in strands cannot load this rsl_rl checkpoint yet (finding IL-X-006: strands' RL actor is Tanh, rsl_rl is ELU
    • obs-normalizer); use the exported TorchScript + the small wrapper shown above.
  • Known provider findings tracked with the PR: IL-X-001 (NaN reward not surfaced), IL-X-004/005 (no stop / play action), IL-X-007 (status text hid a crash traceback).

License

Card choice: license: other — our generated data / weights under CC-BY-4.0, plus NVIDIA notices. Why:

  • The recorded trajectories, rendered camera video, playback clips and the trained policy weights are user-generated content produced with NVIDIA Isaac Sim / Isaac Lab. The NVIDIA Omniverse License Agreement (which governs Isaac Sim 6.1, shipped as isaacsim/LICENSE.txt) §2.1 explicitly allows you to "distribute user generated content that you develop using Omniverse, such as video, audio, stills, models, 3D assets and screen captures". We release that content under CC-BY-4.0.
  • No NVIDIA Content is redistributed: the Shadow Hand USD (Robots_Multiphysics/ShadowRobot/ShadowHandMultiPhysics_v0) and scene assets come from the Isaac Lab / Isaac Sim asset packs on NVIDIA's asset server and are not in this repo; params/env.yaml only references their paths. To reproduce you download them under your own NVIDIA EULA acceptance. "Shadow Hand" is a design of The Shadow Robot Company; no endorsement by Shadow Robot or NVIDIA is implied.
  • params/*.yaml are Isaac Lab task / agent configurations (Isaac Lab is BSD-3-Clause); the example script is Apache-2.0 like strands-robots. Running Isaac Sim itself requires accepting the NVIDIA Isaac Sim / Omniverse EULA.
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Dataset used to train cagataydev/strands-isaaclab-shadow-reorient-policy

Evaluation results

  • success rate (last training iteration) on Isaac Lab Isaac-Reorient-Cube-Shadow (8192 envs, PhysX)
    self-reported
    0.934
  • mean episode reward (last iteration) on Isaac Lab Isaac-Reorient-Cube-Shadow (8192 envs, PhysX)
    self-reported
    91.930