Robotics
ONNX
LeRobot
rsl_rl
isaaclab
isaac-sim
strands-robots
reinforcement-learning
franka
reach
ppo
agents
sim2sim
Eval Results (legacy)
How to use from the
Use from the
LeRobot library
# No code snippets available yet for this library.

# To use this model, check the repository files and the library's documentation.

# Want to help? PRs adding snippets are welcome at:
# https://github.com/huggingface/huggingface.js

strands-isaaclab-reach-franka-policy

rsl_rl PPO policy for Franka Panda pose reaching (Isaac-Reach-Franka), trained by a Strands Agent through strands-robots' isaaclab train_policy provider (PR #4227) — and deployed unchanged on strands' MuJoCo backend (sim-to-sim).

playback: 4 recorded envs

Isaac Lab rollouts from the strands camera — mp4. Datasets: Isaac Lab rollouts cagataydev/strands-isaaclab-reach-franka · strands MuJoCo rollouts cagataydev/strands-isaaclab-reach-franka-mujoco-sim2sim.

Results

parallel envs × iterations 4096 × 300 (29.5 M env steps), Newton/MJWarp
wall time 4 min 15 s on 1× NVIDIA L40S
env-steps/s median 117 k, max 126 k
success rate (Isaac Lab) 0.983 last iteration; recorded rollout 11 / 12 reached
sim-to-sim (strands MuJoCo) 10 / 12 reached (< 5 cm, < 0.2 rad); mean final position error 10.6 cm (keeps acting after success)
selection best of a 6-run agent sweep (below)
PPO iteration mean reward success rate mean ep. length (of 360) env-steps/s
0 -0.25 0.000 21 36,907
30 -2.87 0.017 360 114,322
75 -1.41 0.309 312 114,751
150 -0.40 0.827 147 122,203
225 -0.37 0.939 104 117,760
299 0.02 0.983 47 119,700

Agent sweep — examples/agent_rl_research_lead.py; the agent's full answer is agent_answer.md and the transcript agent_transcript.json:

lr seed job final mean reward (what the agent judged) success rate (Isaac Lab metric, last it) ee position error
3e-4 1 …edcec0d28a0b -0.35 0.914 8.0 cm
3e-4 2 …b06eca8dd857 -0.29 0.965 6.5 cm
1e-3 1 …05a8c9c0dbae (this policy) +0.02 0.983 6.0 cm
1e-3 2 …116af3a5a1bb -0.51 0.877 9.0 cm
3e-3 1 …41a69ccff38b -0.34 0.922 8.4 cm
3e-3 2 …e67c06646018 -0.47 0.939 7.5 cm

Files

file what
model_299.pt final rsl_rl checkpoint (OnPolicyRunner.load)
exported/policy.pt · exported/policy.onnx actor incl. observation normalizer; input (N, 32) → (N, 7)
params/env.yaml, params/agent.yaml exact Isaac Lab env + rsl_rl agent configs (seed 1, lr 1e-3)
train_curve.json · record.json · verify.json · sim2sim_mujoco.json curve · recording metadata · dataset checks · sim-to-sim result
agent_answer.md · agent_transcript.json the Strands Agent's final report + full tool-call transcript
examples/ agent_rl_research_lead.py (sweep), record_trained_policy.py (Isaac Lab → strands dataset), sim2sim_mujoco.py (strands MuJoCo)

How it was made with strands-robots

Setup

# Isaac Lab in its OWN venv (its pins clash with strands; strands never imports it)
uv venv --python 3.12 ~/il && uv pip install --python ~/il/bin/python --prerelease=allow \
  --index https://pypi.nvidia.com --index-strategy unsafe-best-match "isaaclab[rsl-rl,isaacsim]==3.0.0rc1"
export ISAACLAB_PYTHON=~/il/bin/python
export OMNI_KIT_ACCEPT_EULA=YES          # you accept the NVIDIA Omniverse / Isaac Sim EULA yourself
pip install "git+https://github.com/cagataycali/robots@feat/isaaclab-trainer" strands-agents   # strands-robots with PR #4227

Train (agent-driven)

This policy was trained by a Strands Agent: one natural-language request → Agent(tools=[train_policy, gpu_status, wait_minutes, stop_job]) (Bedrock Claude) ran a 3 learning-rate × 2 seed sweep of 300-iteration PPO runs through the isaaclab provider, one at a time on the shared GPU, polled them and picked a winner. Script: examples/agent_rl_research_lead.py.

from strands import Agent
from strands_robots.tools.train_policy import train_policy

agent = Agent(tools=[train_policy])          # the example adds gpu_status / wait_minutes / stop_job helpers
agent("Sweep learning rate 3e-4, 1e-3, 3e-3 x seeds 1, 2 on Isaac-Reach-Franka, 4096 envs, 300 iterations each, "
      "one run at a time; poll with status and pick the best.")

The exact tool call the agent made for this run (from agent_transcript.json) — equally usable as plain Python:

train_policy(action="train", provider="isaaclab", steps=300, seed=1, learning_rate=1e-3,
             output_dir="runs/c3_agent_sweep/1e-3_s1",
             extra={"task": "Isaac-Reach-Franka", "num_envs": 4096, "timeout_s": 1800})
train_policy(action="status", provider="isaaclab", job_id="isaaclab-20260929-050417-05a8c9c0dbae")

→ python -m isaaclab train --rl_library rsl_rl --task Isaac-Reach-Franka --max_iterations 300 --num_envs 4096 --seed 1 agent.algorithm.learning_rate=0.001 in $ISAACLAB_PYTHON (task-default physics: Newton / MJWarp). Docs: docs/learn/training/isaaclab.md · PR: strands-labs/robots#4227.

Record (dataset repo)

Recorded with examples/record_trained_policy.py: rebuild the task env (play mode) with an RTX camera → load model_299.pt with rsl_rl and export TorchScript/ONNX → wrap the actor as a strands Policy (RslRlJitPolicy, max |Δa| vs rsl_rl = 3.0e-07) → step with policy.get_actions_sync(...) → write every frame via strands DatasetRecorder (LeRobot v3) → verify with strands verify_dataset.

OMNI_KIT_ACCEPT_EULA=YES PYTHONPATH=/path/to/strands-robots $ISAACLAB_PYTHON examples/record_trained_policy.py \
  --task Isaac-Reach-Franka --checkpoint model_299.pt --episodes 12 --frames 360 \
  --cam fixed --cam_name front --eye 1.6,0.9,0.9 --target 0.4,0,0.3 \
  --task_str "reach the commanded end-effector pose" --robot_type franka_panda --root out/ds --repo_id cagataydev/strands-isaaclab-reach-franka

Use it

Play it in Isaac Lab:

huggingface-cli download cagataydev/strands-isaaclab-reach-franka-policy --local-dir reach_policy
OMNI_KIT_ACCEPT_EULA=YES $ISAACLAB_PYTHON -m isaaclab play --rl_library rsl_rl --task Isaac-Reach-Franka \
  --num_envs 16 --checkpoint reach_policy/model_299.pt        # trained on the task-default Newton/MJWarp physics

As a strands Policy — create_policy("rl") cannot load rsl_rl checkpoints yet (IL-X-006), so wrap the exported actor:

import json, numpy as np, torch
from strands_robots.policies.base import Policy

class RslRlJitPolicy(Policy):
    def __init__(self, jit_path, action_names, device="cpu"):
        self.net = torch.jit.load(jit_path, map_location=device).eval()
        self.device, self.action_names, self.robot_state_keys = device, list(action_names), []
    @property
    def provider_name(self): return "isaaclab_rsl_rl_jit"
    def set_robot_state_keys(self, keys): self.robot_state_keys = list(keys)
    def reset(self, seed=None): self.net.reset()
    async def get_actions(self, observation_dict, instruction, **kw):
        x = torch.as_tensor(np.asarray(observation_dict["policy_obs"], np.float32), device=self.device)
        with torch.inference_mode():
            y = self.net(x.unsqueeze(0))[0].cpu().numpy()
        return [dict(zip(self.action_names, map(float, y)))]

names = json.load(open("reach_policy/record.json"))["action_names"]
policy = RslRlJitPolicy("reach_policy/exported/policy.pt", names)
chunk = policy.get_actions_sync({"policy_obs": obs_32}, "reach the pose")   # [{"arm_action.panda_joint1": ..., ...}]

On strands' MuJoCo backend (no Isaac needed) — examples/sim2sim_mujoco.py rebuilds the 32-D Isaac Lab observation (joint pos − default | joint vel | pose command | last action) from the strands observation and drives the menagerie panda through the standard strands simulation API:

from strands_robots.simulation.factory import create_simulation

sim = create_simulation("mujoco")
sim.create_world(timestep=1.0 / 120.0)
sim.add_robot("arm", data_config="panda")
sim.add_camera("front", position=[1.6, 0.9, 0.9], target=[0.4, 0.0, 0.3], width=320, height=240)
sim.start_recording(repo_id="me/reach-sim2sim", root="ds", fps=30, task="reach the commanded end-effector pose", cameras=["front"])
sim.run_policy("arm", policy_object=IsaacReachOnStrands(jit, sim, seed=0),   # the adapter class in the example
               instruction="reach the commanded end-effector pose", control_frequency=30.0, control_substeps=4,
               n_steps=120, n_episodes=12, reset_between=True, seed=0)
sim.stop_recording()
MUJOCO_GL=egl python examples/sim2sim_mujoco.py --jit reach_policy/exported/policy.pt --episodes 12 --root ds --repo_id me/reach-sim2sim

Provenance

  • strands-robots: feat/isaaclab-trainer @ fa66fc68 — strands-labs/robots#4227 (isaaclab train_policy provider; DatasetRecorder; verify_dataset)
  • Isaac Lab 3.0.0rc1 · Isaac Sim 6.1.0.0 · Newton / MJWarp (task default physics) · rsl-rl-lib 5.4.1 (PPO) · lerobot 0.6.1
  • GPU: 1× NVIDIA L40S (46 GB), shared with other trainings during the sweep
  • Seeds: training seed 1, lr 1e-3 (params/agent.yaml); recording seed 7
  • Training job: isaaclab-20260929-050417-05a8c9c0dbae (best of 6 agent-launched runs), 2026-09-29; agent: Strands Agent on Amazon Bedrock (Claude)

Limitations

  • Simulation only; no real Franka was used.
  • Release candidates: Isaac Lab 3.0.0rc1 / Isaac Sim 6.1.0.0. Trained on Newton/MJWarp; PhysX results may differ (IL-X-011: the physics preset is not remembered by the checkpoint — pass the same one when replaying).
  • Short run (300 iterations, ~4 min). Mean reward ends near 0 and falls mid-training because the task's curriculum ramps the action-rate/joint-velocity penalties and successful episodes terminate early; judge it by success rate. The agent itself misread this and concluded "none of the 6 runs clearly learned" (finding IL-X-012) — its full answer is in agent_answer.md.
  • Recording: episodes end at success, so lengths vary (9 – 360 frames); 11/12 reached the target, 1 timed out at 12 s.
  • create_policy("rl") cannot load rsl_rl checkpoints yet (IL-X-006); use the exported TorchScript + the wrapper below.
  • Sim-to-sim is to MuJoCo Menagerie's panda (not the Isaac Lab USD): 2 / 12 episodes missed (30 – 34 cm); the adapter hard-codes Isaac Lab's action scale 0.5 and default joint pose.

License

Card choice: license: other — our generated data / weights under CC-BY-4.0, plus NVIDIA notices. Why:

  • Trajectories, rendered camera video, playback clips, trained weights and agent logs are user-generated content made with NVIDIA Isaac Sim / Isaac Lab; the NVIDIA Omniverse License Agreement (governs Isaac Sim 6.1, isaacsim/LICENSE.txt) §2.1 allows distributing "user generated content that you develop using Omniverse, such as video, audio, stills, models, 3D assets and screen captures". We release it under CC-BY-4.0.
  • No NVIDIA Content is redistributed: the Franka Panda USD and scene assets come from the Isaac Lab / Isaac Sim asset packs and are not in this repo (params/env.yaml only references their paths). "Franka" is a trademark of Franka Robotics; no endorsement by Franka Robotics or NVIDIA is implied.
  • params/*.yaml are Isaac Lab configurations (BSD-3-Clause); examples/*.py are Apache-2.0 like strands-robots. Running Isaac Sim requires your own acceptance of the NVIDIA Isaac Sim / Omniverse EULA. Full text: LICENSE.md.
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Datasets used to train cagataydev/strands-isaaclab-reach-franka-policy

Evaluation results

  • success rate (last training iteration) on Isaac Lab Isaac-Reach-Franka (4096 envs, Newton/MJWarp)
    self-reported
    0.983
  • reached pose in 10/12 episodes on strands-robots MuJoCo backend, menagerie panda
    self-reported
    0.833