GR00T N1.6 โ€” pointwam assemble-tissue (tissue roll onto holder)

RLWRLD OpenArm + RH56F1 bimanual platform finetune of GR00T N1.6 3B. 30,000 steps on 2x H100.

Training setup

base model nvidia/GR00T-N1.6-3B
embodiment tag new_embodiment
steps 30,000
GPUs 2x H100 80GB
global batch size 64 (32 per GPU)
learning rate 1e-4, cosine, warmup ratio 0.05
weight decay 1e-5
action chunk 30 (1.0 s at 30 fps)
action representation absolute joint targets
wall clock 4h 04m

Observation and action space

video    ego_left, ego_right      320x180 stereo ego pair
         -> letterbox padded to square, then 256x256, aspect preserved (~40% black bars)
state    28-D  openarm_state_joints[0:16], left_hand_joints[16:22], right_hand_joints[22:28]
action   28-D  head[0:2], arm[2:16], left_hand[16:22], right_hand[22:28]

State and action group their dimensions differently. That is the dataset's own layout, not a mistake: GR00T reads the two sides independently, so they do not need to line up.

What is in here

The final export only: model-0000{1,2}-of-00002.safetensors, config.json, processor/, experiment_cfg/. Enough for inference and evaluation.

Optimizer state and the intermediate checkpoints (10000 / 20000) are not included.

Evaluation notes

  • Do not reimplement the preprocessing. letterbox -> SmallestMaxSize(256) -> center crop(0.95) -> SmallestMaxSize(256) is baked into the processor in this checkpoint. Calling processor.eval() reproduces training exactly; hand-rolling it is the usual way eval silently diverges.
  • Camera order is ego_left then ego_right. Swapping them costs accuracy quietly.
  • Feed the native 320x180. Pre-resizing to 256x256 gets letterboxed a second time and produces a different image than training saw.
  • Use the bundled processor/statistics.json for normalization. Do not recompute from a dataset.
  • Do not request chunks longer than 30; the model never saw that range.

Data

34,564 frames / 50 episodes at 30 fps. Internal dataset, not public.

Downloads last month
13
Safetensors
Model size
3B params
Tensor type
BF16
ยท
Video Preview
loading

Model tree for arunos728/gr00t-n16-pointwam-assembletissue-chunk30-2gpu-b64-30k

Finetuned
(87)
this model