DPP 6-point β€” human keypoints (100k)

Dexterous Point Policy (DPP) finetuned on 80 anyh2r pick-and-place episodes, using HaWoR human as the source of the 6 hand keypoints.

One of three models that differ only in how the 6 keypoints (wrist + 5 fingertips) were derived. Everything else β€” object point clouds, per-finger contact, camera pose, language, validity mask, episode set, hyper-parameters, pretrained init β€” is byte-identical across the three.

repo keypoint source
dpp-6pt-anyh2r-fk-100k IDM wristIK robot state β†’ MuJoCo FK (ours)
dpp-6pt-anyh2r-human-100k HaWoR human hand (baseline)
dpp-6pt-anyh2r-retgt-100k dex-retargeting β†’ Inspire RH56F1 (baseline)

This model: HaWoR human hand keypoints, used as-is (baseline)

Files

file what
100000.pt final snapshot (323 MB) β€” state_dicts + optimizer + normalisation stats
train_config.yaml / train_overrides.yaml the exact resolved Hydra config of the run
train_log.csv per-100-step training log

Training

Code: beomjun02/dex-point-policy @ e972ba5 (private). Init: vitra_6points_final/snapshot/100000.pt (VITRA video pretrain, 6points, no contact head β€” the head is added at finetune and loaded with strict=False).

agent=dpp  suite.action_mode=6points  suite.history_len=1  num_queries=16
dataloader.bc_dataset.normalize=false          # raw metres, matches the pretrain run
dataloader.bc_dataset.hand=right
dataloader.bc_dataset.max_num_objects=4
teacher_forcing=true
use_contact_head=true  contact_loss_coef=1.0  contact_detach=true
uniform_contact_pos_weight=true                # scalar pos_weight 4.38
geom_drop_p=0  contact_drop_p=0  robot_noise_std=0
batch_size=64  lr=1e-4  steps=100000  finetune=true

1Γ— GPU, ~55 min. Final training loss (mean over steps 95k–100k): 0.1083.

Training losses are not comparable across the three models β€” each fits a different target sequence. Use rollout success, not this number.

Data

80 episodes (ball 22 / bottle 20 / box 22 / doll 16), 22,454 frames @ 20 fps, 95.6% kept. Built from dpp_ego3d_v1/openarm_inspire (pt512 pack) for objects / contact / camera / language, intersected with the anyh2r wristIK synthetic set (idm_v4form_rh56f1_vla_stereo-depth-wristIK_20260916) so all three variants cover exactly the same clips.

  • Frame: ego_right_optical_frame (x right, y down, z forward), metres, static camera.
  • Keypoints: MANO-21 indices [0, 4, 8, 12, 16, 20] = wrist, thumb, index, middle, ring, little tip.
  • Objects: 128 points per slot, 2 slots (task object + white container).
  • Contact: per-finger binary, thresholded at 0.3 on the HaWoR-tactile probability.

Verification of the FK path: our forward kinematics reproduces ego3d's own robot/*/hands21.npz to 0.00 mm (dohyeon calibrated ego extrinsic, rig in MJCF body pitch_motor_rotor_2). Mean disagreement between the human and FK keypoints on the same frames: wrist 36 mm, thumb 42, index 43, middle 46, ring 47, little 41 β€” note the wrist points are different physical landmarks (MANO wrist joint vs RH56F1 palm site), so a steady offset there is geometry, not error.

License

Inherits the dex-point-policy / NVIDIA Source Code License terms of the upstream point_bridge project. Research use.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading