rldx_1_robodojo_taco_b64_30k

Partial run of the RoboDojo TACO low-level policy: RLWRLD/RLDX-1-PT fine-tuned on Myungkyu/RoboDojo-taco-gemini (8 long-horizon bimanual tasks, 100 demonstrations each, dense subtask labels from the task-specific context). The training was preempted at step 30000 of the planned 60000 (2026-09-12); the checkpoints are kept for resumption and analysis.

  • Architecture: RLDX-1-PT (video length 4, three live camera views, no keyframe slot)
  • Optimizer batch 64 (launched as 128 = 64 x gradient accumulation 2; the DeepSpeed ZeRO-2 build applied the accumulation without the intended scaling, so the effective batch was 64), cosine schedule for 60000 steps, seed 42
  • checkpoint-30000/: complete checkpoint — weights + global_step30000/ DeepSpeed optimizer states + RNG states + latest (resume with the RLDX-1 trainer)
  • checkpoint-20000/, checkpoint-25000/: weights only (config, index, safetensors shards, experiment_cfg/, processor/)
  • Each checkpoint folder carries a SHA256SUMS of its weight files

Configs reference the base backbone / tokenizer by hub id or by the training site's local path — point them at your local copies before loading.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for Myungkyu/rldx_1_robodojo_taco_b64_30k

Finetuned
RLWRLD/RLDX-1-PT
Finetuned
(17)
this model

Dataset used to train Myungkyu/rldx_1_robodojo_taco_b64_30k