Instructions to use arunos728/fastwam-robodojo-precision8-vert576-8gpu-b64-step80k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Wan2.2
How to use arunos728/fastwam-robodojo-precision8-vert576-8gpu-b64-step80k with Wan2.2:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
FastWAM โ RoboDojo Precision suite, vertical 576x256
FastWAM (Wan2.2-TI2V-5B backbone) trained on the RoboDojo Precision suite, 8 tasks. The full 100,000-step run, plus the 80k checkpoint kept for comparison.
Use step_100000.pt. The cosine schedule anneals hard at the end: val_loss went
0.1011 at 80k to 0.0272 at 100k. The 80k file was uploaded while the run was still in
flight and is kept only so the difference is checkable, not because it is a useful
alternative.
Training setup
| base model | Wan-AI/Wan2.2-TI2V-5B |
| ActionDiT backbone | linear-interpolated from the Wan2.2 DiT, alpha-scaled, 1024 hidden |
| steps | 100,000 |
| GPUs | 8x H100 80GB, DeepSpeed ZeRO-1 |
| batch | 8 per GPU = global batch 64 |
| lr | 1e-4 cosine, weight decay 1e-2 |
| grad accumulation | 1 |
| val_loss | 0.0272 at 100k (0.1011 at 80k) |
Canvas and cameras
The vertical 3-camera stack, not the T ("robotwin") composite:
576x256 canvas = three 192x256 tiles stacked top to bottom
cam_high
cam_left_wrist
cam_right_wrist
Each tile is a uniform 0.8x of the 240x320 source, so the 4:3 aspect is preserved exactly and the canvas equals the tile sum โ no crop. The alternative T layout on this dataset squeezes every tile to 0.8:1; this arm exists to avoid that.
num_frames 33, action_video_freq_ratio 4 -> 9 video frames + 32 action steps
at 25 fps that is 1.28 s of future action per window
Data
RoboDojo Precision suite, 8 tasks x 100 episodes = 800 episodes / 368,459 frames at 25 fps, 3 cameras at 320x240 H.264, 14-D action and state (left arm 0:6, left gripper 6:7, right arm 7:13, right gripper 13:14).
Built from RoboDojo-Benchmark/RoboDojo by pulling the 8 Precision-dimension blocks and
re-indexing to a standalone dataset. Tasks: build_tower, deposit_coin, fasten_screws,
insert_key, insert_tubes, play_Xylophone, plug_in_charger, pour_balls_into_vase.
What is in here
step_100000.pt final weights (12 GB) -- use this one
step_080000.pt intermediate, kept for comparison
dataset_stats.json z-score normalization computed over the 800-episode subset
config.yaml the resolved training config for this run
dataset_stats.json is required at inference. The normalization is z-score computed
from this dataset; feeding un-normalized actions, or stats from a different corpus, gives
garbage. The 80 GB state/ directory (optimizer moments, ZeRO shards, RNG) is not
included โ this checkpoint is for inference, not for resuming.
Notes
- Trained with
mot_checkpoint_mixed_attn: falseand the uncond (no action conditioning on the video DiT) recipe. - The run was interrupted once at step ~58.9k and resumed from the 55k state checkpoint โ optimizer moments, LR schedule position and step counter all restored, so the cosine schedule is intact. Resuming from weights alone would have restarted the schedule and is not what happened here.
- Downloads last month
- -
Model tree for arunos728/fastwam-robodojo-precision8-vert576-8gpu-b64-step80k
Base model
Wan-AI/Wan2.2-TI2V-5B