LingBot-VA grasp-bag-strap DR v1

LingBot-VA transformer fine-tuned on the deformable-bench/grasp_bag_strap_dr_v1 dataset.

Training setup

  • Base model: robbyant/lingbot-va-base
  • Dataset revision: 239700e0cc56136cdc03a1d3625c1064e6fd48b9
  • Dataset size: 200 episodes and 83,566 frames
  • Cameras: static, left hand, and right hand
  • Action/state dimension: 14
  • Latent extraction frame stride: 4
  • Latent extraction shards: 8
  • DataLoader workers: 16
  • GPUs: 8 H20 GPUs
  • Batch size: 1 per rank
  • Learning rate: 1e-5
  • Precision: bf16
  • CFG dropout probability: 0.1
  • Training length: 30,000 optimizer steps

Training loss

The run completed normally at step 30,000. Mean logged losses over selected windows were:

Window Latent Action
Steps 10-100 0.202769 0.024912
Steps 14,901-15,000 0.025104 0.000480
Final 200 steps 0.010343 0.000362
Final logged step 0.008182 0.000339

Both latent and action losses converged steadily without a late rebound.

Files

  • transformer/diffusion_pytorch_model.safetensors: final transformer at step 30,000
  • transformer/config.json: transformer architecture configuration
  • train.log: complete training log

Only the fine-tuned transformer is included. The tokenizer, text encoder, and VAE should be loaded from robbyant/lingbot-va-base.

Intended use

This checkpoint is intended for research and evaluation with the matching LingBot-VA codebase and the grasp_bag_strap_dr_v1 observation/action schema. It has not been validated for safety-critical or out-of-distribution deployment.

Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Video Preview
loading

Model tree for deformable-bench/lingbot-grasp-bag-strap-dr-v1

Finetuned
(13)
this model

Dataset used to train deformable-bench/lingbot-grasp-bag-strap-dr-v1