pengyue-polaron commited on
Commit
62ccdbd
·
verified ·
1 Parent(s): eb541f8

Make model card concise and move training details out of the overview

Browse files
Files changed (2) hide show
  1. README.md +10 -32
  2. TRAINING.md +22 -0
README.md CHANGED
@@ -19,21 +19,13 @@ tags:
19
 
20
  # LingBot-VA — Galaxea A1 Mango-to-Plate EEF — Step 100
21
 
22
- This is a full-parameter fine-tune of
23
- [`robbyant/lingbot-va-base`](https://huggingface.co/robbyant/lingbot-va-base)
24
- for placing a red mango into a blue plate with a Galaxea A1 robot using
25
- episode-relative end-effector (EEF) actions. The model jointly predicts video
26
- latents and robot actions.
27
 
28
- The tokenizer, text encoder, and VAE are inherited unchanged from LingBot-VA
29
- base; `transformer/` contains the fine-tuned step-100 weights.
30
 
31
- ## Training data
32
-
33
- The model was trained on the red-mango-to-blue-plate task from
34
- [`pengyue-polaron/nyush-galaxea-a1-fruit-placement-eef-v21`](https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-fruit-placement-eef-v21),
35
- a LeRobot Dataset v2.1 collection with episode-relative EEF pose actions and a
36
- continuous normalized gripper command.
37
 
38
  | Field | Value |
39
  | --- | --- |
@@ -43,27 +35,13 @@ continuous normalized gripper command.
43
  | Cameras | `front` 480×480 RGB; `wrist` 640×480 RGB |
44
  | Action | Episode-relative EEF pose + continuous normalized gripper |
45
 
46
- ## Training
47
-
48
- | Setting | Value |
49
- | --- | --- |
50
- | Optimizer steps | 100 |
51
- | Hardware | 2 × NVIDIA A100 |
52
- | Distributed strategy | Full-parameter FSDP |
53
- | Precision | bfloat16 |
54
- | Effective global batch size | 16 |
55
- | Optimizer | Fused AdamW |
56
- | Learning rate | 1e-5 with 10-step warmup, then constant |
57
- | Adam betas / weight decay | (0.9, 0.95) / 0.1 |
58
- | Objective | Video latent loss + action loss |
59
 
60
- Across the 100 optimizer steps, the mean video-latent loss was `0.161962` and
61
- the mean action loss was `0.020129`.
 
62
 
63
- The eight source action values are mapped to LingBot-VA action channels
64
- `[0, 1, 2, 3, 4, 5, 6, 28]`; channel 28 carries the gripper command. The exact
65
- normalization statistics and model settings are included in
66
- `configs/va_a1_cfg.py`.
67
 
68
  ## License
69
 
 
19
 
20
  # LingBot-VA — Galaxea A1 Mango-to-Plate EEF — Step 100
21
 
22
+ [LingBot-VA](https://huggingface.co/robbyant/lingbot-va-base) fine-tuned for
23
+ red-mango placement onto a blue plate with a Galaxea A1 arm.
24
+ The model predicts robot actions and video latents.
 
 
25
 
26
+ ## Data
 
27
 
28
+ [Training demonstrations](https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-fruit-placement-eef-v21).
 
 
 
 
 
29
 
30
  | Field | Value |
31
  | --- | --- |
 
35
  | Cameras | `front` 480×480 RGB; `wrist` 640×480 RGB |
36
  | Action | Episode-relative EEF pose + continuous normalized gripper |
37
 
38
+ ## Files and configuration
 
 
 
 
 
 
 
 
 
 
 
 
39
 
40
+ `transformer/` contains the step-100 weights. The base tokenizer, text encoder,
41
+ and VAE are included. Use [configs/va_a1_cfg.py](configs/va_a1_cfg.py) for the
42
+ EEF action mapping and normalization.
43
 
44
+ [Training details](TRAINING.md) · [Training summary](training_summary.json)
 
 
 
45
 
46
  ## License
47
 
TRAINING.md ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Training
2
+
3
+ | Setting | Value |
4
+ | --- | --- |
5
+ | Optimizer steps | 100 |
6
+ | Hardware | 2 × NVIDIA A100 |
7
+ | Distributed strategy | Full-parameter FSDP |
8
+ | Precision | bfloat16 |
9
+ | Effective global batch size | 16 |
10
+ | Optimizer | Fused AdamW |
11
+ | Learning rate | 1e-5 with 10-step warmup, then constant |
12
+ | Adam betas / weight decay | (0.9, 0.95) / 0.1 |
13
+ | Objective | Video latent loss + action loss |
14
+
15
+ Across the 100 optimizer steps, the mean video-latent loss was `0.161962` and
16
+ the mean action loss was `0.020129`.
17
+
18
+ The eight source action values are mapped to LingBot-VA action channels
19
+ `[0, 1, 2, 3, 4, 5, 6, 28]`; channel 28 carries the gripper command. The exact
20
+ normalization statistics and model settings are included in
21
+ `configs/va_a1_cfg.py`.
22
+