AdrianLlopart commited on
Commit
ea28e83
·
verified ·
1 Parent(s): 31a209d

docs: HF model card for OpenRAL/rskill-lingbot_va_a1-galaxea_a1-fruit_placement-bf16 v0.1.0

Browse files
Files changed (1) hide show
  1. README.md +155 -0
README.md ADDED
@@ -0,0 +1,155 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: apache-2.0
5
+ pipeline_tag: robotics
6
+ tags:
7
+ - OpenRAL
8
+ - rskill
9
+ - lingbot_va_a1
10
+ - vision-language-action
11
+ - galaxea_a1
12
+ base_model:
13
+ - robbyant/lingbot-va-base
14
+ base_model_relation: finetune
15
+ datasets:
16
+ - pengyue-polaron/nyush-galaxea-a1-fruit-placement-eef-v21
17
+ inference: false
18
+ ---
19
+
20
+ # rskill-lingbot_va_a1-galaxea_a1-fruit_placement-bf16
21
+
22
+ > **OpenRAL rSkill** — a LingBot-VA fruit-placement policy for the Galaxea A1,
23
+ > deployed through OpenRAL's observation, typed-action, safety-kernel, and HAL
24
+ > contracts.
25
+
26
+ This package points to the public checkpoint at
27
+ [`pengyue-polaron/lingbot-va-galaxea-a1-fruit-placement-eef`](https://huggingface.co/pengyue-polaron/lingbot-va-galaxea-a1-fruit-placement-eef)
28
+ and does not copy model weights into the OpenRAL repository.
29
+
30
+ ## Preview
31
+
32
+ ![Galaxea A1 fruit-placement scene](https://huggingface.co/pengyue-polaron/lingbot-va-galaxea-a1-fruit-placement-eef/resolve/90e017bdbc6afac2e441b4634c9192776bbcb8b7/assets/fruit_placement_agent_view_labeled.png)
33
+
34
+ ## What this skill does
35
+
36
+ The policy picks fruit, including a mango, from a tabletop and places it into a
37
+ bowl or plate. It consumes synchronized front and wrist RGB observations and
38
+ predicts episode-relative end-effector pose plus a continuous normalized gripper
39
+ command.
40
+
41
+ | Field | Value |
42
+ | --- | --- |
43
+ | Actions | `pick`, `place` |
44
+ | Objects | `fruit`, `mango`, `bowl`, `plate` |
45
+ | Scene | `tabletop` |
46
+ | Embodiment | `galaxea_a1` |
47
+
48
+ ## Upstream model and training
49
+
50
+ The checkpoint is a full-parameter fine-tune of
51
+ [`robbyant/lingbot-va-base`](https://huggingface.co/robbyant/lingbot-va-base).
52
+ It jointly predicts video latents and robot-action channels. Training used 130
53
+ episodes and 44,824 frames at 30 FPS from the revision-pinned
54
+ [`nyush-galaxea-a1-fruit-placement-eef-v21`](https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-fruit-placement-eef-v21)
55
+ dataset. The run used 1,000 optimizer steps, two NVIDIA H100 80 GB GPUs,
56
+ full-parameter FSDP, bfloat16, and an effective global batch size of 16.
57
+
58
+ The model emits 16 EEF/gripper steps per chunk. Quantile normalization and the
59
+ action-channel map `[0, 1, 2, 3, 4, 5, 6, 28]` are applied in the external
60
+ LingBot server from the checkpoint's `configs/va_a1_cfg.py`; this is not a
61
+ LeRobot `PolicyProcessorPipeline`.
62
+
63
+ The A1 Runtime policy gateway validates each physical EEF target, solves IK,
64
+ and emits six absolute joint targets plus one normalized gripper target. If an
65
+ IK solution is farther than the rSkill's feedback-relative joint-step bound,
66
+ the gateway advances toward that same solution on subsequent 30 Hz ticks and
67
+ does not consume the next model action until the full solved target can be
68
+ dispatched. It then writes the dispatched target's FK result into the LingBot
69
+ KV cache. OpenRAL's thin adapter validates the gateway model and robot contract,
70
+ then routes the typed proposal through the normal candidate-action, C++ safety
71
+ kernel, safe-action, and Galaxea A1 HAL path. Runtime's IK is constructed with
72
+ the active OpenRAL robot manifest's ordered joint limits, so Runtime calibration
73
+ margins cannot widen the official command envelope.
74
+
75
+ ## Sensors and observation contract
76
+
77
+ | Direction | Key | Shape | Notes |
78
+ | --- | --- | --- | --- |
79
+ | in | `observation.images.front` | `(480, 480, 3)` RGB uint8 | Cropped D455 front view from the A1 Runtime Camera Bridge |
80
+ | in | `observation.images.wrist` | `(480, 640, 3)` RGB uint8 | D405 wrist view from the same paired Camera Bridge |
81
+ | in | `observation.state` | `(6,)` float32 | Six A1 arm joints in radians |
82
+ | out | action | `(7,)` float32 | Six absolute joint targets in radians and one normalized gripper target |
83
+
84
+ The A1 Runtime remains the sole camera-device owner. OpenRAL connects to its
85
+ versioned paired Camera Bridge and policy gateway over private per-user Unix
86
+ sockets; it neither imports the Runtime checkout nor opens either RealSense
87
+ device.
88
+
89
+ ## Supported robots
90
+
91
+ | Robot | Embodiment tag | Status | Notes |
92
+ | --- | --- | --- | --- |
93
+ | Galaxea A1, original arm | `galaxea_a1` | Hardware-in-the-loop integration | Joint/gripper round trips and one visually verified model-driven lemon pick-and-place have passed through OpenRAL; automatic task adjudication remains pending |
94
+
95
+ This rSkill is specific to the six-joint Galaxea A1 contract in
96
+ `robots/galaxea_a1/robot.yaml`. It is not a generic Cartesian HAL and does not
97
+ enable the vendor AnyGrasp/AnyEffector path.
98
+
99
+ ## Manifest summary
100
+
101
+ | Field | Value |
102
+ | --- | --- |
103
+ | `name` | `OpenRAL/rskill-lingbot_va_a1-galaxea_a1-fruit_placement-bf16` |
104
+ | `version` | `0.1.0` |
105
+ | `license` | `apache-2.0` |
106
+ | `model_family` | `lingbot_va_a1` |
107
+ | `embodiment_tags` | `galaxea_a1` |
108
+ | `runtime` / precision | `pytorch` / `bf16` |
109
+ | `weights_uri` | revision-pinned public LingBot-VA A1 checkpoint |
110
+ | `state_contract.dim` / `action_contract.dim` | `6` / `7` |
111
+ | `chunk_size` / `n_action_steps` | `16` / `8` |
112
+ | `latency_budget.per_chunk_ms` | `6000` |
113
+ | `commercial_use_allowed` | `true` |
114
+
115
+ Full schema: [`openral_core.schemas.RSkillManifest`](../../python/core/src/openral_core/schemas.py).
116
+
117
+ ## Quick start
118
+
119
+ The software-only compatibility check does not initialize ROS or hardware:
120
+
121
+ ```bash
122
+ uv run --group lingbot openral rskill check \
123
+ rskills/lingbot-va-galaxea-a1-fruit-placement/rskill.yaml \
124
+ --robot robots/galaxea_a1/robot.yaml
125
+ ```
126
+
127
+ For real deployment, follow the owner-separated startup sequence in
128
+ [`docs/methods/01-hal.md`](../../docs/methods/01-hal.md): start the A1 Runtime
129
+ camera owner, LingBot server, and policy gateway; start the isolated OpenRAL
130
+ ROS1 sidecar; then run the OpenRAL deployment scene. Do not start the A1
131
+ Runtime joint execution bridge at the same time.
132
+
133
+ ## Evaluation
134
+
135
+ No formal automatic task-success result is shipped yet. Validation has
136
+ exercised the real paired cameras, real model server, manifest-to-policy
137
+ construction, EEF validation and IK, joint-step subdivision, and OpenRAL typed
138
+ joint/gripper dispatch. Separate real-hardware joint/gripper round trips and a
139
+ visually verified lemon pick-and-place have traversed candidate action, the C++
140
+ safety kernel, safe action, HAL, ROS 1 relay, and the official driver. The
141
+ non-terminating VLA continued after the visible placement and was stopped by
142
+ the unchanged joint-solution jump guard; a success detector or bounded episode
143
+ termination is still needed for a formal task-success result.
144
+
145
+ ## License
146
+
147
+ This rSkill package and the revision-pinned model weights are Apache-2.0. The
148
+ model repository contains the authoritative `LICENSE.txt`; the OpenRAL package
149
+ references the weights and does not redistribute them.
150
+
151
+ ## See also
152
+
153
+ - [`robots/galaxea_a1/robot.yaml`](../../robots/galaxea_a1/robot.yaml)
154
+ - [`scenes/deploy/galaxea_a1_bench.yaml`](../../scenes/deploy/galaxea_a1_bench.yaml)
155
+ - [`docs/methods/01-hal.md`](../../docs/methods/01-hal.md)