jerseyjerry commited on
Commit
7d66dc6
·
verified ·
1 Parent(s): 6bf4763

Uploading model files

Browse files
ATTRIBUTION.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # Attribution
2
+
3
+ This policy includes weights derived from `facebook/metaclip-b16-fullcc2.5b` by Meta Platforms, Inc.
4
+
5
+ MetaCLIP model page: https://huggingface.co/facebook/metaclip-b16-fullcc2.5b
6
+
7
+ MetaCLIP paper: https://arxiv.org/abs/2309.16671
8
+
9
+ Licensed under Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0).
CC-BY-NC-4.0.txt ADDED
@@ -0,0 +1,408 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Attribution-NonCommercial 4.0 International
2
+
3
+ =======================================================================
4
+
5
+ Creative Commons Corporation ("Creative Commons") is not a law firm and
6
+ does not provide legal services or legal advice. Distribution of
7
+ Creative Commons public licenses does not create a lawyer-client or
8
+ other relationship. Creative Commons makes its licenses and related
9
+ information available on an "as-is" basis. Creative Commons gives no
10
+ warranties regarding its licenses, any material licensed under their
11
+ terms and conditions, or any related information. Creative Commons
12
+ disclaims all liability for damages resulting from their use to the
13
+ fullest extent possible.
14
+
15
+ Using Creative Commons Public Licenses
16
+
17
+ Creative Commons public licenses provide a standard set of terms and
18
+ conditions that creators and other rights holders may use to share
19
+ original works of authorship and other material subject to copyright
20
+ and certain other rights specified in the public license below. The
21
+ following considerations are for informational purposes only, are not
22
+ exhaustive, and do not form part of our licenses.
23
+
24
+ Considerations for licensors: Our public licenses are
25
+ intended for use by those authorized to give the public
26
+ permission to use material in ways otherwise restricted by
27
+ copyright and certain other rights. Our licenses are
28
+ irrevocable. Licensors should read and understand the terms
29
+ and conditions of the license they choose before applying it.
30
+ Licensors should also secure all rights necessary before
31
+ applying our licenses so that the public can reuse the
32
+ material as expected. Licensors should clearly mark any
33
+ material not subject to the license. This includes other CC-
34
+ licensed material, or material used under an exception or
35
+ limitation to copyright. More considerations for licensors:
36
+ wiki.creativecommons.org/Considerations_for_licensors
37
+
38
+ Considerations for the public: By using one of our public
39
+ licenses, a licensor grants the public permission to use the
40
+ licensed material under specified terms and conditions. If
41
+ the licensor's permission is not necessary for any reason--for
42
+ example, because of any applicable exception or limitation to
43
+ copyright--then that use is not regulated by the license. Our
44
+ licenses grant only permissions under copyright and certain
45
+ other rights that a licensor has authority to grant. Use of
46
+ the licensed material may still be restricted for other
47
+ reasons, including because others have copyright or other
48
+ rights in the material. A licensor may make special requests,
49
+ such as asking that all changes be marked or described.
50
+ Although not required by our licenses, you are encouraged to
51
+ respect those requests where reasonable. More considerations
52
+ for the public:
53
+ wiki.creativecommons.org/Considerations_for_licensees
54
+
55
+ =======================================================================
56
+
57
+ Creative Commons Attribution-NonCommercial 4.0 International Public
58
+ License
59
+
60
+ By exercising the Licensed Rights (defined below), You accept and agree
61
+ to be bound by the terms and conditions of this Creative Commons
62
+ Attribution-NonCommercial 4.0 International Public License ("Public
63
+ License"). To the extent this Public License may be interpreted as a
64
+ contract, You are granted the Licensed Rights in consideration of Your
65
+ acceptance of these terms and conditions, and the Licensor grants You
66
+ such rights in consideration of benefits the Licensor receives from
67
+ making the Licensed Material available under these terms and
68
+ conditions.
69
+
70
+
71
+ Section 1 -- Definitions.
72
+
73
+ a. Adapted Material means material subject to Copyright and Similar
74
+ Rights that is derived from or based upon the Licensed Material
75
+ and in which the Licensed Material is translated, altered,
76
+ arranged, transformed, or otherwise modified in a manner requiring
77
+ permission under the Copyright and Similar Rights held by the
78
+ Licensor. For purposes of this Public License, where the Licensed
79
+ Material is a musical work, performance, or sound recording,
80
+ Adapted Material is always produced where the Licensed Material is
81
+ synched in timed relation with a moving image.
82
+
83
+ b. Adapter's License means the license You apply to Your Copyright
84
+ and Similar Rights in Your contributions to Adapted Material in
85
+ accordance with the terms and conditions of this Public License.
86
+
87
+ c. Copyright and Similar Rights means copyright and/or similar rights
88
+ closely related to copyright including, without limitation,
89
+ performance, broadcast, sound recording, and Sui Generis Database
90
+ Rights, without regard to how the rights are labeled or
91
+ categorized. For purposes of this Public License, the rights
92
+ specified in Section 2(b)(1)-(2) are not Copyright and Similar
93
+ Rights.
94
+ d. Effective Technological Measures means those measures that, in the
95
+ absence of proper authority, may not be circumvented under laws
96
+ fulfilling obligations under Article 11 of the WIPO Copyright
97
+ Treaty adopted on December 20, 1996, and/or similar international
98
+ agreements.
99
+
100
+ e. Exceptions and Limitations means fair use, fair dealing, and/or
101
+ any other exception or limitation to Copyright and Similar Rights
102
+ that applies to Your use of the Licensed Material.
103
+
104
+ f. Licensed Material means the artistic or literary work, database,
105
+ or other material to which the Licensor applied this Public
106
+ License.
107
+
108
+ g. Licensed Rights means the rights granted to You subject to the
109
+ terms and conditions of this Public License, which are limited to
110
+ all Copyright and Similar Rights that apply to Your use of the
111
+ Licensed Material and that the Licensor has authority to license.
112
+
113
+ h. Licensor means the individual(s) or entity(ies) granting rights
114
+ under this Public License.
115
+
116
+ i. NonCommercial means not primarily intended for or directed towards
117
+ commercial advantage or monetary compensation. For purposes of
118
+ this Public License, the exchange of the Licensed Material for
119
+ other material subject to Copyright and Similar Rights by digital
120
+ file-sharing or similar means is NonCommercial provided there is
121
+ no payment of monetary compensation in connection with the
122
+ exchange.
123
+
124
+ j. Share means to provide material to the public by any means or
125
+ process that requires permission under the Licensed Rights, such
126
+ as reproduction, public display, public performance, distribution,
127
+ dissemination, communication, or importation, and to make material
128
+ available to the public including in ways that members of the
129
+ public may access the material from a place and at a time
130
+ individually chosen by them.
131
+
132
+ k. Sui Generis Database Rights means rights other than copyright
133
+ resulting from Directive 96/9/EC of the European Parliament and of
134
+ the Council of 11 March 1996 on the legal protection of databases,
135
+ as amended and/or succeeded, as well as other essentially
136
+ equivalent rights anywhere in the world.
137
+
138
+ l. You means the individual or entity exercising the Licensed Rights
139
+ under this Public License. Your has a corresponding meaning.
140
+
141
+
142
+ Section 2 -- Scope.
143
+
144
+ a. License grant.
145
+
146
+ 1. Subject to the terms and conditions of this Public License,
147
+ the Licensor hereby grants You a worldwide, royalty-free,
148
+ non-sublicensable, non-exclusive, irrevocable license to
149
+ exercise the Licensed Rights in the Licensed Material to:
150
+
151
+ a. reproduce and Share the Licensed Material, in whole or
152
+ in part, for NonCommercial purposes only; and
153
+
154
+ b. produce, reproduce, and Share Adapted Material for
155
+ NonCommercial purposes only.
156
+
157
+ 2. Exceptions and Limitations. For the avoidance of doubt, where
158
+ Exceptions and Limitations apply to Your use, this Public
159
+ License does not apply, and You do not need to comply with
160
+ its terms and conditions.
161
+
162
+ 3. Term. The term of this Public License is specified in Section
163
+ 6(a).
164
+
165
+ 4. Media and formats; technical modifications allowed. The
166
+ Licensor authorizes You to exercise the Licensed Rights in
167
+ all media and formats whether now known or hereafter created,
168
+ and to make technical modifications necessary to do so. The
169
+ Licensor waives and/or agrees not to assert any right or
170
+ authority to forbid You from making technical modifications
171
+ necessary to exercise the Licensed Rights, including
172
+ technical modifications necessary to circumvent Effective
173
+ Technological Measures. For purposes of this Public License,
174
+ simply making modifications authorized by this Section 2(a)
175
+ (4) never produces Adapted Material.
176
+
177
+ 5. Downstream recipients.
178
+
179
+ a. Offer from the Licensor -- Licensed Material. Every
180
+ recipient of the Licensed Material automatically
181
+ receives an offer from the Licensor to exercise the
182
+ Licensed Rights under the terms and conditions of this
183
+ Public License.
184
+
185
+ b. No downstream restrictions. You may not offer or impose
186
+ any additional or different terms or conditions on, or
187
+ apply any Effective Technological Measures to, the
188
+ Licensed Material if doing so restricts exercise of the
189
+ Licensed Rights by any recipient of the Licensed
190
+ Material.
191
+
192
+ 6. No endorsement. Nothing in this Public License constitutes or
193
+ may be construed as permission to assert or imply that You
194
+ are, or that Your use of the Licensed Material is, connected
195
+ with, or sponsored, endorsed, or granted official status by,
196
+ the Licensor or others designated to receive attribution as
197
+ provided in Section 3(a)(1)(A)(i).
198
+
199
+ b. Other rights.
200
+
201
+ 1. Moral rights, such as the right of integrity, are not
202
+ licensed under this Public License, nor are publicity,
203
+ privacy, and/or other similar personality rights; however, to
204
+ the extent possible, the Licensor waives and/or agrees not to
205
+ assert any such rights held by the Licensor to the limited
206
+ extent necessary to allow You to exercise the Licensed
207
+ Rights, but not otherwise.
208
+
209
+ 2. Patent and trademark rights are not licensed under this
210
+ Public License.
211
+
212
+ 3. To the extent possible, the Licensor waives any right to
213
+ collect royalties from You for the exercise of the Licensed
214
+ Rights, whether directly or through a collecting society
215
+ under any voluntary or waivable statutory or compulsory
216
+ licensing scheme. In all other cases the Licensor expressly
217
+ reserves any right to collect such royalties, including when
218
+ the Licensed Material is used other than for NonCommercial
219
+ purposes.
220
+
221
+
222
+ Section 3 -- License Conditions.
223
+
224
+ Your exercise of the Licensed Rights is expressly made subject to the
225
+ following conditions.
226
+
227
+ a. Attribution.
228
+
229
+ 1. If You Share the Licensed Material (including in modified
230
+ form), You must:
231
+
232
+ a. retain the following if it is supplied by the Licensor
233
+ with the Licensed Material:
234
+
235
+ i. identification of the creator(s) of the Licensed
236
+ Material and any others designated to receive
237
+ attribution, in any reasonable manner requested by
238
+ the Licensor (including by pseudonym if
239
+ designated);
240
+
241
+ ii. a copyright notice;
242
+
243
+ iii. a notice that refers to this Public License;
244
+
245
+ iv. a notice that refers to the disclaimer of
246
+ warranties;
247
+
248
+ v. a URI or hyperlink to the Licensed Material to the
249
+ extent reasonably practicable;
250
+
251
+ b. indicate if You modified the Licensed Material and
252
+ retain an indication of any previous modifications; and
253
+
254
+ c. indicate the Licensed Material is licensed under this
255
+ Public License, and include the text of, or the URI or
256
+ hyperlink to, this Public License.
257
+
258
+ 2. You may satisfy the conditions in Section 3(a)(1) in any
259
+ reasonable manner based on the medium, means, and context in
260
+ which You Share the Licensed Material. For example, it may be
261
+ reasonable to satisfy the conditions by providing a URI or
262
+ hyperlink to a resource that includes the required
263
+ information.
264
+
265
+ 3. If requested by the Licensor, You must remove any of the
266
+ information required by Section 3(a)(1)(A) to the extent
267
+ reasonably practicable.
268
+
269
+ 4. If You Share Adapted Material You produce, the Adapter's
270
+ License You apply must not prevent recipients of the Adapted
271
+ Material from complying with this Public License.
272
+
273
+
274
+ Section 4 -- Sui Generis Database Rights.
275
+
276
+ Where the Licensed Rights include Sui Generis Database Rights that
277
+ apply to Your use of the Licensed Material:
278
+
279
+ a. for the avoidance of doubt, Section 2(a)(1) grants You the right
280
+ to extract, reuse, reproduce, and Share all or a substantial
281
+ portion of the contents of the database for NonCommercial purposes
282
+ only;
283
+
284
+ b. if You include all or a substantial portion of the database
285
+ contents in a database in which You have Sui Generis Database
286
+ Rights, then the database in which You have Sui Generis Database
287
+ Rights (but not its individual contents) is Adapted Material; and
288
+
289
+ c. You must comply with the conditions in Section 3(a) if You Share
290
+ all or a substantial portion of the contents of the database.
291
+
292
+ For the avoidance of doubt, this Section 4 supplements and does not
293
+ replace Your obligations under this Public License where the Licensed
294
+ Rights include other Copyright and Similar Rights.
295
+
296
+
297
+ Section 5 -- Disclaimer of Warranties and Limitation of Liability.
298
+
299
+ a. UNLESS OTHERWISE SEPARATELY UNDERTAKEN BY THE LICENSOR, TO THE
300
+ EXTENT POSSIBLE, THE LICENSOR OFFERS THE LICENSED MATERIAL AS-IS
301
+ AND AS-AVAILABLE, AND MAKES NO REPRESENTATIONS OR WARRANTIES OF
302
+ ANY KIND CONCERNING THE LICENSED MATERIAL, WHETHER EXPRESS,
303
+ IMPLIED, STATUTORY, OR OTHER. THIS INCLUDES, WITHOUT LIMITATION,
304
+ WARRANTIES OF TITLE, MERCHANTABILITY, FITNESS FOR A PARTICULAR
305
+ PURPOSE, NON-INFRINGEMENT, ABSENCE OF LATENT OR OTHER DEFECTS,
306
+ ACCURACY, OR THE PRESENCE OR ABSENCE OF ERRORS, WHETHER OR NOT
307
+ KNOWN OR DISCOVERABLE. WHERE DISCLAIMERS OF WARRANTIES ARE NOT
308
+ ALLOWED IN FULL OR IN PART, THIS DISCLAIMER MAY NOT APPLY TO YOU.
309
+
310
+ b. TO THE EXTENT POSSIBLE, IN NO EVENT WILL THE LICENSOR BE LIABLE
311
+ TO YOU ON ANY LEGAL THEORY (INCLUDING, WITHOUT LIMITATION,
312
+ NEGLIGENCE) OR OTHERWISE FOR ANY DIRECT, SPECIAL, INDIRECT,
313
+ INCIDENTAL, CONSEQUENTIAL, PUNITIVE, EXEMPLARY, OR OTHER LOSSES,
314
+ COSTS, EXPENSES, OR DAMAGES ARISING OUT OF THIS PUBLIC LICENSE OR
315
+ USE OF THE LICENSED MATERIAL, EVEN IF THE LICENSOR HAS BEEN
316
+ ADVISED OF THE POSSIBILITY OF SUCH LOSSES, COSTS, EXPENSES, OR
317
+ DAMAGES. WHERE A LIMITATION OF LIABILITY IS NOT ALLOWED IN FULL OR
318
+ IN PART, THIS LIMITATION MAY NOT APPLY TO YOU.
319
+
320
+ c. The disclaimer of warranties and limitation of liability provided
321
+ above shall be interpreted in a manner that, to the extent
322
+ possible, most closely approximates an absolute disclaimer and
323
+ waiver of all liability.
324
+
325
+
326
+ Section 6 -- Term and Termination.
327
+
328
+ a. This Public License applies for the term of the Copyright and
329
+ Similar Rights licensed here. However, if You fail to comply with
330
+ this Public License, then Your rights under this Public License
331
+ terminate automatically.
332
+
333
+ b. Where Your right to use the Licensed Material has terminated under
334
+ Section 6(a), it reinstates:
335
+
336
+ 1. automatically as of the date the violation is cured, provided
337
+ it is cured within 30 days of Your discovery of the
338
+ violation; or
339
+
340
+ 2. upon express reinstatement by the Licensor.
341
+
342
+ For the avoidance of doubt, this Section 6(b) does not affect any
343
+ right the Licensor may have to seek remedies for Your violations
344
+ of this Public License.
345
+
346
+ c. For the avoidance of doubt, the Licensor may also offer the
347
+ Licensed Material under separate terms or conditions or stop
348
+ distributing the Licensed Material at any time; however, doing so
349
+ will not terminate this Public License.
350
+
351
+ d. Sections 1, 5, 6, 7, and 8 survive termination of this Public
352
+ License.
353
+
354
+
355
+ Section 7 -- Other Terms and Conditions.
356
+
357
+ a. The Licensor shall not be bound by any additional or different
358
+ terms or conditions communicated by You unless expressly agreed.
359
+
360
+ b. Any arrangements, understandings, or agreements regarding the
361
+ Licensed Material not stated herein are separate from and
362
+ independent of the terms and conditions of this Public License.
363
+
364
+
365
+ Section 8 -- Interpretation.
366
+
367
+ a. For the avoidance of doubt, this Public License does not, and
368
+ shall not be interpreted to, reduce, limit, restrict, or impose
369
+ conditions on any use of the Licensed Material that could lawfully
370
+ be made without permission under this Public License.
371
+
372
+ b. To the extent possible, if any provision of this Public License is
373
+ deemed unenforceable, it shall be automatically reformed to the
374
+ minimum extent necessary to make it enforceable. If the provision
375
+ cannot be reformed, it shall be severed from this Public License
376
+ without affecting the enforceability of the remaining terms and
377
+ conditions.
378
+
379
+ c. No term or condition of this Public License will be waived and no
380
+ failure to comply consented to unless expressly agreed to by the
381
+ Licensor.
382
+
383
+ d. Nothing in this Public License constitutes or may be interpreted
384
+ as a limitation upon, or waiver of, any privileges and immunities
385
+ that apply to the Licensor or You, including from the legal
386
+ processes of any jurisdiction or authority.
387
+
388
+ =======================================================================
389
+
390
+ Creative Commons is not a party to its public
391
+ licenses. Notwithstanding, Creative Commons may elect to apply one of
392
+ its public licenses to material it publishes and in those instances
393
+ will be considered the “Licensor.” The text of the Creative Commons
394
+ public licenses is dedicated to the public domain under the CC0 Public
395
+ Domain Dedication. Except for the limited purpose of indicating that
396
+ material is shared under a Creative Commons public license or as
397
+ otherwise permitted by the Creative Commons policies published at
398
+ creativecommons.org/policies, Creative Commons does not authorize the
399
+ use of the trademark "Creative Commons" or any other trademark or logo
400
+ of Creative Commons without its prior written consent including,
401
+ without limitation, in connection with any unauthorized modifications
402
+ to any of its public licenses or any other arrangements,
403
+ understandings, or agreements concerning use of licensed material. For
404
+ the avoidance of doubt, this paragraph does not form part of the
405
+ public licenses.
406
+
407
+ Creative Commons may be contacted at creativecommons.org.
408
+
README.md CHANGED
@@ -1,10 +1,8 @@
1
- # DINOv3 temporal action-chunk robotics VLA
2
-
3
- Self-contained policy produced by `train_dinov3_action_chunk_vla.py`.
4
 
5
 
6
  Upload every file in this directory. The adapter loads `model.safetensors`
7
  strictly and performs no model-hub or network access.
8
 
9
- Built with DINOv3. Use and redistribution are governed by the included
10
- `LICENSE.md`.
 
1
+ # MetaCLIP temporal action-chunk robotics VLA
 
 
2
 
3
 
4
  Upload every file in this directory. The adapter loads `model.safetensors`
5
  strictly and performs no model-hub or network access.
6
 
7
+ MetaCLIP is licensed under CC BY-NC 4.0. Commercial use is not permitted by
8
+ that license. See `CC-BY-NC-4.0.txt` and `ATTRIBUTION.md`.
flock_robotics_adapter.py CHANGED
@@ -1,9 +1,9 @@
1
- """Transactional FLock adapter for a self-contained DINOv3 action-chunk policy.
2
 
3
  The exporter copies this file to ``flock_robotics_adapter.py`` next to
4
- ``dinov3_action_chunk_model.py``, ``vla_config.json``, and
5
  ``model.safetensors``. Runtime loading is deliberately local-only: the complete
6
- DINOv3 backbone and policy head must already be present in the safetensors file.
7
 
8
  The validator can retry the same ``policy.act(obs)`` request. State is therefore
9
  copy-on-write and is committed only after a valid action has been produced. An
@@ -294,7 +294,7 @@ def _temporal_ensemble(
294
  return result, tuple(retained)
295
 
296
 
297
- class DINOv3ActionChunkPolicy:
298
  """Stateful action-chunk policy with transactional validator semantics."""
299
 
300
  def __init__(
@@ -304,14 +304,14 @@ class DINOv3ActionChunkPolicy:
304
  device: torch.device,
305
  dtype: torch.dtype,
306
  *,
307
- text_vector_fn: Callable[[str, int], np.ndarray],
308
  phase_vector_fn: Callable[[int, int], np.ndarray],
309
  ) -> None:
310
  self.model = model
311
  self.config = dict(config)
312
  self.device = torch.device(device)
313
  self.dtype = dtype
314
- self.text_vector_fn = text_vector_fn
315
  self.phase_vector_fn = phase_vector_fn
316
 
317
  self.history = int(self.config.get("history", DEFAULT_HISTORY))
@@ -345,6 +345,7 @@ class DINOv3ActionChunkPolicy:
345
  raise ValueError("gripper_hysteresis must be finite and non-negative")
346
 
347
  self._state: _RuntimeState | None = None
 
348
  self.model.to(device=self.device, dtype=self.dtype)
349
  self.model.eval()
350
 
@@ -406,16 +407,26 @@ class DINOv3ActionChunkPolicy:
406
  base.phase_history if base else (), phase_tensor, self.history
407
  )
408
 
409
- text = (
410
- f"task {parsed.task} difficulty {parsed.difficulty} "
411
- f"instruction {parsed.instruction}"
412
- )
413
- text_array = _fixed_feature_array(
414
- self.text_vector_fn(text, self.text_dim), self.text_dim, "text"
415
- )
416
- text_features = torch.as_tensor(
417
- text_array, device=self.device, dtype=self.dtype
418
- ).unsqueeze(0)
 
 
 
 
 
 
 
 
 
 
419
  task_id = self.task_to_id.get(parsed.task, len(self.task_to_id))
420
  difficulty_id = self.difficulty_to_id.get(
421
  parsed.difficulty, len(self.difficulty_to_id)
@@ -470,6 +481,8 @@ class DINOv3ActionChunkPolicy:
470
  gripper_state=gripper_state,
471
  last_action=cached_action,
472
  )
 
 
473
  return action.copy()
474
 
475
 
@@ -506,6 +519,31 @@ def _single_encoded_frame(
506
  return result
507
 
508
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
509
  def _validate_id_map(value: Any, name: str) -> dict[str, int]:
510
  if not isinstance(value, Mapping):
511
  raise ValueError(f"vla_config {name} must be an object")
@@ -523,49 +561,124 @@ def _validate_id_map(value: Any, name: str) -> dict[str, int]:
523
  return result
524
 
525
 
526
- def load_policy(model_dir: str, device: str, dtype: str) -> DINOv3ActionChunkPolicy:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
527
  """Load a complete local checkpoint without any Hub/network fallback."""
528
  model_root = Path(model_dir).expanduser().resolve()
529
  config_path = model_root / "vla_config.json"
530
  weights_path = model_root / "model.safetensors"
531
- runtime_path = model_root / "dinov3_action_chunk_model.py"
532
- for path in (config_path, weights_path, runtime_path):
 
 
533
  if not path.is_file():
534
  raise FileNotFoundError(
535
- f"self-contained DINOv3 submission is missing {path.name}"
536
  )
537
 
538
  config = json.loads(config_path.read_text(encoding="utf-8"))
539
- if not isinstance(config, dict) or not isinstance(
540
- config.get("vision_config"), dict
541
- ):
542
- raise ValueError("vla_config.json must contain an inline vision_config object")
543
  if str(model_root) not in sys.path:
544
  sys.path.insert(0, str(model_root))
545
 
546
  # These imports are intentionally local. The adapter remains easy to test
547
  # with a fake model, and production has no model-Hub loading path.
548
  from safetensors.torch import load_file
549
- from dinov3_action_chunk_model import (
550
- DinoV3ActionChunkModel,
 
551
  phase_vector,
552
- text_vector,
553
  )
554
 
555
  torch_device = _resolve_device(device)
556
  torch_dtype = _resolve_dtype(dtype, torch_device)
557
- model = DinoV3ActionChunkModel(config)
558
  state = load_file(str(weights_path), device="cpu")
559
  model.load_state_dict(state, strict=True)
560
  del state
561
- return DINOv3ActionChunkPolicy(
 
 
 
 
 
 
562
  model=model,
563
  config=config,
564
  device=torch_device,
565
  dtype=torch_dtype,
566
- text_vector_fn=text_vector,
567
  phase_vector_fn=phase_vector,
568
  )
569
 
570
 
571
- __all__ = ["DINOv3ActionChunkPolicy", "load_policy"]
 
1
+ """Transactional FLock adapter for a self-contained MetaCLIP action-chunk policy.
2
 
3
  The exporter copies this file to ``flock_robotics_adapter.py`` next to
4
+ ``metaclip_action_chunk_model.py``, ``vla_config.json``, and
5
  ``model.safetensors``. Runtime loading is deliberately local-only: the complete
6
+ MetaCLIP backbone and policy head must already be present in the safetensors file.
7
 
8
  The validator can retry the same ``policy.act(obs)`` request. State is therefore
9
  copy-on-write and is committed only after a valid action has been produced. An
 
294
  return result, tuple(retained)
295
 
296
 
297
+ class MetaCLIPActionChunkPolicy:
298
  """Stateful action-chunk policy with transactional validator semantics."""
299
 
300
  def __init__(
 
304
  device: torch.device,
305
  dtype: torch.dtype,
306
  *,
307
+ tokenizer: Any,
308
  phase_vector_fn: Callable[[int, int], np.ndarray],
309
  ) -> None:
310
  self.model = model
311
  self.config = dict(config)
312
  self.device = torch.device(device)
313
  self.dtype = dtype
314
+ self.tokenizer = tokenizer
315
  self.phase_vector_fn = phase_vector_fn
316
 
317
  self.history = int(self.config.get("history", DEFAULT_HISTORY))
 
345
  raise ValueError("gripper_hysteresis must be finite and non-negative")
346
 
347
  self._state: _RuntimeState | None = None
348
+ self._text_cache: dict[str, torch.Tensor] = {}
349
  self.model.to(device=self.device, dtype=self.dtype)
350
  self.model.eval()
351
 
 
407
  base.phase_history if base else (), phase_tensor, self.history
408
  )
409
 
410
+ text_key = parsed.instruction.strip()
411
+ pending_text_cache: torch.Tensor | None = None
412
+ cached_text = self._text_cache.get(text_key)
413
+ if cached_text is None:
414
+ encoded = self.tokenizer(
415
+ text_key,
416
+ padding="max_length",
417
+ truncation=True,
418
+ max_length=77,
419
+ return_tensors="pt",
420
+ )
421
+ text_features = self.model.encode_text(
422
+ encoded["input_ids"], encoded["attention_mask"]
423
+ )
424
+ text_features = _single_text_feature(
425
+ text_features, self.text_dim, self.device, self.dtype
426
+ )
427
+ pending_text_cache = text_features.detach().clone()
428
+ else:
429
+ text_features = cached_text
430
  task_id = self.task_to_id.get(parsed.task, len(self.task_to_id))
431
  difficulty_id = self.difficulty_to_id.get(
432
  parsed.difficulty, len(self.difficulty_to_id)
 
481
  gripper_state=gripper_state,
482
  last_action=cached_action,
483
  )
484
+ if pending_text_cache is not None:
485
+ self._text_cache[text_key] = pending_text_cache
486
  return action.copy()
487
 
488
 
 
519
  return result
520
 
521
 
522
+ def _single_text_feature(
523
+ value: Any,
524
+ expected_dim: int,
525
+ device: torch.device,
526
+ dtype: torch.dtype,
527
+ ) -> torch.Tensor:
528
+ if (
529
+ not isinstance(value, torch.Tensor)
530
+ or value.ndim != 2
531
+ or tuple(value.shape) != (1, expected_dim)
532
+ ):
533
+ shape = (
534
+ tuple(value.shape)
535
+ if isinstance(value, torch.Tensor)
536
+ else type(value).__name__
537
+ )
538
+ raise RuntimeError(
539
+ f"text encoder output must be [1,{expected_dim}], got {shape}"
540
+ )
541
+ result = value.detach().to(device=device, dtype=dtype)
542
+ if not bool(torch.isfinite(result.float()).all().item()):
543
+ raise RuntimeError("text encoder output contains NaN or Inf")
544
+ return result
545
+
546
+
547
  def _validate_id_map(value: Any, name: str) -> dict[str, int]:
548
  if not isinstance(value, Mapping):
549
  raise ValueError(f"vla_config {name} must be an object")
 
561
  return result
562
 
563
 
564
+ def _build_local_clip_tokenizer(
565
+ tokenizer_cls: Any,
566
+ vocab_path: Path,
567
+ merges_path: Path,
568
+ clip_config: Mapping[str, Any],
569
+ ) -> Any:
570
+ """Build the bundled byte-BPE tokenizer across Transformers 4.x/5.x."""
571
+
572
+ text_config = clip_config.get("text_config")
573
+ if not isinstance(text_config, Mapping):
574
+ raise ValueError("clip_config.text_config must be an object")
575
+ try:
576
+ expected_vocab_size = int(text_config["vocab_size"])
577
+ expected_bos = int(text_config.get("bos_token_id", expected_vocab_size - 2))
578
+ expected_eos = int(text_config.get("eos_token_id", expected_vocab_size - 1))
579
+ except (KeyError, TypeError, ValueError) as exc:
580
+ raise ValueError(
581
+ "clip_config.text_config has an invalid tokenizer contract"
582
+ ) from exc
583
+ common = {"model_max_length": 77}
584
+ attempts = (
585
+ (
586
+ "transformers_v5",
587
+ {"vocab": str(vocab_path), "merges": str(merges_path)},
588
+ ),
589
+ (
590
+ "transformers_v4",
591
+ {
592
+ "vocab_file": str(vocab_path),
593
+ "merges_file": str(merges_path),
594
+ },
595
+ ),
596
+ )
597
+ failures: list[str] = []
598
+ for label, asset_kwargs in attempts:
599
+ try:
600
+ tokenizer = tokenizer_cls(**asset_kwargs, **common)
601
+ observed = {
602
+ "vocab_size": int(tokenizer.vocab_size),
603
+ "get_vocab_size": len(tokenizer.get_vocab()),
604
+ "bos_token_id": int(tokenizer.bos_token_id),
605
+ "eos_token_id": int(tokenizer.eos_token_id),
606
+ "pad_token_id": int(tokenizer.pad_token_id),
607
+ "unk_token_id": int(tokenizer.unk_token_id),
608
+ }
609
+ except (AttributeError, OSError, TypeError, ValueError) as exc:
610
+ failures.append(f"{label}: {type(exc).__name__}: {exc}")
611
+ continue
612
+ expected = {
613
+ "vocab_size": expected_vocab_size,
614
+ "get_vocab_size": expected_vocab_size,
615
+ "bos_token_id": expected_bos,
616
+ "eos_token_id": expected_eos,
617
+ "pad_token_id": expected_eos,
618
+ "unk_token_id": expected_eos,
619
+ }
620
+ mismatches = {
621
+ name: {"expected": expected[name], "actual": value}
622
+ for name, value in observed.items()
623
+ if value != expected[name]
624
+ }
625
+ if not mismatches:
626
+ return tokenizer
627
+ failures.append(f"{label}: {json.dumps(mismatches, sort_keys=True)}")
628
+ raise ValueError(
629
+ "Could not construct the bundled MetaCLIP tokenizer; " + " | ".join(failures)
630
+ )
631
+
632
+
633
+ def load_policy(model_dir: str, device: str, dtype: str) -> MetaCLIPActionChunkPolicy:
634
  """Load a complete local checkpoint without any Hub/network fallback."""
635
  model_root = Path(model_dir).expanduser().resolve()
636
  config_path = model_root / "vla_config.json"
637
  weights_path = model_root / "model.safetensors"
638
+ runtime_path = model_root / "metaclip_action_chunk_model.py"
639
+ vocab_path = model_root / "vocab.json"
640
+ merges_path = model_root / "merges.txt"
641
+ for path in (config_path, weights_path, runtime_path, vocab_path, merges_path):
642
  if not path.is_file():
643
  raise FileNotFoundError(
644
+ f"self-contained MetaCLIP submission is missing {path.name}"
645
  )
646
 
647
  config = json.loads(config_path.read_text(encoding="utf-8"))
648
+ if not isinstance(config, dict) or not isinstance(config.get("clip_config"), dict):
649
+ raise ValueError("vla_config.json must contain an inline clip_config object")
 
 
650
  if str(model_root) not in sys.path:
651
  sys.path.insert(0, str(model_root))
652
 
653
  # These imports are intentionally local. The adapter remains easy to test
654
  # with a fake model, and production has no model-Hub loading path.
655
  from safetensors.torch import load_file
656
+ from transformers import CLIPTokenizer
657
+ from metaclip_action_chunk_model import (
658
+ MetaCLIPActionChunkModel,
659
  phase_vector,
 
660
  )
661
 
662
  torch_device = _resolve_device(device)
663
  torch_dtype = _resolve_dtype(dtype, torch_device)
664
+ model = MetaCLIPActionChunkModel(config)
665
  state = load_file(str(weights_path), device="cpu")
666
  model.load_state_dict(state, strict=True)
667
  del state
668
+ tokenizer = _build_local_clip_tokenizer(
669
+ CLIPTokenizer,
670
+ vocab_path,
671
+ merges_path,
672
+ config["clip_config"],
673
+ )
674
+ return MetaCLIPActionChunkPolicy(
675
  model=model,
676
  config=config,
677
  device=torch_device,
678
  dtype=torch_dtype,
679
+ tokenizer=tokenizer,
680
  phase_vector_fn=phase_vector,
681
  )
682
 
683
 
684
+ __all__ = ["MetaCLIPActionChunkPolicy", "load_policy"]
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
metaclip_action_chunk_model.py ADDED
@@ -0,0 +1,730 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Self-contained MetaCLIP temporal action-chunk policy definition.
2
+
3
+ This module deliberately contains no model-hub access. A complete
4
+ ``clip_config`` dictionary is embedded in the policy configuration, so
5
+ constructing :class:`MetaCLIPActionChunkModel` only creates modules. Callers are
6
+ responsible for loading a local state dict afterwards.
7
+
8
+ The image, phase, and token helpers are shared by feature-cache creation,
9
+ training, and the submission adapter. Keeping those operations here prevents
10
+ subtle train/deployment preprocessing drift.
11
+ """
12
+
13
+ from __future__ import annotations
14
+
15
+ import copy
16
+ import math
17
+ from collections.abc import Mapping
18
+ from typing import Any
19
+
20
+ import numpy as np
21
+ import torch
22
+ import torch.nn as nn
23
+ import torch.nn.functional as F
24
+ from transformers import CLIPConfig, CLIPModel
25
+
26
+ ACTION_DIM = 7
27
+ DEFAULT_TEXT_DIM = 512
28
+ DEFAULT_PHASE_DIM = 4
29
+ DEFAULT_SPATIAL_HEADS = 8
30
+ DEFAULT_DIFFICULTIES = ("low", "medium", "hard", "very_high")
31
+ TEXT_FEATURE_VERSION = "metaclip_clip_bpe_projected_l2_text_v2"
32
+ METACLIP_IMAGE_MEAN = (0.48145466, 0.4578275, 0.40821073)
33
+ METACLIP_IMAGE_STD = (0.26862954, 0.26130258, 0.27577711)
34
+
35
+ _REQUIRED_CONFIG_KEYS = (
36
+ "clip_config",
37
+ "image_size",
38
+ "spatial_grid",
39
+ "proprio_dim",
40
+ "text_dim",
41
+ "phase_dim",
42
+ "hidden_dim",
43
+ "history",
44
+ "action_chunk",
45
+ "ensemble_heads",
46
+ "dropout",
47
+ "task_to_id",
48
+ "difficulty_to_id",
49
+ )
50
+
51
+
52
+ def _require_plain_int(value: Any, name: str, *, minimum: int = 1) -> int:
53
+ if isinstance(value, bool) or not isinstance(value, int) or value < minimum:
54
+ raise ValueError(f"{name} must be an integer >= {minimum}, got {value!r}")
55
+ return int(value)
56
+
57
+
58
+ def _require_probability(value: Any, name: str) -> float:
59
+ if isinstance(value, bool) or not isinstance(value, (int, float)):
60
+ raise ValueError(f"{name} must be a number in [0, 1), got {value!r}")
61
+ value = float(value)
62
+ if not math.isfinite(value) or not 0.0 <= value < 1.0:
63
+ raise ValueError(f"{name} must be a finite number in [0, 1), got {value!r}")
64
+ return value
65
+
66
+
67
+ def _validate_id_map(value: Any, name: str) -> dict[str, int]:
68
+ if not isinstance(value, Mapping) or not value:
69
+ raise ValueError(f"{name} must be a non-empty mapping of strings to IDs")
70
+ result: dict[str, int] = {}
71
+ for key, item in value.items():
72
+ if not isinstance(key, str) or not key.strip():
73
+ raise ValueError(f"{name} contains an invalid key: {key!r}")
74
+ if isinstance(item, bool) or not isinstance(item, int) or item < 0:
75
+ raise ValueError(f"{name}[{key!r}] must be a non-negative integer")
76
+ if key in result:
77
+ raise ValueError(f"{name} contains duplicate key {key!r}")
78
+ result[key] = int(item)
79
+ expected = list(range(len(result)))
80
+ actual = sorted(result.values())
81
+ if actual != expected:
82
+ raise ValueError(
83
+ f"{name} IDs must be unique and contiguous 0..{len(result) - 1}, got {actual}"
84
+ )
85
+ return result
86
+
87
+
88
+ def validate_policy_config(config: Mapping[str, Any]) -> dict[str, Any]:
89
+ """Validate and return an isolated copy of a policy configuration.
90
+
91
+ Deployment/training metadata outside the architectural keys is retained,
92
+ but all fields that affect tensor shapes or preprocessing are checked. The
93
+ returned object can therefore be safely stored on a model without later
94
+ mutations by the caller changing its behavior.
95
+ """
96
+
97
+ if not isinstance(config, Mapping):
98
+ raise TypeError(f"config must be a mapping, got {type(config).__name__}")
99
+ missing = [key for key in _REQUIRED_CONFIG_KEYS if key not in config]
100
+ if missing:
101
+ raise ValueError(f"policy config is missing required keys: {missing}")
102
+
103
+ validated = copy.deepcopy(dict(config))
104
+ clip_config = validated["clip_config"]
105
+ if not isinstance(clip_config, Mapping) or not clip_config:
106
+ raise ValueError("clip_config must be a non-empty mapping")
107
+ clip_config = copy.deepcopy(dict(clip_config))
108
+ if clip_config.get("model_type") != "clip":
109
+ raise ValueError(
110
+ "clip_config.model_type must be 'clip', got "
111
+ f"{clip_config.get('model_type')!r}"
112
+ )
113
+ if (
114
+ _require_plain_int(
115
+ clip_config.get("projection_dim"), "clip_config.projection_dim"
116
+ )
117
+ != DEFAULT_TEXT_DIM
118
+ ):
119
+ raise ValueError(
120
+ f"MetaCLIP projection_dim must be {DEFAULT_TEXT_DIM}, "
121
+ f"got {clip_config.get('projection_dim')!r}"
122
+ )
123
+
124
+ vision_config = clip_config.get("vision_config")
125
+ if not isinstance(vision_config, Mapping) or not vision_config:
126
+ raise ValueError("clip_config.vision_config must be a non-empty mapping")
127
+ vision_config = copy.deepcopy(dict(vision_config))
128
+ for key in (
129
+ "hidden_size",
130
+ "intermediate_size",
131
+ "image_size",
132
+ "patch_size",
133
+ "num_hidden_layers",
134
+ "num_attention_heads",
135
+ ):
136
+ if key not in vision_config:
137
+ raise ValueError(
138
+ f"clip_config.vision_config is missing required key {key!r}"
139
+ )
140
+ _require_plain_int(vision_config[key], f"clip_config.vision_config.{key}")
141
+ if vision_config.get("model_type") != "clip_vision_model":
142
+ raise ValueError(
143
+ "clip_config.vision_config.model_type must be 'clip_vision_model', got "
144
+ f"{vision_config.get('model_type')!r}"
145
+ )
146
+ if (
147
+ int(vision_config["image_size"]) != 224
148
+ or int(vision_config["patch_size"]) != 16
149
+ ):
150
+ raise ValueError(
151
+ "The audited MetaCLIP B/16 contract requires image_size=224 and "
152
+ f"patch_size=16, got {vision_config['image_size']!r} and "
153
+ f"{vision_config['patch_size']!r}"
154
+ )
155
+ if int(vision_config["hidden_size"]) % int(vision_config["num_attention_heads"]):
156
+ raise ValueError(
157
+ "vision_config.hidden_size must be divisible by "
158
+ "vision_config.num_attention_heads"
159
+ )
160
+ if "num_channels" in vision_config and int(vision_config["num_channels"]) != 3:
161
+ raise ValueError(
162
+ "only three-channel MetaCLIP vision configurations are supported"
163
+ )
164
+
165
+ text_config = clip_config.get("text_config")
166
+ if not isinstance(text_config, Mapping) or not text_config:
167
+ raise ValueError("clip_config.text_config must be a non-empty mapping")
168
+ text_config = copy.deepcopy(dict(text_config))
169
+ for key in (
170
+ "hidden_size",
171
+ "intermediate_size",
172
+ "max_position_embeddings",
173
+ "num_hidden_layers",
174
+ "num_attention_heads",
175
+ "vocab_size",
176
+ ):
177
+ if key not in text_config:
178
+ raise ValueError(f"clip_config.text_config is missing required key {key!r}")
179
+ _require_plain_int(text_config[key], f"clip_config.text_config.{key}")
180
+ if text_config.get("model_type") != "clip_text_model":
181
+ raise ValueError(
182
+ "clip_config.text_config.model_type must be 'clip_text_model', got "
183
+ f"{text_config.get('model_type')!r}"
184
+ )
185
+ if int(text_config["max_position_embeddings"]) != 77:
186
+ raise ValueError(
187
+ "MetaCLIP text max_position_embeddings must be 77, got "
188
+ f"{text_config['max_position_embeddings']!r}"
189
+ )
190
+ clip_config["vision_config"] = vision_config
191
+ clip_config["text_config"] = text_config
192
+ validated["clip_config"] = clip_config
193
+
194
+ image_size = _require_plain_int(validated["image_size"], "image_size")
195
+ if image_size != int(vision_config["image_size"]):
196
+ raise ValueError(
197
+ "policy image_size must match vision_config.image_size, got "
198
+ f"{image_size} and {vision_config['image_size']!r}"
199
+ )
200
+ patch_size = int(vision_config["patch_size"])
201
+ if image_size % patch_size:
202
+ raise ValueError(
203
+ f"image_size={image_size} must be divisible by MetaCLIP patch_size={patch_size}"
204
+ )
205
+ patch_side = image_size // patch_size
206
+ spatial_grid = _require_plain_int(validated["spatial_grid"], "spatial_grid")
207
+ if spatial_grid > patch_side:
208
+ raise ValueError(
209
+ f"spatial_grid={spatial_grid} cannot exceed the {patch_side}x{patch_side} "
210
+ "MetaCLIP input patch grid"
211
+ )
212
+
213
+ for key in (
214
+ "proprio_dim",
215
+ "text_dim",
216
+ "phase_dim",
217
+ "hidden_dim",
218
+ "history",
219
+ "action_chunk",
220
+ "ensemble_heads",
221
+ ):
222
+ validated[key] = _require_plain_int(validated[key], key)
223
+ if validated["text_dim"] != int(clip_config["projection_dim"]):
224
+ raise ValueError(
225
+ "text_dim must equal clip_config.projection_dim, got "
226
+ f"{validated['text_dim']} and {clip_config['projection_dim']!r}"
227
+ )
228
+ if validated["phase_dim"] != DEFAULT_PHASE_DIM:
229
+ raise ValueError(
230
+ f"phase_dim must be {DEFAULT_PHASE_DIM} for phase_vector(), "
231
+ f"got {validated['phase_dim']}"
232
+ )
233
+ if validated["hidden_dim"] % DEFAULT_SPATIAL_HEADS:
234
+ raise ValueError(
235
+ f"hidden_dim must be divisible by {DEFAULT_SPATIAL_HEADS} spatial heads"
236
+ )
237
+ validated["dropout"] = _require_probability(validated["dropout"], "dropout")
238
+ validated["task_to_id"] = _validate_id_map(validated["task_to_id"], "task_to_id")
239
+ validated["difficulty_to_id"] = _validate_id_map(
240
+ validated["difficulty_to_id"], "difficulty_to_id"
241
+ )
242
+
243
+ if "action_dim" in validated and validated["action_dim"] != ACTION_DIM:
244
+ raise ValueError(
245
+ f"action_dim must be {ACTION_DIM}, got {validated['action_dim']!r}"
246
+ )
247
+ if "rgb_dim" in validated and validated["rgb_dim"] != 3:
248
+ raise ValueError(f"rgb_dim must be 3, got {validated['rgb_dim']!r}")
249
+ if (
250
+ "text_feature_version" in validated
251
+ and validated["text_feature_version"] != TEXT_FEATURE_VERSION
252
+ ):
253
+ raise ValueError(
254
+ f"text_feature_version must be {TEXT_FEATURE_VERSION!r}, got "
255
+ f"{validated['text_feature_version']!r}"
256
+ )
257
+ return validated
258
+
259
+
260
+ def phase_vector(step: int, horizon: int) -> "np.ndarray":
261
+ """Encode episode progress using the exact train/runtime four-vector."""
262
+
263
+ if isinstance(step, bool) or not isinstance(step, (int, np.integer)):
264
+ raise ValueError(f"step must be an integer, got {step!r}")
265
+ if isinstance(horizon, bool) or not isinstance(horizon, (int, np.integer)):
266
+ raise ValueError(f"horizon must be an integer, got {horizon!r}")
267
+ horizon_f = max(float(horizon), 1.0)
268
+ progress = float(np.clip(float(step) / horizon_f, 0.0, 1.0))
269
+ return np.asarray(
270
+ [
271
+ progress,
272
+ 1.0 - progress,
273
+ math.sin(math.pi * progress),
274
+ math.cos(math.pi * progress),
275
+ ],
276
+ dtype=np.float32,
277
+ )
278
+
279
+
280
+ def _images_to_nchw_rgb(images: torch.Tensor) -> torch.Tensor:
281
+ if not isinstance(images, torch.Tensor):
282
+ raise TypeError(f"images must be a torch.Tensor, got {type(images).__name__}")
283
+ if images.ndim != 4:
284
+ raise ValueError(
285
+ f"expected a four-dimensional image tensor, got {tuple(images.shape)}"
286
+ )
287
+
288
+ # Prefer an unambiguous channel-first interpretation, then NHWC. Normal
289
+ # robotics images are 224x224, so both layouts are unambiguous in practice.
290
+ if images.shape[1] in (1, 3, 4) and images.shape[-1] not in (1, 3, 4):
291
+ nchw = images
292
+ elif images.shape[-1] in (1, 3, 4):
293
+ nchw = images.permute(0, 3, 1, 2)
294
+ elif images.shape[1] in (1, 3, 4):
295
+ nchw = images
296
+ else:
297
+ raise ValueError(
298
+ f"cannot determine image channels for shape {tuple(images.shape)}"
299
+ )
300
+
301
+ if nchw.shape[1] == 1:
302
+ nchw = nchw.repeat(1, 3, 1, 1)
303
+ elif nchw.shape[1] == 4:
304
+ nchw = nchw[:, :3]
305
+ if nchw.shape[1] != 3:
306
+ raise ValueError(
307
+ f"expected one, three, or four image channels, got {nchw.shape[1]}"
308
+ )
309
+ return nchw
310
+
311
+
312
+ def images_to_unit_rgb(images: torch.Tensor, image_size: int = 224) -> torch.Tensor:
313
+ """Convert NHWC/NCHW uint8-like images to resized NCHW RGB in ``[0, 1]``."""
314
+
315
+ image_size = _require_plain_int(image_size, "image_size")
316
+ rgb = _images_to_nchw_rgb(images).float()
317
+ if rgb.numel() and float(rgb.detach().amax().cpu()) > 2.0:
318
+ rgb = rgb / 255.0
319
+ if not bool(torch.isfinite(rgb).all().detach().cpu()):
320
+ raise ValueError("images contain NaN or Inf")
321
+ if tuple(rgb.shape[-2:]) != (image_size, image_size):
322
+ rgb = F.interpolate(
323
+ rgb,
324
+ size=(image_size, image_size),
325
+ mode="bicubic",
326
+ align_corners=False,
327
+ antialias=True,
328
+ )
329
+ return rgb
330
+
331
+
332
+ def normalize_images(images: torch.Tensor, image_size: int = 224) -> torch.Tensor:
333
+ """Prepare image pixels for the MetaCLIP vision encoder."""
334
+
335
+ rgb = images_to_unit_rgb(images, image_size=image_size)
336
+ mean = rgb.new_tensor(METACLIP_IMAGE_MEAN).view(1, 3, 1, 1)
337
+ std = rgb.new_tensor(METACLIP_IMAGE_STD).view(1, 3, 1, 1)
338
+ return (rgb - mean) / std
339
+
340
+
341
+ def rgb_grid_tokens(
342
+ images: torch.Tensor, spatial_grid: int, image_size: int = 224
343
+ ) -> torch.Tensor:
344
+ """Return global RGB plus a row-major spatial grid, shaped ``[B,1+G²,3]``."""
345
+
346
+ spatial_grid = _require_plain_int(spatial_grid, "spatial_grid")
347
+ rgb = images_to_unit_rgb(images, image_size=image_size)
348
+ global_rgb = rgb.mean(dim=(-2, -1)).unsqueeze(1)
349
+ grid_rgb = F.adaptive_avg_pool2d(rgb, (spatial_grid, spatial_grid))
350
+ grid_rgb = grid_rgb.flatten(2).transpose(1, 2)
351
+ return torch.cat([global_rgb, grid_rgb], dim=1)
352
+
353
+
354
+ def pool_metaclip_tokens(
355
+ hidden_states: torch.Tensor,
356
+ spatial_grid: int,
357
+ ) -> torch.Tensor:
358
+ """Pool MetaCLIP patch tokens to ``G x G`` and retain the CLS token."""
359
+
360
+ spatial_grid = _require_plain_int(spatial_grid, "spatial_grid")
361
+ if not isinstance(hidden_states, torch.Tensor):
362
+ raise TypeError("hidden_states must be a torch.Tensor")
363
+ if hidden_states.ndim != 3 or hidden_states.shape[1] <= 1:
364
+ raise ValueError(
365
+ f"unexpected MetaCLIP output shape: {tuple(hidden_states.shape)}"
366
+ )
367
+ cls_token = hidden_states[:, :1]
368
+ patches = hidden_states[:, 1:]
369
+ side = math.isqrt(int(patches.shape[1]))
370
+ if side * side != int(patches.shape[1]):
371
+ raise ValueError(f"MetaCLIP patch count {patches.shape[1]} is not a square")
372
+ patches = patches.transpose(1, 2).reshape(
373
+ patches.shape[0], patches.shape[2], side, side
374
+ )
375
+ patches = F.adaptive_avg_pool2d(patches, (spatial_grid, spatial_grid))
376
+ patches = patches.flatten(2).transpose(1, 2)
377
+ return torch.cat([cls_token, patches], dim=1)
378
+
379
+
380
+ class MetaCLIPActionChunkHead(nn.Module):
381
+ """Task-conditioned spatial pooling followed by a short temporal policy."""
382
+
383
+ def __init__(
384
+ self,
385
+ *,
386
+ vision_dim: int,
387
+ proprio_dim: int,
388
+ text_dim: int,
389
+ phase_dim: int,
390
+ num_tasks: int,
391
+ num_difficulties: int,
392
+ hidden_dim: int = 256,
393
+ history: int = 4,
394
+ action_chunk: int = 8,
395
+ ensemble_heads: int = 3,
396
+ dropout: float = 0.10,
397
+ ) -> None:
398
+ super().__init__()
399
+ self.vision_dim = _require_plain_int(vision_dim, "vision_dim")
400
+ self.proprio_dim = _require_plain_int(proprio_dim, "proprio_dim")
401
+ self.text_dim = _require_plain_int(text_dim, "text_dim")
402
+ self.phase_dim = _require_plain_int(phase_dim, "phase_dim")
403
+ self.hidden_dim = _require_plain_int(hidden_dim, "hidden_dim")
404
+ self.history = _require_plain_int(history, "history")
405
+ self.action_chunk = _require_plain_int(action_chunk, "action_chunk")
406
+ self.ensemble_heads = _require_plain_int(ensemble_heads, "ensemble_heads")
407
+ num_tasks = _require_plain_int(num_tasks, "num_tasks")
408
+ num_difficulties = _require_plain_int(num_difficulties, "num_difficulties")
409
+ dropout = _require_probability(dropout, "dropout")
410
+ if self.hidden_dim % DEFAULT_SPATIAL_HEADS:
411
+ raise ValueError(
412
+ f"hidden_dim must be divisible by {DEFAULT_SPATIAL_HEADS} spatial heads"
413
+ )
414
+
415
+ self.vision_proj = nn.Sequential(
416
+ nn.LayerNorm(self.vision_dim), nn.Linear(self.vision_dim, self.hidden_dim)
417
+ )
418
+ self.rgb_proj = nn.Sequential(nn.Linear(3, self.hidden_dim), nn.SiLU())
419
+ self.proprio_proj = nn.Sequential(
420
+ nn.LayerNorm(self.proprio_dim + self.phase_dim),
421
+ nn.Linear(self.proprio_dim + self.phase_dim, self.hidden_dim),
422
+ nn.SiLU(),
423
+ nn.Dropout(dropout),
424
+ )
425
+ self.text_proj = nn.Sequential(
426
+ nn.LayerNorm(self.text_dim),
427
+ nn.Linear(self.text_dim, self.hidden_dim),
428
+ nn.SiLU(),
429
+ )
430
+ # The last row of each embedding is the trained unknown/fallback ID.
431
+ self.task_embedding = nn.Embedding(num_tasks + 1, self.hidden_dim)
432
+ self.difficulty_embedding = nn.Embedding(num_difficulties + 1, self.hidden_dim)
433
+ self.task_scale = nn.Parameter(torch.tensor(0.5))
434
+ self.difficulty_scale = nn.Parameter(torch.tensor(0.25))
435
+ self.condition_norm = nn.LayerNorm(self.hidden_dim)
436
+ self.spatial_attention = nn.MultiheadAttention(
437
+ embed_dim=self.hidden_dim,
438
+ num_heads=DEFAULT_SPATIAL_HEADS,
439
+ dropout=dropout,
440
+ batch_first=True,
441
+ )
442
+ self.frame_fusion = nn.Sequential(
443
+ nn.Linear(self.hidden_dim * 2, self.hidden_dim),
444
+ nn.SiLU(),
445
+ nn.LayerNorm(self.hidden_dim),
446
+ nn.Dropout(dropout),
447
+ )
448
+ self.temporal_gru = nn.GRU(
449
+ input_size=self.hidden_dim,
450
+ hidden_size=self.hidden_dim,
451
+ num_layers=2,
452
+ dropout=dropout,
453
+ batch_first=True,
454
+ )
455
+ self.output_heads = nn.ModuleList(
456
+ [
457
+ nn.Sequential(
458
+ nn.LayerNorm(self.hidden_dim),
459
+ nn.Linear(self.hidden_dim, self.hidden_dim),
460
+ nn.SiLU(),
461
+ nn.Dropout(dropout),
462
+ nn.Linear(self.hidden_dim, self.action_chunk * ACTION_DIM),
463
+ )
464
+ for _ in range(self.ensemble_heads)
465
+ ]
466
+ )
467
+
468
+ def _validate_inputs(
469
+ self,
470
+ visual_tokens: torch.Tensor,
471
+ rgb_tokens: torch.Tensor,
472
+ proprio: torch.Tensor,
473
+ phase: torch.Tensor,
474
+ text_features: torch.Tensor,
475
+ task_ids: torch.Tensor,
476
+ difficulty_ids: torch.Tensor,
477
+ ) -> tuple[int, int, int]:
478
+ if visual_tokens.ndim != 4:
479
+ raise ValueError(
480
+ f"visual_tokens must have shape [B,T,V,D], got {tuple(visual_tokens.shape)}"
481
+ )
482
+ batch, timesteps, token_count, vision_dim = visual_tokens.shape
483
+ if timesteps != self.history:
484
+ raise ValueError(f"expected history={self.history}, got {timesteps}")
485
+ if vision_dim != self.vision_dim:
486
+ raise ValueError(f"expected vision_dim={self.vision_dim}, got {vision_dim}")
487
+ if rgb_tokens.shape != (batch, timesteps, token_count, 3):
488
+ raise ValueError(
489
+ "rgb_tokens must align with visual_tokens and end in RGB, got "
490
+ f"{tuple(rgb_tokens.shape)}"
491
+ )
492
+ if proprio.shape != (batch, timesteps, self.proprio_dim):
493
+ raise ValueError(
494
+ f"proprio must have shape {(batch, timesteps, self.proprio_dim)}, "
495
+ f"got {tuple(proprio.shape)}"
496
+ )
497
+ if phase.shape != (batch, timesteps, self.phase_dim):
498
+ raise ValueError(
499
+ f"phase must have shape {(batch, timesteps, self.phase_dim)}, "
500
+ f"got {tuple(phase.shape)}"
501
+ )
502
+ if text_features.shape != (batch, self.text_dim):
503
+ raise ValueError(
504
+ f"text_features must have shape {(batch, self.text_dim)}, "
505
+ f"got {tuple(text_features.shape)}"
506
+ )
507
+ for name, ids in (("task_ids", task_ids), ("difficulty_ids", difficulty_ids)):
508
+ if ids.shape != (batch,):
509
+ raise ValueError(
510
+ f"{name} must have shape {(batch,)}, got {tuple(ids.shape)}"
511
+ )
512
+ if ids.dtype not in (torch.int32, torch.int64):
513
+ raise ValueError(
514
+ f"{name} must contain integer IDs, got dtype={ids.dtype}"
515
+ )
516
+ return batch, timesteps, token_count
517
+
518
+ def forward_cached(
519
+ self,
520
+ visual_tokens: torch.Tensor,
521
+ rgb_tokens: torch.Tensor,
522
+ proprio: torch.Tensor,
523
+ phase: torch.Tensor,
524
+ text_features: torch.Tensor,
525
+ task_ids: torch.Tensor,
526
+ difficulty_ids: torch.Tensor,
527
+ ) -> torch.Tensor:
528
+ """Return raw action logits shaped ``[ensemble, batch, chunk, 7]``."""
529
+
530
+ batch, timesteps, token_count = self._validate_inputs(
531
+ visual_tokens,
532
+ rgb_tokens,
533
+ proprio,
534
+ phase,
535
+ text_features,
536
+ task_ids,
537
+ difficulty_ids,
538
+ )
539
+ visual = self.vision_proj(visual_tokens) + self.rgb_proj(rgb_tokens)
540
+ state = self.proprio_proj(torch.cat([proprio, phase], dim=-1))
541
+ text = self.text_proj(text_features)
542
+ task = self.task_embedding(task_ids)
543
+ difficulty = self.difficulty_embedding(difficulty_ids)
544
+ condition = self.condition_norm(
545
+ state
546
+ + text[:, None]
547
+ + torch.tanh(self.task_scale) * task[:, None]
548
+ + torch.tanh(self.difficulty_scale) * difficulty[:, None]
549
+ )
550
+
551
+ flat_visual = visual.reshape(batch * timesteps, token_count, self.hidden_dim)
552
+ flat_query = condition.reshape(batch * timesteps, 1, self.hidden_dim)
553
+ attended, _ = self.spatial_attention(
554
+ flat_query, flat_visual, flat_visual, need_weights=False
555
+ )
556
+ attended = attended.reshape(batch, timesteps, self.hidden_dim)
557
+ frames = self.frame_fusion(torch.cat([attended, condition], dim=-1))
558
+ temporal, _ = self.temporal_gru(frames)
559
+ final = temporal[:, -1]
560
+ outputs = [
561
+ head(final).reshape(batch, self.action_chunk, ACTION_DIM)
562
+ for head in self.output_heads
563
+ ]
564
+ return torch.stack(outputs, dim=0)
565
+
566
+ def forward(
567
+ self,
568
+ visual_tokens: torch.Tensor,
569
+ rgb_tokens: torch.Tensor,
570
+ proprio: torch.Tensor,
571
+ phase: torch.Tensor,
572
+ text_features: torch.Tensor,
573
+ task_ids: torch.Tensor,
574
+ difficulty_ids: torch.Tensor,
575
+ ) -> torch.Tensor:
576
+ return self.forward_cached(
577
+ visual_tokens,
578
+ rgb_tokens,
579
+ proprio,
580
+ phase,
581
+ text_features,
582
+ task_ids,
583
+ difficulty_ids,
584
+ )
585
+
586
+
587
+ class MetaCLIPActionChunkModel(nn.Module):
588
+ """Complete submission model containing frozen MetaCLIP and the policy head."""
589
+
590
+ def __init__(self, config: Mapping[str, Any]) -> None:
591
+ super().__init__()
592
+ self.policy_config = validate_policy_config(config)
593
+ self.spatial_grid = int(self.policy_config["spatial_grid"])
594
+
595
+ # Offline construction only: this creates a model from the embedded
596
+ # architecture. It never resolves a repository or downloads weights.
597
+ clip_config = CLIPConfig.from_dict(self.policy_config["clip_config"])
598
+ self.clip = CLIPModel(clip_config)
599
+ self.head = MetaCLIPActionChunkHead(
600
+ vision_dim=int(clip_config.vision_config.hidden_size),
601
+ proprio_dim=int(self.policy_config["proprio_dim"]),
602
+ text_dim=int(self.policy_config["text_dim"]),
603
+ phase_dim=int(self.policy_config["phase_dim"]),
604
+ num_tasks=len(self.policy_config["task_to_id"]),
605
+ num_difficulties=len(self.policy_config["difficulty_to_id"]),
606
+ hidden_dim=int(self.policy_config["hidden_dim"]),
607
+ history=int(self.policy_config["history"]),
608
+ action_chunk=int(self.policy_config["action_chunk"]),
609
+ ensemble_heads=int(self.policy_config["ensemble_heads"]),
610
+ dropout=float(self.policy_config["dropout"]),
611
+ )
612
+
613
+ def freeze_backbone(self) -> None:
614
+ """Freeze both MetaCLIP towers and keep them in inference mode."""
615
+
616
+ self.clip.requires_grad_(False)
617
+ self.clip.eval()
618
+
619
+ def train(self, mode: bool = True) -> "MetaCLIPActionChunkModel":
620
+ # A caller may train the complete wrapper for convenience. If the
621
+ # backbone has been frozen, do not accidentally switch it back to train
622
+ # mode through nn.Module.train() recursion.
623
+ super().train(mode)
624
+ if not any(parameter.requires_grad for parameter in self.clip.parameters()):
625
+ self.clip.eval()
626
+ return self
627
+
628
+ def encode_images(self, images: torch.Tensor) -> tuple[torch.Tensor, torch.Tensor]:
629
+ """Encode an image batch into aligned MetaCLIP and raw-RGB tokens."""
630
+
631
+ image_size = int(self.policy_config["image_size"])
632
+ rgb_tokens = rgb_grid_tokens(
633
+ images, spatial_grid=self.spatial_grid, image_size=image_size
634
+ )
635
+ pixels = normalize_images(images, image_size=image_size)
636
+ vision_parameter = next(self.clip.vision_model.parameters())
637
+ pixels = pixels.to(device=vision_parameter.device, dtype=vision_parameter.dtype)
638
+ rgb_tokens = rgb_tokens.to(
639
+ device=vision_parameter.device, dtype=vision_parameter.dtype
640
+ )
641
+ hidden = self.clip.vision_model(pixel_values=pixels).last_hidden_state
642
+ hidden = self.clip.vision_model.post_layernorm(hidden)
643
+ return (
644
+ pool_metaclip_tokens(hidden, self.spatial_grid),
645
+ rgb_tokens,
646
+ )
647
+
648
+ def encode_text(
649
+ self, input_ids: torch.Tensor, attention_mask: torch.Tensor
650
+ ) -> torch.Tensor:
651
+ """Return normalized projected MetaCLIP instruction embeddings."""
652
+
653
+ parameter = next(self.clip.text_model.parameters())
654
+ input_ids = input_ids.to(device=parameter.device)
655
+ attention_mask = attention_mask.to(device=parameter.device)
656
+ outputs = self.clip.text_model(
657
+ input_ids=input_ids,
658
+ attention_mask=attention_mask,
659
+ return_dict=True,
660
+ )
661
+ features = self.clip.text_projection(outputs.pooler_output)
662
+ return F.normalize(features.float(), dim=-1).to(dtype=parameter.dtype)
663
+
664
+ def forward_cached(
665
+ self,
666
+ visual_tokens: torch.Tensor,
667
+ rgb_tokens: torch.Tensor,
668
+ proprio: torch.Tensor,
669
+ phase: torch.Tensor,
670
+ text_features: torch.Tensor,
671
+ task_ids: torch.Tensor,
672
+ difficulty_ids: torch.Tensor,
673
+ ) -> torch.Tensor:
674
+ return self.head.forward_cached(
675
+ visual_tokens,
676
+ rgb_tokens,
677
+ proprio,
678
+ phase,
679
+ text_features,
680
+ task_ids,
681
+ difficulty_ids,
682
+ )
683
+
684
+ def forward(
685
+ self,
686
+ visual_tokens: torch.Tensor,
687
+ rgb_tokens: torch.Tensor,
688
+ proprio: torch.Tensor,
689
+ phase: torch.Tensor,
690
+ text_features: torch.Tensor,
691
+ task_ids: torch.Tensor,
692
+ difficulty_ids: torch.Tensor,
693
+ ) -> torch.Tensor:
694
+ return self.forward_cached(
695
+ visual_tokens,
696
+ rgb_tokens,
697
+ proprio,
698
+ phase,
699
+ text_features,
700
+ task_ids,
701
+ difficulty_ids,
702
+ )
703
+
704
+
705
+ # Compatibility aliases make reference checkpoints/code easy to compare while
706
+ # retaining descriptive names in the new trainer.
707
+ CompetitivePolicyHead = MetaCLIPActionChunkHead
708
+ CompetitiveVLAModel = MetaCLIPActionChunkModel
709
+
710
+
711
+ __all__ = [
712
+ "ACTION_DIM",
713
+ "DEFAULT_DIFFICULTIES",
714
+ "DEFAULT_PHASE_DIM",
715
+ "DEFAULT_SPATIAL_HEADS",
716
+ "DEFAULT_TEXT_DIM",
717
+ "METACLIP_IMAGE_MEAN",
718
+ "METACLIP_IMAGE_STD",
719
+ "TEXT_FEATURE_VERSION",
720
+ "CompetitivePolicyHead",
721
+ "CompetitiveVLAModel",
722
+ "MetaCLIPActionChunkHead",
723
+ "MetaCLIPActionChunkModel",
724
+ "images_to_unit_rgb",
725
+ "normalize_images",
726
+ "phase_vector",
727
+ "pool_metaclip_tokens",
728
+ "rgb_grid_tokens",
729
+ "validate_policy_config",
730
+ ]
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:b833cd1cee8d6070d162883e211cb13cdb455ed74633b5b06de199b2035be5ef
3
- size 1219514848
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c8cda3c02b88b7b1a1e164e66bf4168a25a5b022bf3466f791108c1185974217
3
+ size 605614380
vla_config.json CHANGED
@@ -10,8 +10,104 @@
10
  "difficulty",
11
  "horizon"
12
  ],
13
- "backbone_id_for_provenance": "facebook/dinov3-vitl16-pretrain-lvd1689m",
14
- "difficulty_condition_dropout": 0.45,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
15
  "difficulty_to_id": {
16
  "low": 0,
17
  "medium": 1
@@ -29,17 +125,17 @@
29
  "history": 4,
30
  "image_size": 224,
31
  "license": {
32
- "file": "LICENSE.md",
33
- "id": "dinov3-license",
34
- "source_sha256": "25d122eb8f5b880fd23c736fb6ea8018ee45c12237e00b8a86d14c653904999e"
35
  },
36
- "model_type": "frozen_dinov3_rgb_temporal_action_chunk_bc",
37
  "phase_dim": 4,
38
  "proprio_dim": 25,
39
  "rgb_dim": 3,
40
- "schema_version": "flock_robotics_dinov3_action_chunk_v1",
41
  "spatial_grid": 8,
42
- "task_condition_dropout": 0.3,
43
  "task_to_id": {
44
  "lift_cube": 0,
45
  "pick_place_bread": 1,
@@ -49,89 +145,13 @@
49
  "stack_blocks": 5
50
  },
51
  "temporal_ensemble_decay": 0.55,
52
- "text_dim": 128,
53
- "text_feature_version": "signed_hash_subtokens_v2",
54
- "unknown_condition_strategy": "trained_fallback_embedding",
55
- "vision_config": {
56
- "_name_or_path": "facebook/dinov3-vitl16-pretrain-lvd1689m",
57
- "apply_layernorm": true,
58
- "architectures": [
59
- "DINOv3ViTModel"
60
- ],
61
- "attention_dropout": 0.0,
62
- "chunk_size_feed_forward": 0,
63
- "drop_path_rate": 0.0,
64
- "dtype": "float32",
65
- "hidden_act": "gelu",
66
- "hidden_size": 1024,
67
- "id2label": {
68
- "0": "LABEL_0",
69
- "1": "LABEL_1"
70
- },
71
- "image_size": 224,
72
- "initializer_range": 0.02,
73
- "intermediate_size": 4096,
74
- "is_encoder_decoder": false,
75
- "key_bias": false,
76
- "label2id": {
77
- "LABEL_0": 0,
78
- "LABEL_1": 1
79
- },
80
- "layer_norm_eps": 1e-05,
81
- "layerscale_value": 1.0,
82
- "mlp_bias": true,
83
- "model_type": "dinov3_vit",
84
- "num_attention_heads": 16,
85
- "num_channels": 3,
86
- "num_hidden_layers": 24,
87
- "num_register_tokens": 4,
88
- "out_features": [
89
- "stage24"
90
- ],
91
- "out_indices": [
92
- 24
93
- ],
94
- "output_attentions": false,
95
- "output_hidden_states": false,
96
- "patch_size": 16,
97
- "pos_embed_jitter": null,
98
- "pos_embed_rescale": 2.0,
99
- "pos_embed_shift": null,
100
- "problem_type": null,
101
- "proj_bias": true,
102
- "query_bias": true,
103
- "reshape_hidden_states": true,
104
- "return_dict": true,
105
- "rope_theta": 100.0,
106
- "stage_names": [
107
- "stem",
108
- "stage1",
109
- "stage2",
110
- "stage3",
111
- "stage4",
112
- "stage5",
113
- "stage6",
114
- "stage7",
115
- "stage8",
116
- "stage9",
117
- "stage10",
118
- "stage11",
119
- "stage12",
120
- "stage13",
121
- "stage14",
122
- "stage15",
123
- "stage16",
124
- "stage17",
125
- "stage18",
126
- "stage19",
127
- "stage20",
128
- "stage21",
129
- "stage22",
130
- "stage23",
131
- "stage24"
132
- ],
133
- "transformers_version": "5.4.0",
134
- "use_gated_mlp": false,
135
- "value_bias": true
136
- }
137
  }
 
10
  "difficulty",
11
  "horizon"
12
  ],
13
+ "backbone_id_for_provenance": "facebook/metaclip-b16-fullcc2.5b",
14
+ "clip_config": {
15
+ "_name_or_path": "facebook/metaclip-b16-fullcc2.5b",
16
+ "architectures": [
17
+ "CLIPModel"
18
+ ],
19
+ "chunk_size_feed_forward": 0,
20
+ "dtype": "float32",
21
+ "id2label": {
22
+ "0": "LABEL_0",
23
+ "1": "LABEL_1"
24
+ },
25
+ "initializer_factor": 1.0,
26
+ "is_encoder_decoder": false,
27
+ "label2id": {
28
+ "LABEL_0": 0,
29
+ "LABEL_1": 1
30
+ },
31
+ "logit_scale_init_value": 2.6592,
32
+ "model_type": "clip",
33
+ "output_attentions": false,
34
+ "output_hidden_states": false,
35
+ "problem_type": null,
36
+ "projection_dim": 512,
37
+ "return_dict": true,
38
+ "text_config": {
39
+ "_name_or_path": "",
40
+ "architectures": null,
41
+ "attention_dropout": 0.0,
42
+ "bos_token_id": 49406,
43
+ "chunk_size_feed_forward": 0,
44
+ "dtype": "float32",
45
+ "eos_token_id": 49407,
46
+ "heads": 8,
47
+ "hidden_act": "quick_gelu",
48
+ "hidden_size": 512,
49
+ "id2label": {
50
+ "0": "LABEL_0",
51
+ "1": "LABEL_1"
52
+ },
53
+ "initializer_factor": 1.0,
54
+ "initializer_range": 0.02,
55
+ "intermediate_size": 2048,
56
+ "is_encoder_decoder": false,
57
+ "label2id": {
58
+ "LABEL_0": 0,
59
+ "LABEL_1": 1
60
+ },
61
+ "layer_norm_eps": 1e-05,
62
+ "layers": 12,
63
+ "max_position_embeddings": 77,
64
+ "model_type": "clip_text_model",
65
+ "num_attention_heads": 8,
66
+ "num_hidden_layers": 12,
67
+ "output_attentions": false,
68
+ "output_hidden_states": false,
69
+ "pad_token_id": 1,
70
+ "problem_type": null,
71
+ "projection_dim": 512,
72
+ "return_dict": true,
73
+ "vocab_size": 49408
74
+ },
75
+ "transformers_version": "5.4.0",
76
+ "vision_config": {
77
+ "_name_or_path": "",
78
+ "architectures": null,
79
+ "attention_dropout": 0.0,
80
+ "chunk_size_feed_forward": 0,
81
+ "dtype": "float32",
82
+ "hidden_act": "quick_gelu",
83
+ "hidden_size": 768,
84
+ "id2label": {
85
+ "0": "LABEL_0",
86
+ "1": "LABEL_1"
87
+ },
88
+ "image_size": 224,
89
+ "initializer_factor": 1.0,
90
+ "initializer_range": 0.02,
91
+ "intermediate_size": 3072,
92
+ "is_encoder_decoder": false,
93
+ "label2id": {
94
+ "LABEL_0": 0,
95
+ "LABEL_1": 1
96
+ },
97
+ "layer_norm_eps": 1e-05,
98
+ "model_type": "clip_vision_model",
99
+ "num_attention_heads": 12,
100
+ "num_channels": 3,
101
+ "num_hidden_layers": 12,
102
+ "output_attentions": false,
103
+ "output_hidden_states": false,
104
+ "patch_size": 16,
105
+ "problem_type": null,
106
+ "projection_dim": 512,
107
+ "return_dict": true
108
+ }
109
+ },
110
+ "difficulty_condition_dropout": 0.5,
111
  "difficulty_to_id": {
112
  "low": 0,
113
  "medium": 1
 
125
  "history": 4,
126
  "image_size": 224,
127
  "license": {
128
+ "file": "CC-BY-NC-4.0.txt",
129
+ "id": "cc-by-nc-4.0",
130
+ "source_sha256": "41003d4a74749c0220e33dd415042164b5a1093ed401f36277234f772d22d3d0"
131
  },
132
+ "model_type": "frozen_metaclip_vision_text_temporal_action_chunk_bc",
133
  "phase_dim": 4,
134
  "proprio_dim": 25,
135
  "rgb_dim": 3,
136
+ "schema_version": "flock_robotics_metaclip_action_chunk_v2",
137
  "spatial_grid": 8,
138
+ "task_condition_dropout": 0.5,
139
  "task_to_id": {
140
  "lift_cube": 0,
141
  "pick_place_bread": 1,
 
145
  "stack_blocks": 5
146
  },
147
  "temporal_ensemble_decay": 0.55,
148
+ "text_dim": 512,
149
+ "text_feature_version": "metaclip_clip_bpe_projected_l2_text_v2",
150
+ "tokenizer": {
151
+ "merges_file": "merges.txt",
152
+ "model_max_length": 77,
153
+ "type": "CLIPTokenizer",
154
+ "vocab_file": "vocab.json"
155
+ },
156
+ "unknown_condition_strategy": "trained_fallback_embedding"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
157
  }
vocab.json ADDED
The diff for this file is too large to render. See raw diff