Update model card
Browse files
README.md
ADDED
|
@@ -0,0 +1,135 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
library_name: jepa.cpp
|
| 4 |
+
tags:
|
| 5 |
+
- jepa
|
| 6 |
+
- ggml
|
| 7 |
+
- gguf
|
| 8 |
+
- jepa.cpp
|
| 9 |
+
- v-jepa
|
| 10 |
+
- v-jepa-2
|
| 11 |
+
- world-model
|
| 12 |
+
- robotics
|
| 13 |
+
- video
|
| 14 |
+
- planning
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
# V-JEPA 2-AC ViT-g (action-conditioned world model) β GGUF for jepa.cpp
|
| 18 |
+
|
| 19 |
+
Meta's **V-JEPA 2-AC** β the action-conditioned world model behind the paper's zero-shot robot
|
| 20 |
+
planning β converted to GGUF for [jepa.cpp](https://github.com/aselimc/jepa.cpp), a ggml C/C++ engine that runs it on a plain CPU
|
| 21 |
+
with no Python and no PyTorch. One bundle: the frozen ViT-g/16 encoder from `vjepa2-ac-vitg.pt` plus
|
| 22 |
+
the 24-layer, 1024-dim predictor that takes a 7-d end-effector action and a 7-d pose per frame and
|
| 23 |
+
predicts the next frame's latents, block-causally over frames. `jepa_ac_rollout` scores K candidate
|
| 24 |
+
action sequences in one graph per step, and `jepa_ac_energy` is the L1 planning energy a CEM planner
|
| 25 |
+
minimises.
|
| 26 |
+
|
| 27 |
+
**Note on the encoder:** it is *not* `facebook/vjepa2-vitg-fpc64-256`. Meta's AC checkpoint carries its
|
| 28 |
+
own frozen ViT-g, which agrees with the HF release only to cosine ~0.998 per tensor; `encoder` and
|
| 29 |
+
`target_encoder` inside it are bit-identical. This bundle ships the checkpoint's own, which is what
|
| 30 |
+
`vjepa2_ac_vit_giant` loads.
|
| 31 |
+
|
| 32 |
+
**1317 M** parameters; D = 1408, 40 layers, 22 heads, patch 16, tubelet 2, 256x256. Everything the engine needs β dimensions, positional scheme,
|
| 33 |
+
preprocessing recipe, and class labels where there are any β travels inside the file, so inference needs
|
| 34 |
+
one binary and one GGUF and nothing else.
|
| 35 |
+
|
| 36 |
+
## Run it
|
| 37 |
+
|
| 38 |
+
```bash
|
| 39 |
+
git clone --recursive https://github.com/aselimc/jepa.cpp && cd jepa.cpp
|
| 40 |
+
cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release && cmake --build build -j
|
| 41 |
+
hf download jepacpp/vjepa2-ac-vitg-GGUF vjepa2-ac-vitg-f16.gguf --local-dir models/gguf
|
| 42 |
+
|
| 43 |
+
# encode a frame, roll 4 candidate action sequences out 2 steps, score them against a goal
|
| 44 |
+
build/jepa-worldmodel --ac -m vjepa2-ac-vitg-f16.gguf --image ctx.png --goal goal.png \
|
| 45 |
+
--actions-npy actions.npy # float32 [K, H, 7]
|
| 46 |
+
```
|
| 47 |
+
|
| 48 |
+
`--pool` selects `mean`, `cls`, `lewm` or `none` (the full token map); `-o` writes a `.npy`.
|
| 49 |
+
`scripts/download_models.sh` fetches whole sets at once. The C API is one header,
|
| 50 |
+
[`include/jepa.h`](https://github.com/aselimc/jepa.cpp/blob/main/include/jepa.h) β full reference on the [C API page](https://aselimc.github.io/jepa.cpp/api/).
|
| 51 |
+
|
| 52 |
+
## Files
|
| 53 |
+
|
| 54 |
+
| file | size | sha256 (first 16) | tier | measured against the PyTorch reference |
|
| 55 |
+
|---|---|---|---|---|
|
| 56 |
+
| `vjepa2-ac-vitg-f32.gguf` | 5025.5 MiB | `234b7e775da2ff28` | exact | β |
|
| 57 |
+
| `vjepa2-ac-vitg-f16.gguf` | 2519.0 MiB | `d57cc1760f90f779` | parity | β |
|
| 58 |
+
| `vjepa2-ac-vitg-q8_0.gguf` | 1344.1 MiB | `3dffe7130d7f7efd` | parity, **below the bar** | β |
|
| 59 |
+
| `vjepa2-ac-vitg-q4_0.gguf` | 717.4 MiB | `554f3a9f293a2b80` | advisory, **below the bar** | β |
|
| 60 |
+
| `vjepa2-ac-vitg-q4_k.gguf` | 717.4 MiB | `0d2fb4e022e2d1ca` | advisory, **below the bar** | β |
|
| 61 |
+
|
| 62 |
+
<sub>α΅ `tests/test-parity` on the CPU backend, stored reference input, 32 threads, worst sample β [docs/parity.md](https://aselimc.github.io/jepa.cpp/parity/). α΅ `scripts/gguf_dequant_selftest.py`: the dequantized weights through the numpy reference graph at f32 activations, so the figure is the weight error alone β [docs/quantization.md](https://aselimc.github.io/jepa.cpp/quantization/). `cos mean` is the mean per-token cosine of `last_hidden_state`, `worst` its single worst token.</sub>
|
| 63 |
+
|
| 64 |
+
**Tiers.** `exact` β reproduces the PyTorch reference to the printed precision on the CPU. `parity` β
|
| 65 |
+
passes its family's `test-parity` thresholds. `advisory` β below 8 bits per weight, which is not a parity
|
| 66 |
+
configuration: the results are reported, only the derived tensors and the top-1 are gated. Which file to
|
| 67 |
+
ship: [Accuracy β which dtype](https://aselimc.github.io/jepa.cpp/accuracy/#which-dtype-to-ship).
|
| 68 |
+
`vjepa2-ac-vitg-q8_0.gguf` multi-step rollouts.
|
| 69 |
+
`vjepa2-ac-vitg-q4_0.gguf` multi-step rollouts.
|
| 70 |
+
`vjepa2-ac-vitg-q4_k.gguf` multi-step rollouts.
|
| 71 |
+
|
| 72 |
+
Full checksums:
|
| 73 |
+
|
| 74 |
+
```
|
| 75 |
+
234b7e775da2ff28e29d6b12a48fd28f6b0e888b1854581a684dc8e8e84848bb vjepa2-ac-vitg-f32.gguf
|
| 76 |
+
d57cc1760f90f779eba12d90d04b5028334a95f9b6960da6a98239b9a6f2441c vjepa2-ac-vitg-f16.gguf
|
| 77 |
+
3dffe7130d7f7efd0805d2d714e72a232632077cc59a9036b7896435f6c40ef6 vjepa2-ac-vitg-q8_0.gguf
|
| 78 |
+
554f3a9f293a2b80d3e0f9dd40f033e4ca1a291640712606a727a419ee0f4425 vjepa2-ac-vitg-q4_0.gguf
|
| 79 |
+
0d2fb4e022e2d1ca9a3aaf93d5ae3ad51b1d41244b04f634561ba2ac48a5dfbc vjepa2-ac-vitg-q4_k.gguf
|
| 80 |
+
```
|
| 81 |
+
|
| 82 |
+
Verify a download with `sha256sum -c`. The other types `jepa-quantize` can produce (`q4_1`, `q5_0`,
|
| 83 |
+
`q5_1`, `q5_k`, `q6_k`, measured in [quantization](https://aselimc.github.io/jepa.cpp/quantization/)) are not published here; make
|
| 84 |
+
them locally with `build/jepa-quantize vjepa2-ac-vitg-f16.gguf out.gguf q6_k -t 32`.
|
| 85 |
+
|
| 86 |
+
## Measured
|
| 87 |
+
|
| 88 |
+
Every figure below is read from a committed artifact of jepa.cpp [`c952229`](https://github.com/aselimc/jepa.cpp/commit/c952229) by `scripts/hf_publish.py` β [parity](https://aselimc.github.io/jepa.cpp/parity/), [quantization](https://aselimc.github.io/jepa.cpp/quantization/), [accuracy](https://aselimc.github.io/jepa.cpp/accuracy/), [performance](https://aselimc.github.io/jepa.cpp/performance/) and `tests/results/*.json`.
|
| 89 |
+
|
| 90 |
+
**Plan with f16.** A rollout compounds β step h's input is step h-1's output β so the worst predicted
|
| 91 |
+
token of a 2-step rollout falls from 0.9925 at f16 to 0.9368 at q8_0 to 0.5429 at q4_k, and at q4_k on
|
| 92 |
+
a GPU the planning energy misranks the candidates and the model picks a different action. The encoder
|
| 93 |
+
half of the bundle passes at every tier on both backends. At f32 on the CPU the predictor is exact to
|
| 94 |
+
cosine 1.0000000 against Meta's own world-model dump, and K candidates batched on the graph's batch
|
| 95 |
+
axis are bit-identical to K sequential rollouts. Full tables in [parity](https://aselimc.github.io/jepa.cpp/parity/).
|
| 96 |
+
|
| 97 |
+
## Source, licence and attribution
|
| 98 |
+
|
| 99 |
+
Converted from [`dl.fbaipublicfiles.com/vjepa2/vjepa2-ac-vitg.pt`](https://dl.fbaipublicfiles.com/vjepa2/vjepa2-ac-vitg.pt).
|
| 100 |
+
|
| 101 |
+
**MIT.** The checkpoint is published by **Meta AI (FAIR)** at
|
| 102 |
+
`https://dl.fbaipublicfiles.com/vjepa2/vjepa2-ac-vitg.pt`; the licence is the
|
| 103 |
+
[LICENSE](https://github.com/facebookresearch/vjepa2/blob/main/LICENSE) of `facebookresearch/vjepa2`
|
| 104 |
+
(Copyright (c) Meta Platforms, Inc. and affiliates). No gating, no acceptable-use policy. These GGUF
|
| 105 |
+
files are the same weights re-serialised, quantized where the file name says so. Cite the
|
| 106 |
+
[V-JEPA 2 paper](https://arxiv.org/abs/2506.09985).
|
| 107 |
+
|
| 108 |
+
The licence travels inside every GGUF as `general.license` and the origin as `general.source_url`;
|
| 109 |
+
`build/jepa-info <file> --kv` prints them. jepa.cpp's own code is MIT.
|
| 110 |
+
|
| 111 |
+
## Conversion
|
| 112 |
+
|
| 113 |
+
Produced by jepa.cpp [`c952229`](https://github.com/aselimc/jepa.cpp/commit/c952229):
|
| 114 |
+
|
| 115 |
+
```bash
|
| 116 |
+
scripts/download_models.sh --convert vjepa2-ac
|
| 117 |
+
python scripts/convert.py --family vjepa2_ac --src models/vjepa2_ac/vjepa2-ac-vitg.pt --out models/gguf/vjepa2-ac-vitg-f16.gguf --ftype f16
|
| 118 |
+
# ... and again with --ftype f32 --out models/gguf/vjepa2-ac-vitg-f32.gguf for the f32 file
|
| 119 |
+
|
| 120 |
+
for q in q8_0 q4_0 q4_k; do
|
| 121 |
+
build/jepa-quantize models/gguf/vjepa2-ac-vitg-f16.gguf \
|
| 122 |
+
models/gguf/vjepa2-ac-vitg-$q.gguf $q -t 32
|
| 123 |
+
done
|
| 124 |
+
```
|
| 125 |
+
|
| 126 |
+
`jepa-quantize` re-types only the 2-D attention / FFN / projection / classifier matrices; patch
|
| 127 |
+
embeddings, position tables, norms and biases keep the source type. The rules are in
|
| 128 |
+
[docs/gguf-schema.md](https://aselimc.github.io/jepa.cpp/gguf-schema/).
|
| 129 |
+
|
| 130 |
+
## Links
|
| 131 |
+
|
| 132 |
+
- Code: <https://github.com/aselimc/jepa.cpp>
|
| 133 |
+
- Documentation: <https://aselimc.github.io/jepa.cpp/>
|
| 134 |
+
- Parity fixtures: <https://huggingface.co/datasets/jepacpp/jepa.cpp-fixtures>
|
| 135 |
+
- All jepa.cpp GGUFs: <https://huggingface.co/jepacpp>
|