aselimc commited on
Commit
ea6f6cc
Β·
verified Β·
1 Parent(s): fbb844b

Update model card

Browse files
Files changed (1) hide show
  1. README.md +135 -0
README.md ADDED
@@ -0,0 +1,135 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ library_name: jepa.cpp
4
+ tags:
5
+ - jepa
6
+ - ggml
7
+ - gguf
8
+ - jepa.cpp
9
+ - v-jepa
10
+ - v-jepa-2
11
+ - world-model
12
+ - robotics
13
+ - video
14
+ - planning
15
+ ---
16
+
17
+ # V-JEPA 2-AC ViT-g (action-conditioned world model) β€” GGUF for jepa.cpp
18
+
19
+ Meta's **V-JEPA 2-AC** β€” the action-conditioned world model behind the paper's zero-shot robot
20
+ planning β€” converted to GGUF for [jepa.cpp](https://github.com/aselimc/jepa.cpp), a ggml C/C++ engine that runs it on a plain CPU
21
+ with no Python and no PyTorch. One bundle: the frozen ViT-g/16 encoder from `vjepa2-ac-vitg.pt` plus
22
+ the 24-layer, 1024-dim predictor that takes a 7-d end-effector action and a 7-d pose per frame and
23
+ predicts the next frame's latents, block-causally over frames. `jepa_ac_rollout` scores K candidate
24
+ action sequences in one graph per step, and `jepa_ac_energy` is the L1 planning energy a CEM planner
25
+ minimises.
26
+
27
+ **Note on the encoder:** it is *not* `facebook/vjepa2-vitg-fpc64-256`. Meta's AC checkpoint carries its
28
+ own frozen ViT-g, which agrees with the HF release only to cosine ~0.998 per tensor; `encoder` and
29
+ `target_encoder` inside it are bit-identical. This bundle ships the checkpoint's own, which is what
30
+ `vjepa2_ac_vit_giant` loads.
31
+
32
+ **1317 M** parameters; D = 1408, 40 layers, 22 heads, patch 16, tubelet 2, 256x256. Everything the engine needs β€” dimensions, positional scheme,
33
+ preprocessing recipe, and class labels where there are any β€” travels inside the file, so inference needs
34
+ one binary and one GGUF and nothing else.
35
+
36
+ ## Run it
37
+
38
+ ```bash
39
+ git clone --recursive https://github.com/aselimc/jepa.cpp && cd jepa.cpp
40
+ cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release && cmake --build build -j
41
+ hf download jepacpp/vjepa2-ac-vitg-GGUF vjepa2-ac-vitg-f16.gguf --local-dir models/gguf
42
+
43
+ # encode a frame, roll 4 candidate action sequences out 2 steps, score them against a goal
44
+ build/jepa-worldmodel --ac -m vjepa2-ac-vitg-f16.gguf --image ctx.png --goal goal.png \
45
+ --actions-npy actions.npy # float32 [K, H, 7]
46
+ ```
47
+
48
+ `--pool` selects `mean`, `cls`, `lewm` or `none` (the full token map); `-o` writes a `.npy`.
49
+ `scripts/download_models.sh` fetches whole sets at once. The C API is one header,
50
+ [`include/jepa.h`](https://github.com/aselimc/jepa.cpp/blob/main/include/jepa.h) β€” full reference on the [C API page](https://aselimc.github.io/jepa.cpp/api/).
51
+
52
+ ## Files
53
+
54
+ | file | size | sha256 (first 16) | tier | measured against the PyTorch reference |
55
+ |---|---|---|---|---|
56
+ | `vjepa2-ac-vitg-f32.gguf` | 5025.5 MiB | `234b7e775da2ff28` | exact | β€” |
57
+ | `vjepa2-ac-vitg-f16.gguf` | 2519.0 MiB | `d57cc1760f90f779` | parity | β€” |
58
+ | `vjepa2-ac-vitg-q8_0.gguf` | 1344.1 MiB | `3dffe7130d7f7efd` | parity, **below the bar** | β€” |
59
+ | `vjepa2-ac-vitg-q4_0.gguf` | 717.4 MiB | `554f3a9f293a2b80` | advisory, **below the bar** | β€” |
60
+ | `vjepa2-ac-vitg-q4_k.gguf` | 717.4 MiB | `0d2fb4e022e2d1ca` | advisory, **below the bar** | β€” |
61
+
62
+ <sub>α΅– `tests/test-parity` on the CPU backend, stored reference input, 32 threads, worst sample β€” [docs/parity.md](https://aselimc.github.io/jepa.cpp/parity/). ᡈ `scripts/gguf_dequant_selftest.py`: the dequantized weights through the numpy reference graph at f32 activations, so the figure is the weight error alone β€” [docs/quantization.md](https://aselimc.github.io/jepa.cpp/quantization/). `cos mean` is the mean per-token cosine of `last_hidden_state`, `worst` its single worst token.</sub>
63
+
64
+ **Tiers.** `exact` β€” reproduces the PyTorch reference to the printed precision on the CPU. `parity` β€”
65
+ passes its family's `test-parity` thresholds. `advisory` β€” below 8 bits per weight, which is not a parity
66
+ configuration: the results are reported, only the derived tensors and the top-1 are gated. Which file to
67
+ ship: [Accuracy β†’ which dtype](https://aselimc.github.io/jepa.cpp/accuracy/#which-dtype-to-ship).
68
+ `vjepa2-ac-vitg-q8_0.gguf` multi-step rollouts.
69
+ `vjepa2-ac-vitg-q4_0.gguf` multi-step rollouts.
70
+ `vjepa2-ac-vitg-q4_k.gguf` multi-step rollouts.
71
+
72
+ Full checksums:
73
+
74
+ ```
75
+ 234b7e775da2ff28e29d6b12a48fd28f6b0e888b1854581a684dc8e8e84848bb vjepa2-ac-vitg-f32.gguf
76
+ d57cc1760f90f779eba12d90d04b5028334a95f9b6960da6a98239b9a6f2441c vjepa2-ac-vitg-f16.gguf
77
+ 3dffe7130d7f7efd0805d2d714e72a232632077cc59a9036b7896435f6c40ef6 vjepa2-ac-vitg-q8_0.gguf
78
+ 554f3a9f293a2b80d3e0f9dd40f033e4ca1a291640712606a727a419ee0f4425 vjepa2-ac-vitg-q4_0.gguf
79
+ 0d2fb4e022e2d1ca9a3aaf93d5ae3ad51b1d41244b04f634561ba2ac48a5dfbc vjepa2-ac-vitg-q4_k.gguf
80
+ ```
81
+
82
+ Verify a download with `sha256sum -c`. The other types `jepa-quantize` can produce (`q4_1`, `q5_0`,
83
+ `q5_1`, `q5_k`, `q6_k`, measured in [quantization](https://aselimc.github.io/jepa.cpp/quantization/)) are not published here; make
84
+ them locally with `build/jepa-quantize vjepa2-ac-vitg-f16.gguf out.gguf q6_k -t 32`.
85
+
86
+ ## Measured
87
+
88
+ Every figure below is read from a committed artifact of jepa.cpp [`c952229`](https://github.com/aselimc/jepa.cpp/commit/c952229) by `scripts/hf_publish.py` β€” [parity](https://aselimc.github.io/jepa.cpp/parity/), [quantization](https://aselimc.github.io/jepa.cpp/quantization/), [accuracy](https://aselimc.github.io/jepa.cpp/accuracy/), [performance](https://aselimc.github.io/jepa.cpp/performance/) and `tests/results/*.json`.
89
+
90
+ **Plan with f16.** A rollout compounds β€” step h's input is step h-1's output β€” so the worst predicted
91
+ token of a 2-step rollout falls from 0.9925 at f16 to 0.9368 at q8_0 to 0.5429 at q4_k, and at q4_k on
92
+ a GPU the planning energy misranks the candidates and the model picks a different action. The encoder
93
+ half of the bundle passes at every tier on both backends. At f32 on the CPU the predictor is exact to
94
+ cosine 1.0000000 against Meta's own world-model dump, and K candidates batched on the graph's batch
95
+ axis are bit-identical to K sequential rollouts. Full tables in [parity](https://aselimc.github.io/jepa.cpp/parity/).
96
+
97
+ ## Source, licence and attribution
98
+
99
+ Converted from [`dl.fbaipublicfiles.com/vjepa2/vjepa2-ac-vitg.pt`](https://dl.fbaipublicfiles.com/vjepa2/vjepa2-ac-vitg.pt).
100
+
101
+ **MIT.** The checkpoint is published by **Meta AI (FAIR)** at
102
+ `https://dl.fbaipublicfiles.com/vjepa2/vjepa2-ac-vitg.pt`; the licence is the
103
+ [LICENSE](https://github.com/facebookresearch/vjepa2/blob/main/LICENSE) of `facebookresearch/vjepa2`
104
+ (Copyright (c) Meta Platforms, Inc. and affiliates). No gating, no acceptable-use policy. These GGUF
105
+ files are the same weights re-serialised, quantized where the file name says so. Cite the
106
+ [V-JEPA 2 paper](https://arxiv.org/abs/2506.09985).
107
+
108
+ The licence travels inside every GGUF as `general.license` and the origin as `general.source_url`;
109
+ `build/jepa-info <file> --kv` prints them. jepa.cpp's own code is MIT.
110
+
111
+ ## Conversion
112
+
113
+ Produced by jepa.cpp [`c952229`](https://github.com/aselimc/jepa.cpp/commit/c952229):
114
+
115
+ ```bash
116
+ scripts/download_models.sh --convert vjepa2-ac
117
+ python scripts/convert.py --family vjepa2_ac --src models/vjepa2_ac/vjepa2-ac-vitg.pt --out models/gguf/vjepa2-ac-vitg-f16.gguf --ftype f16
118
+ # ... and again with --ftype f32 --out models/gguf/vjepa2-ac-vitg-f32.gguf for the f32 file
119
+
120
+ for q in q8_0 q4_0 q4_k; do
121
+ build/jepa-quantize models/gguf/vjepa2-ac-vitg-f16.gguf \
122
+ models/gguf/vjepa2-ac-vitg-$q.gguf $q -t 32
123
+ done
124
+ ```
125
+
126
+ `jepa-quantize` re-types only the 2-D attention / FFN / projection / classifier matrices; patch
127
+ embeddings, position tables, norms and biases keep the source type. The rules are in
128
+ [docs/gguf-schema.md](https://aselimc.github.io/jepa.cpp/gguf-schema/).
129
+
130
+ ## Links
131
+
132
+ - Code: <https://github.com/aselimc/jepa.cpp>
133
+ - Documentation: <https://aselimc.github.io/jepa.cpp/>
134
+ - Parity fixtures: <https://huggingface.co/datasets/jepacpp/jepa.cpp-fixtures>
135
+ - All jepa.cpp GGUFs: <https://huggingface.co/jepacpp>