V-JEPA 2.1 ViT-B/16 @384 β€” GGUF for jepa.cpp

Meta's V-JEPA 2.1 ViT-B/16 encoder at 384x384, distilled from ViT-g, with its masked latent predictor β€” converted to GGUF for jepa.cpp, a ggml C/C++ engine that runs it on a plain CPU with no Python and no PyTorch. The one model here that handles both modalities: a still image goes through the image tokenizer and its own modality vector, a clip through the tubelet tokenizer.

110 M parameters; D = 768, 12 layers, 12 heads, patch 16, tubelet 2, 384x384. Everything the engine needs β€” dimensions, positional scheme, preprocessing recipe, and class labels where there are any β€” travels inside the file, so inference needs one binary and one GGUF and nothing else.

Run it

git clone --recursive https://github.com/aselimc/jepa.cpp && cd jepa.cpp
cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release && cmake --build build -j
hf download jepacpp/vjepa2_1-vitb-384-GGUF vjepa2_1-vitb-384-f16.gguf --local-dir models/gguf

# an image and a clip through the same file
build/jepa-embed -m vjepa2_1-vitb-384-f16.gguf -i photo.jpg          --pool mean -t 32
build/jepa-embed -m vjepa2_1-vitb-384-f16.gguf --frames-npy clip.npy --pool mean -t 32 -o feat.npy

--pool selects mean, cls, lewm or none (the full token map); -o writes a .npy. scripts/download_models.sh fetches whole sets at once. The C API is one header, include/jepa.h β€” full reference on the C API page.

Files

file size sha256 (first 16) tier measured against the PyTorch reference
vjepa2_1-vitb-384-f32.gguf 418.5 MiB cb9dd396b2d27ac5 exact cos mean 1.000000, median 1.000000, worst 1.000000, pooled_mean 1.000000, rel_max 8.7e-05 α΅–
vjepa2_1-vitb-384-f16.gguf 209.7 MiB 54cef7642f0ef4d5 parity cos mean 0.999952, median 0.999991, worst 0.9697, pooled_mean 1.000000 α΅–
vjepa2_1-vitb-384-q8_0.gguf 113.3 MiB 2c6fa6bed8349e06 parity cos mean 0.999076, median 0.999578, worst 0.8302, pooled_mean 0.999986 α΅–
vjepa2_1-vitb-384-q4_0.gguf 61.9 MiB de6dc6336ed5959e advisory cos mean 0.976396, worst 0.734594, pooled_mean 0.993697 ᡈ
vjepa2_1-vitb-384-q4_k.gguf 61.9 MiB 73ed90dab19b2293 advisory cos mean 0.990260, worst 0.751045, pooled_mean 0.998225 ᡈ

α΅– tests/test-parity on the CPU backend, stored reference input, 32 threads, worst sample β€” docs/parity.md. ᡈ scripts/gguf_dequant_selftest.py: the dequantized weights through the numpy reference graph at f32 activations, so the figure is the weight error alone β€” docs/quantization.md. cos mean is the mean per-token cosine of last_hidden_state, worst its single worst token.

Tiers. exact β€” reproduces the PyTorch reference to the printed precision on the CPU. parity β€” passes its family's test-parity thresholds. advisory β€” below 8 bits per weight, which is not a parity configuration: the results are reported, only the derived tensors and the top-1 are gated. Which file to ship: Accuracy β†’ which dtype.

Full checksums:

cb9dd396b2d27ac5b85c78883cb0bcda7082f106ff5fde396e5de50130c3b504  vjepa2_1-vitb-384-f32.gguf
54cef7642f0ef4d59a76f1a100e6a6dc0585beb21086368a528ecbcad36359aa  vjepa2_1-vitb-384-f16.gguf
2c6fa6bed8349e06f8b2f63e0e7550791bea16971aba0f66493e4dbcca673d95  vjepa2_1-vitb-384-q8_0.gguf
de6dc6336ed5959ee22f128247e162a9bdde7f63653cddddbb62060284fdb6e9  vjepa2_1-vitb-384-q4_0.gguf
73ed90dab19b22938754f7fea085fc520acffd30ddb8999e85b70d7afdaccdc6  vjepa2_1-vitb-384-q4_k.gguf

Verify a download with sha256sum -c. The other types jepa-quantize can produce (q4_1, q5_0, q5_1, q5_k, q6_k, measured in quantization) are not published here; make them locally with build/jepa-quantize vjepa2_1-vitb-384-f16.gguf out.gguf q6_k -t 32.

Measured

Every figure below is read from a committed artifact of jepa.cpp 00bfd4e by scripts/hf_publish.py β€” parity, quantization, accuracy, performance and tests/results/*.json.

UCF-101 k-NN β€” 10 classes, 105 query clips (val+test) against a gallery of 300, 16 frames per clip, k = 20 cosine vote over frozen features. Nothing is trained.

backend dtype k-NN top-1 % centroid top-1 % k-NN agreement % centroid agreement % feature cosine
pytorch f32 88.57 86.67 β€” β€” β€”
jepa.cpp f32 88.57 86.67 100.00 100.00 1.000000
jepa.cpp f16 89.52 86.67 99.05 100.00 1.000000
jepa.cpp q8_0 89.52 86.67 99.05 100.00 0.999988

Speed β€” the encoder graph at f16 on 32 threads (AMD Ryzen Threadripper PRO 7995WX 96-Cores): 60 ms per image against PyTorch's 110 ms; 853 ms per 16-frame clip against PyTorch's 908 ms; 9036 ms per 64-frame clip. The same shape on NVIDIA RTX 4500 Ada Generation: 4.4 ms. Peak RSS at f16: 948 MiB.

Far more forgiving at reduced precision than the V-JEPA 2 ViT-L encoders β€” its worst-token column stays high at f16 β€” so f16 is a per-token-grade configuration here, not only a pooled-feature-grade one.

Source, licence and attribution

Converted from dl.fbaipublicfiles.com/vjepa2/vjepa2_1_vitb_dist_vitG_384.pt.

MIT. The source is a bare .pt checkpoint published by Meta AI (FAIR) from the V-JEPA 2.1 checkpoint table of facebookresearch/vjepa2, whose LICENSE is MIT (Copyright (c) Meta Platforms, Inc. and affiliates). The file itself carries no licence metadata, so MIT here is the repository-level grant that publishes it rather than a per-file one; there is no gating and no acceptable-use policy. These GGUF files are the same weights re-serialised into the GGUF container, quantized where the file name says so. Cite the V-JEPA 2 paper.

The licence travels inside every GGUF as general.license and the origin as general.source_url; build/jepa-info <file> --kv prints them. jepa.cpp's own code is MIT.

Conversion

Produced by jepa.cpp 00bfd4e:

scripts/download_models.sh --convert vjepa21
python scripts/convert.py --family vjepa2_1 --src models/vjepa2_1/vjepa2_1_vitb_dist_vitG_384.pt \
                          --out models/gguf/vjepa2_1-vitb-384-f16.gguf
#   ... and again with --ftype f32 --out models/gguf/vjepa2_1-vitb-384-f32.gguf for the f32 file

for q in q8_0 q4_0 q4_k; do
  build/jepa-quantize models/gguf/vjepa2_1-vitb-384-f16.gguf \
      models/gguf/vjepa2_1-vitb-384-$q.gguf $q -t 32
done

jepa-quantize re-types only the 2-D attention / FFN / projection / classifier matrices; patch embeddings, position tables, norms and biases keep the source type. The rules are in docs/gguf-schema.md.

Links

Downloads last month
29
GGUF
Model size
0.1B params
Architecture
jepa
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for jepacpp/vjepa2_1-vitb-384-GGUF