harrier-0.6b vector-quantised query clients (.vqw, 105–120 MiB)

Query encoders for microsoft/harrier-oss-v1-0.6b at 1.8–2.1 bits per weight, below the smallest format llama.cpp offers (IQ2_XS, 177.5 MiB). They are meant for the same use as the GGUF clients: your document index built with harrier-0.6b stays as it is, only the query encoder moves to the client. These files are read by the thinletter WebGPU runtime, which is open source under Apache-2.0 in the public repository (client/vqweb). The quickest way to try them is the public demo at https://thinletter.io/demo (pick vq2.0 or vq1.75 in the Client selector); locally, serve the repository with python client/browser/serve.py --mount vqw=<directory of the file> and open http://localhost:8765/client/vqweb/bench.html?vqw=/vqw/<file>.vqw&dataset=scidocs&split=test (parameters documented in client/vqweb/bench.js and client/vqweb/README.md), or run the whole benchmark with python scripts/browser_run.py --runtime vqweb --vqw <file.vqw> --tag <tag>.

file bits/weight (blocks) nDCG@10 on SciDocs (% of fp32) browser check
harrier-0.6b-vq2.0-scidocs.vqw (127.0 MiB) 2.10 (K = 256, 4-d codebooks) 0.2150 = 94.7 % of fp32 0.2269 (cosine 0.935), torch simulation on 500 test queries browser (WebGPU, integrated GPU) reproduces the simulation on 50 queries to βˆ’0.0002, 49/50 identical
harrier-0.6b-vq1.75-scidocs.vqw (112.8 MiB) 1.83 (K = 128) 0.2073 = 91.3 % (cosine 0.906) browser +0.0009, 47/50 identical
harrier-0.6b-vq2.0-scifact.vqw (119.5 MiB) 2.10 (K = 256) SciFact: 0.7465 = 96.7 % of fp32 0.7723 on all 1 109 queries (cosine 0.854), torch simulation; test split 0.7380 browser 0.7375 on the 300 test queries (Ξ” βˆ’0.0005 vs simulation), p50 142 ms under load / 103 ms idle on an integrated GPU
harrier-0.6b-vq1.75-scifact.vqw (105.3 MiB) 1.83 (K = 128) SciFact: 0.7243 = 93.8 % (cosine 0.771); test split 0.7109 browser 0.7109 (Ξ” 0.0000)

The two SciFact-calibrated files (calibration: 2 000 SciFact abstracts, vocabulary trimmed to the SciFact corpus + synthetic queries) were added on 2026-09-16 when the runtime was open-sourced; they are the files of report Β§3.6 and vq_results.md Β§1–2. The SciFact corpus itself is CC BY-NC, which is why the public demo runs the SciDocs files; the weights are derived from the MIT base model. tokenizer.json and tokenizer_config.json (the harrier tokenizer, referenced by every container header) are in this repository so that the runtime can load a file directly from here.

Measured on SciFact with the earlier (non-distributable) SciFact-calibrated containers of the same recipe: 2.10 bpw 0.7375 in the browser (97.6 % of fp32 0.7559; torch simulation 0.7380), 1.83 bpw 0.7109 (94.0 %); llama.cpp's IQ2_XS at 177.5 MiB: 0.7097. On one laptop with an Intel integrated GPU, idle, 300 queries: 2.10 bpw p50 103 ms / p95 113 ms, load 2.4 s, 181 MiB + 151 MiB GPU buffers in the tab, peak process memory 1 592 MiB; llama.cpp's WebGPU path in wllama with the 235 MiB Q3_K file: 444 / 754 ms, 4.2 s, 530 MiB, 2 080 MiB. Full tables and caveats: https://github.com/rosecky/embedding-quantization-public/blob/main/docs/release/vq_results.md.

Prompt-K/V variants (added 2026-09-16; need the prefix path of the runtime)

Four more files carry the query prompt's keys and values in full precision (2 MB) and were calibrated with them in place (report Β§3.7, "prefix-aware calibration"). They are re-exports of 2026-09-16 with a fresh synthetic-query draw (the originals were lost), so they are an independent replication of the published effect; one seed each. The runtime reads the stored prefix automatically (client/vqweb, M6); a runtime that ignores it gets 53–90 % of fp32 from these files, which is why the plain files above remain the default. Numbers: torch simulation of the file with the prefix, and the browser on the first 50 test queries.

file MiB bits/weight (blocks) nDCG@10 with prefix (% of fp32) / cosine / overlap browser (WebGPU, integrated GPU)
harrier-0.6b-vq2.0p-scidocs.vqw 129.1 2.10 SciDocs test: 0.2191 = 96.6 % / 0.942 / 0.784 (plain file: 94.7 %) 0.2058 vs 0.2065 simulation on 50 queries (Ξ” βˆ’0.0007), p50 67 ms
harrier-0.6b-vq1.75p-scidocs.vqw 114.9 1.83 SciDocs test: 0.2129 = 93.8 % / 0.929 / 0.760 (plain file: 91.3 %; calibration on title-shaped queries) 0.1920 vs 0.1926 (Ξ” βˆ’0.0006), p50 75 ms
harrier-0.6b-vq2.0p-scifact.vqw 121.6 2.10 SciFact, all 1 109 queries: 0.7571 = 98.1 % / 0.941 / 0.805 not run
harrier-0.6b-vq1.75p-scifact.vqw 107.3 1.83 SciFact: 0.7397 = 95.8 % / 0.913 / 0.754 (plain file 93.8 %) not run

Reading: at 1.8 bits the prompt-aware calibration is worth +2 points on SciFact and +2.5 on SciDocs over the plain files; at 2.1 bits +1.4 (SciFact, against the corpus-calibrated plain file) and +1.9 (SciDocs). Single-seed numbers within Β±0.01 of the three-seed results published in the report. Demo: SciDocs corpus, clients vq2.0p / vq1.75p.

Quickstart in the browser (@thinletterio/vqweb)

import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js';   // or: npm install @thinletterio/vqweb

const client = await loadVqwClient('https://huggingface.co/thinletter/harrier-0.6b-vq-clients/resolve/main/harrier-0.6b-vq2.0p-scifact.vqw');
const q = await client.embed('Does vitamin D reduce the risk of respiratory infections?');   // Float32Array(1024), L2-normalised: dot it with your microsoft/harrier-oss-v1-0.6b index

The package (Apache-2.0, npm, source) downloads the file with progress, keeps it in the browser's private file system for the next visit, uploads it to the GPU, takes the tokenizer and the prompt from the container header and compiles the WebGPU pipelines; client.embed(text) takes the raw query. Needs WebGPU with shader-f16 (Chrome / Edge; probeWebGPU() says why not), no CPU fallback. About 100 ms per query on an integrated GPU, 2–3 s to load after the download.

What is in a .vqw file

One binary: an 8-byte magic, a JSON header (model shape, tokenizer and prompt settings, rotation seeds, a tensor table) and 64-byte-aligned tensor data. Per block linear: f16 codebooks (one per 256-column block of the rotated input space, K Γ— 4 entries), packed 6/7/8-bit indices (one per 4 consecutive columns), f16 per-row scales. Norm weights f16, the token table in llama.cpp's Q2_K block layout trimmed to the vocabulary of the calibration corpus plus a byte fallback, and the sign/permutation data of the input-side structured rotation. The header is readable with any JSON parser; the quantiser that wrote it (scripts/gptvq_harrier.py, scripts/vq_export_web.py) and the WebGPU kernels that consume it (client/vqweb/) are in the public repository under Apache-2.0.

Compatibility

Document vectors from microsoft/harrier-oss-v1-0.6b, no document prompt, last-token pooling, L2, 1024-d. Query prompt (stored in the header): the E5-style web-search instruction, EOS appended. Not compatible with other models' indexes.

Limits

One model, two corpora; quality below 2.2 bits per weight on a Czech index with other models of the same architecture was 83 % (jina-v5-small) and 64 % (Qwen3-Embedding-0.6B), so this rate is for English corpora and this model until a ~3-bit variant exists. Latency and memory were measured on one integrated GPU; no discrete-GPU or mobile numbers.

sha256: e5891a4a54351156fa63681aec4c3d864b22d66b2299bd2cec7385543ef56cdf (2.10 bpw), 0228e0ab4d820f7983f5c124dab84a8c2367f8ad002b48f3a19bc34cac548326 (1.83 bpw). SciFact files: 01f7bef3c9caa5c1d5e48558b8a09e6d83931169387b003d9d40ac2224e9a4e9 (2.10 bpw), 2711df96566c6c43984666955ee16d203b1f29eef27b01e431f540451bab3883 (1.83 bpw). For comparison on the same index, the scalar GGUF clients of this model: Q3_K 235 MiB 99.2 %, Q2_K (SciDocs-synthetic) 192 MiB 94.5 %, i.e. the 2.10 bpw container matches the 2.6-bit GGUF at two thirds of its size.

The prompt's K/V in full precision: measured, not an improvement for these files (2026-09-10)

The query prompt is the same 19 tokens in front of every query, so its keys and values in every block can be computed once by the full-precision model and shipped with the client (β‰ˆ 2.2 MB; the first token alone 112 KB). On the same quantized model this gives +0.009 to +0.020 nDCG@10 at 1.84 bpw (three seeds, every paired CI above zero) and +0.037 at 1.58 bpw (87.5 β†’ 92.3 % of fp32 on SciFact); the first token alone carries most of it (the 2-bit model forms the attention sink wrongly, an exact first token repairs it). The containers here do not carry the prompt K/V yet and the runtime does not read it yet; the numbers are in the report (Β§3.7) and in results/tables/prefix_kv.md of the public repository. On the scalar GGUF clients the same trick is worth +0.001 and is not planned.

Addendum (same day, evening): those gains were measured on quantizations calibrated on documents. The files in this repository are calibrated on synthetic queries (the prompt included), and on them the same prompt K/V gives βˆ’0.003 [βˆ’0.007; βˆ’0.000] (1.83 bpw) and +0.000 [βˆ’0.002; +0.003] (2.10 bpw): a model calibrated with the prompt already reproduces it. So this is not an improvement pending for these files; whether calibrating with the exact prompt in place changes that is being measured, and the files here will only be replaced if it does (β‰₯ +0.005 nDCG@10, paired CI above zero).

Addendum 2 (2026-09-11): measured. Calibrating every block with the exact prompt in place and shipping it gives +0.008 to +0.013 nDCG@10 over the best query-calibrated quantization at 1.83 bpw on SciFact (three seeds, every paired CI above zero), but on SciDocs, the corpus of these files, +0.002 / +0.004 / +0.000 on three seeds, nothing on NFCorpus (βˆ’0.000 / +0.002) or ArguAna (+0.001 / βˆ’0.009), and nothing at 2.10 bpw. The cosine to fp32 rises everywhere; the ranking follows on SciFact and on a Czech legal index (+0.007 / +0.016 / +0.010, jina-v5-small), i.e. where the synthetic calibration queries have the shape of the real ones. Re-calibrating SciDocs on title-shaped queries brings the prefix-aware 1.83 bpw container to 93.3–94.3 % (browser-verified), +0.0045 [βˆ’0.002; +0.011] over the file here: not enough for the release rule, so these files stay as they are; the recipe lives in the runtime as an option. Numbers: report Β§3.7 addenda 2–3, docs/release/vq_results.md Β§6.

Licence

Weights derive from microsoft/harrier-oss-v1-0.6b (MIT) and are released under MIT; calibration text: synthetic queries generated from SciDocs documents (SciDocs, CC BY 4.0). The container format, its compiler and the WebGPU runtime are Apache-2.0 (the public repository's LICENSE; third-party components in its NOTICE).

Questions: info@thinletter.io, or the issues of the public repository.

Downloads last month
90
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for thinletter/harrier-0.6b-vq-clients

Finetuned
(13)
this model