README / index.html
honza-rosecky's picture
org card: quickstart with @thinletterio/vqweb
15159ad verified
Raw History Blame Contribute Delete
9.58 kB
<!doctype html>
<html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width, initial-scale=1"><title>thinletter</title>
<link rel="stylesheet" href="style.css"></head><body><main>
<h1>thinletter — a query encoder as a 120 MiB file, in the browser, open source</h1>
<p>We take an embedding model, leave its document index exactly as it is, and compile the <strong>query side</strong> into a file a tenth of the size that
answers in the browser in about 100 ms. Everything behind the numbers is public under Apache-2.0: the quantiser, the container compiler,
the WebGPU runtime, the verification recipe and every released file. Web: <a href="https://thinletter.io">thinletter.io</a> · code and report:
<a href="https://github.com/rosecky/embedding-quantization-public">github.com/rosecky/embedding-quantization-public</a> · demo:
<a href="https://thinletter.io/demo">thinletter.io/demo</a> or the Space <a href="https://huggingface.co/spaces/thinletter/demo">thinletter/demo</a> · contact: info@thinletter.io</p>
<h2>Quickstart: a released client in your page</h2>
<pre><code class="language-js">import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js'; // or: npm install @thinletterio/vqweb
const client = await loadVqwClient('https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw');
const q = await client.embed('Kolik stojí parkování v centru Prahy?'); // Float32Array(1024), L2-normalised: dot it with your Qwen3-Embedding-0.6B index
</code></pre>
<p><code>@thinletterio/vqweb</code> (<a href="https://www.npmjs.com/package/@thinletterio/vqweb">npm</a>, Apache-2.0) is the open WebGPU runtime as a package: download with
progress, a cache in the browser's private file system, tokenizer and prompt from the container header, pipelines, warm-up. Needs WebGPU
with <code>shader-f16</code> (Chrome / Edge), no CPU fallback; about 100 ms per query on an integrated GPU. The scalar GGUF files run with
<a href="https://github.com/ngxson/wllama">wllama</a> or llama.cpp.</p>
<h2>What holds</h2>
<table>
<thead>
<tr>
<th>number</th>
<th>what it is</th>
<th>where to check</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>98.2 %</strong></td>
<td>of full-precision nDCG@10 from a 120 MiB, 2.1-bit vector-quantised harrier-0.6b client, no retraining: above the trained BitNet-270m (140 MiB, 97.0 %), below any file llama.cpp can produce (178 MiB, 93.9 %). The browser reproduces the simulation (0.7375 vs 0.7380).</td>
<td><a href="https://huggingface.co/thinletter/harrier-0.6b-vq-clients">harrier-0.6b-vq-clients</a>, <a href="https://huggingface.co/datasets/BeIR/scifact">BeIR/scifact</a></td>
</tr>
<tr>
<td><strong>4.3×</strong></td>
<td>faster than llama.cpp's WebGPU path on the same integrated GPU: 103 ms per query for the 119 MiB vector-quantised client against 444 ms for the 235 MiB Q3_K file, in our open WebGPU runtime (<code>client/vqweb</code>).</td>
<td>report §3.4b in the repository</td>
</tr>
<tr>
<td><strong>96.8 %</strong></td>
<td>of fp32 nDCG@10 on the public Czech benchmark WebFAQ-cs from a 237 MiB vector-quantised Qwen3-Embedding-0.6B client (2-d codebooks, 3.6 bits per weight), browser = simulation, 112 ms per query. Czech-calibrated scalar clients: bge-m3 99.2 % (355 MiB), Qwen3 98.7 % (340 MiB) on a Czech legal index.</td>
<td><a href="https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients">qwen3-embedding-0.6b-vq-clients</a>, <a href="https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval">PaDaS-Lab/webfaq-retrieval</a></td>
</tr>
<tr>
<td><strong>+0.010</strong></td>
<td>nDCG@10 at 1.8 bits from a method of ours: ship the query prompt's keys and values in full precision (2 MB) and calibrate the compressed model with them in place; three seeds on SciFact and on a Czech legal index, every paired interval above zero.</td>
<td>report §3.7; the <code>*p</code> files in <a href="https://huggingface.co/thinletter/harrier-0.6b-vq-clients">harrier-0.6b-vq-clients</a></td>
</tr>
<tr>
<td><strong>96–101 %</strong></td>
<td>kept by the 3.4-bit harrier-0.6b scalar client on four English corpora, 235 MiB instead of 1 143 MiB. Not ours to claim: llama.cpp's own quantiser gets there too, and we say so.</td>
<td><a href="https://huggingface.co/thinletter/harrier-0.6b-query-clients">harrier-0.6b-query-clients</a></td>
</tr>
</tbody>
</table>
<h2>Repositories</h2>
<p>Query-side clients: each file is a query encoder for exactly one document encoder and its settings (the model card says which).
The index does not change.</p>
<table>
<thead>
<tr>
<th>repository</th>
<th>base model · licence</th>
<th>files</th>
<th>what holds (test split, the model's own fp32 index)</th>
</tr>
</thead>
<tbody>
<tr>
<td><a href="https://huggingface.co/thinletter/harrier-0.6b-query-clients">harrier-0.6b-query-clients</a></td>
<td><a href="https://huggingface.co/microsoft/harrier-oss-v1-0.6b">microsoft/harrier-oss-v1-0.6b</a> · MIT</td>
<td>3 GGUF, 192–235 MiB</td>
<td>Q3_K generic-English 98.9 / 99.3 / 101 / 99.2 % on <a href="https://huggingface.co/datasets/BeIR/scifact">SciFact</a> / <a href="https://huggingface.co/datasets/BeIR/nfcorpus">NFCorpus</a> / <a href="https://huggingface.co/datasets/BeIR/arguana">ArguAna</a> / <a href="https://huggingface.co/datasets/BeIR/scidocs">SciDocs</a>; Q2_K SciDocs-calibrated 94.5 %</td>
</tr>
<tr>
<td><a href="https://huggingface.co/thinletter/harrier-0.6b-vq-clients">harrier-0.6b-vq-clients</a></td>
<td>microsoft/harrier-oss-v1-0.6b · MIT</td>
<td>8 <code>.vqw</code>, 105–129 MiB, 1.8–2.1 bits per weight</td>
<td>SciDocs 94.7 / 91.3 %, SciFact 98.2 / 96.5 %; prompt-K/V variants SciDocs 96.6 / 93.8 %, SciFact 98.1 / 95.8 %; browser-verified; run in the open WebGPU runtime</td>
</tr>
<tr>
<td><a href="https://huggingface.co/thinletter/qwen3-embedding-0.6b-query-clients">qwen3-embedding-0.6b-query-clients</a></td>
<td><a href="https://huggingface.co/Qwen/Qwen3-Embedding-0.6B">Qwen/Qwen3-Embedding-0.6B</a> · Apache-2.0</td>
<td>4 GGUF, 340–385 MiB, English and Czech calibration</td>
<td>Q4_K_M + 4-bit table: 99.3–100 % on the four English corpora, 98.7 % on a Czech legal index; Q5_K_M 99.5 % Czech</td>
</tr>
<tr>
<td><a href="https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients">qwen3-embedding-0.6b-vq-clients</a></td>
<td>Qwen/Qwen3-Embedding-0.6B · Apache-2.0</td>
<td>1 <code>.vqw</code>, 237 MiB, 2-d codebooks 3.57 bpw + full token table</td>
<td>Czech: 96.8 % of fp32 on <a href="https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval">WebFAQ-cs</a> (cosine 0.946, overlap 0.755), browser = simulation, 112 ms on an integrated GPU</td>
</tr>
<tr>
<td><a href="https://huggingface.co/thinletter/bge-m3-query-clients">bge-m3-query-clients</a></td>
<td><a href="https://huggingface.co/BAAI/bge-m3">BAAI/bge-m3</a> · MIT</td>
<td>2 GGUF, 321–355 MiB, Czech calibration</td>
<td>99.2 % / 98.3 % on the Czech legal index, 99.6–100 % on SciFact</td>
</tr>
<tr>
<td><a href="https://huggingface.co/spaces/thinletter/demo">demo</a> (Space)</td>
<td>—</td>
<td>the browser demo, static</td>
<td>search SciDocs (25 656 abstracts) or WebFAQ-cs (71 529 Czech passages) locally with a scalar or a vector-quantised client; models and indexes download from thinletter.io</td>
</tr>
<tr>
<td><a href="https://huggingface.co/datasets/thinletter/llm-weight-compression-evidence">llm-weight-compression-evidence</a></td>
<td>dataset · MIT</td>
<td>per-window evidence</td>
<td>which objective a post-training quantiser should optimise (companion project)</td>
</tr>
</tbody>
</table>
<p>Each model repository has a summary post in its Community tab with the numbers, what was found along the way, and the limits. The <code>.vqw</code> files load with <code>@thinletterio/vqweb</code> (above); the GGUF files with llama.cpp / wllama.</p>
<h2>How we measure</h2>
<p>Test split, the model's own fp32 index, native runtime (llama.cpp) or the browser. Every run is pre-registered with a prediction and a
kill rule; comparisons change one thing at a time at the same calibration budget; differences are paired bootstrap intervals over queries
(10 000 draws) read against the variance of the calibration draw (~0.01 nDCG@10) and between machines (±0.006): under 0.01 is a tie.
Negative results stay published (a codebook-budget idea that did not work, a 2.6-bit headline that held on one corpus only, the 2-bit cliff
of Qwen3-Embedding: 65–81 % for the same run on two machines).</p>
<h2>Verify on your own index</h2>
<p>The <a href="https://github.com/rosecky/embedding-quantization-public">public repository</a> has the recipe: import your corpus and existing
document vectors, generate synthetic queries, quantise (llama-quantize, the GPTQ exporter, or the vector quantiser + <code>.vqw</code> compiler),
evaluate the client against the unchanged index with a paired interval. Everything in the report is reproducible from it; only the raw
per-run results, the research log and a customer's data stay out.</p>
<h2>Licences</h2>
<p>Code Apache-2.0 (NOTICE lists wllama MIT, @huggingface/tokenizers Apache-2.0, a Q2_K decoder ported from llama.cpp MIT). Weights under
their base model's licence (MIT / Apache-2.0). Demo corpora: SciDocs (BEIR, CC BY 4.0), WebFAQ-cs (CC BY 4.0). Files derived from
non-commercial bases (jina-embeddings-v5, CC BY-NC) are reported as numbers only.</p>
</main></body></html>