Spaces:
Running
Running
File size: 9,583 Bytes
5abcaf4 15159ad 5abcaf4 15159ad 5abcaf4 15159ad 5abcaf4 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 | <!doctype html>
<html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width, initial-scale=1"><title>thinletter</title>
<link rel="stylesheet" href="style.css"></head><body><main>
<h1>thinletter — a query encoder as a 120 MiB file, in the browser, open source</h1>
<p>We take an embedding model, leave its document index exactly as it is, and compile the <strong>query side</strong> into a file a tenth of the size that
answers in the browser in about 100 ms. Everything behind the numbers is public under Apache-2.0: the quantiser, the container compiler,
the WebGPU runtime, the verification recipe and every released file. Web: <a href="https://thinletter.io">thinletter.io</a> · code and report:
<a href="https://github.com/rosecky/embedding-quantization-public">github.com/rosecky/embedding-quantization-public</a> · demo:
<a href="https://thinletter.io/demo">thinletter.io/demo</a> or the Space <a href="https://huggingface.co/spaces/thinletter/demo">thinletter/demo</a> · contact: info@thinletter.io</p>
<h2>Quickstart: a released client in your page</h2>
<pre><code class="language-js">import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js'; // or: npm install @thinletterio/vqweb
const client = await loadVqwClient('https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw');
const q = await client.embed('Kolik stojí parkování v centru Prahy?'); // Float32Array(1024), L2-normalised: dot it with your Qwen3-Embedding-0.6B index
</code></pre>
<p><code>@thinletterio/vqweb</code> (<a href="https://www.npmjs.com/package/@thinletterio/vqweb">npm</a>, Apache-2.0) is the open WebGPU runtime as a package: download with
progress, a cache in the browser's private file system, tokenizer and prompt from the container header, pipelines, warm-up. Needs WebGPU
with <code>shader-f16</code> (Chrome / Edge), no CPU fallback; about 100 ms per query on an integrated GPU. The scalar GGUF files run with
<a href="https://github.com/ngxson/wllama">wllama</a> or llama.cpp.</p>
<h2>What holds</h2>
<table>
<thead>
<tr>
<th>number</th>
<th>what it is</th>
<th>where to check</th>
</tr>
</thead>
<tbody>
<tr>
<td><strong>98.2 %</strong></td>
<td>of full-precision nDCG@10 from a 120 MiB, 2.1-bit vector-quantised harrier-0.6b client, no retraining: above the trained BitNet-270m (140 MiB, 97.0 %), below any file llama.cpp can produce (178 MiB, 93.9 %). The browser reproduces the simulation (0.7375 vs 0.7380).</td>
<td><a href="https://huggingface.co/thinletter/harrier-0.6b-vq-clients">harrier-0.6b-vq-clients</a>, <a href="https://huggingface.co/datasets/BeIR/scifact">BeIR/scifact</a></td>
</tr>
<tr>
<td><strong>4.3×</strong></td>
<td>faster than llama.cpp's WebGPU path on the same integrated GPU: 103 ms per query for the 119 MiB vector-quantised client against 444 ms for the 235 MiB Q3_K file, in our open WebGPU runtime (<code>client/vqweb</code>).</td>
<td>report §3.4b in the repository</td>
</tr>
<tr>
<td><strong>96.8 %</strong></td>
<td>of fp32 nDCG@10 on the public Czech benchmark WebFAQ-cs from a 237 MiB vector-quantised Qwen3-Embedding-0.6B client (2-d codebooks, 3.6 bits per weight), browser = simulation, 112 ms per query. Czech-calibrated scalar clients: bge-m3 99.2 % (355 MiB), Qwen3 98.7 % (340 MiB) on a Czech legal index.</td>
<td><a href="https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients">qwen3-embedding-0.6b-vq-clients</a>, <a href="https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval">PaDaS-Lab/webfaq-retrieval</a></td>
</tr>
<tr>
<td><strong>+0.010</strong></td>
<td>nDCG@10 at 1.8 bits from a method of ours: ship the query prompt's keys and values in full precision (2 MB) and calibrate the compressed model with them in place; three seeds on SciFact and on a Czech legal index, every paired interval above zero.</td>
<td>report §3.7; the <code>*p</code> files in <a href="https://huggingface.co/thinletter/harrier-0.6b-vq-clients">harrier-0.6b-vq-clients</a></td>
</tr>
<tr>
<td><strong>96–101 %</strong></td>
<td>kept by the 3.4-bit harrier-0.6b scalar client on four English corpora, 235 MiB instead of 1 143 MiB. Not ours to claim: llama.cpp's own quantiser gets there too, and we say so.</td>
<td><a href="https://huggingface.co/thinletter/harrier-0.6b-query-clients">harrier-0.6b-query-clients</a></td>
</tr>
</tbody>
</table>
<h2>Repositories</h2>
<p>Query-side clients: each file is a query encoder for exactly one document encoder and its settings (the model card says which).
The index does not change.</p>
<table>
<thead>
<tr>
<th>repository</th>
<th>base model · licence</th>
<th>files</th>
<th>what holds (test split, the model's own fp32 index)</th>
</tr>
</thead>
<tbody>
<tr>
<td><a href="https://huggingface.co/thinletter/harrier-0.6b-query-clients">harrier-0.6b-query-clients</a></td>
<td><a href="https://huggingface.co/microsoft/harrier-oss-v1-0.6b">microsoft/harrier-oss-v1-0.6b</a> · MIT</td>
<td>3 GGUF, 192–235 MiB</td>
<td>Q3_K generic-English 98.9 / 99.3 / 101 / 99.2 % on <a href="https://huggingface.co/datasets/BeIR/scifact">SciFact</a> / <a href="https://huggingface.co/datasets/BeIR/nfcorpus">NFCorpus</a> / <a href="https://huggingface.co/datasets/BeIR/arguana">ArguAna</a> / <a href="https://huggingface.co/datasets/BeIR/scidocs">SciDocs</a>; Q2_K SciDocs-calibrated 94.5 %</td>
</tr>
<tr>
<td><a href="https://huggingface.co/thinletter/harrier-0.6b-vq-clients">harrier-0.6b-vq-clients</a></td>
<td>microsoft/harrier-oss-v1-0.6b · MIT</td>
<td>8 <code>.vqw</code>, 105–129 MiB, 1.8–2.1 bits per weight</td>
<td>SciDocs 94.7 / 91.3 %, SciFact 98.2 / 96.5 %; prompt-K/V variants SciDocs 96.6 / 93.8 %, SciFact 98.1 / 95.8 %; browser-verified; run in the open WebGPU runtime</td>
</tr>
<tr>
<td><a href="https://huggingface.co/thinletter/qwen3-embedding-0.6b-query-clients">qwen3-embedding-0.6b-query-clients</a></td>
<td><a href="https://huggingface.co/Qwen/Qwen3-Embedding-0.6B">Qwen/Qwen3-Embedding-0.6B</a> · Apache-2.0</td>
<td>4 GGUF, 340–385 MiB, English and Czech calibration</td>
<td>Q4_K_M + 4-bit table: 99.3–100 % on the four English corpora, 98.7 % on a Czech legal index; Q5_K_M 99.5 % Czech</td>
</tr>
<tr>
<td><a href="https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients">qwen3-embedding-0.6b-vq-clients</a></td>
<td>Qwen/Qwen3-Embedding-0.6B · Apache-2.0</td>
<td>1 <code>.vqw</code>, 237 MiB, 2-d codebooks 3.57 bpw + full token table</td>
<td>Czech: 96.8 % of fp32 on <a href="https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval">WebFAQ-cs</a> (cosine 0.946, overlap 0.755), browser = simulation, 112 ms on an integrated GPU</td>
</tr>
<tr>
<td><a href="https://huggingface.co/thinletter/bge-m3-query-clients">bge-m3-query-clients</a></td>
<td><a href="https://huggingface.co/BAAI/bge-m3">BAAI/bge-m3</a> · MIT</td>
<td>2 GGUF, 321–355 MiB, Czech calibration</td>
<td>99.2 % / 98.3 % on the Czech legal index, 99.6–100 % on SciFact</td>
</tr>
<tr>
<td><a href="https://huggingface.co/spaces/thinletter/demo">demo</a> (Space)</td>
<td>—</td>
<td>the browser demo, static</td>
<td>search SciDocs (25 656 abstracts) or WebFAQ-cs (71 529 Czech passages) locally with a scalar or a vector-quantised client; models and indexes download from thinletter.io</td>
</tr>
<tr>
<td><a href="https://huggingface.co/datasets/thinletter/llm-weight-compression-evidence">llm-weight-compression-evidence</a></td>
<td>dataset · MIT</td>
<td>per-window evidence</td>
<td>which objective a post-training quantiser should optimise (companion project)</td>
</tr>
</tbody>
</table>
<p>Each model repository has a summary post in its Community tab with the numbers, what was found along the way, and the limits. The <code>.vqw</code> files load with <code>@thinletterio/vqweb</code> (above); the GGUF files with llama.cpp / wllama.</p>
<h2>How we measure</h2>
<p>Test split, the model's own fp32 index, native runtime (llama.cpp) or the browser. Every run is pre-registered with a prediction and a
kill rule; comparisons change one thing at a time at the same calibration budget; differences are paired bootstrap intervals over queries
(10 000 draws) read against the variance of the calibration draw (~0.01 nDCG@10) and between machines (±0.006): under 0.01 is a tie.
Negative results stay published (a codebook-budget idea that did not work, a 2.6-bit headline that held on one corpus only, the 2-bit cliff
of Qwen3-Embedding: 65–81 % for the same run on two machines).</p>
<h2>Verify on your own index</h2>
<p>The <a href="https://github.com/rosecky/embedding-quantization-public">public repository</a> has the recipe: import your corpus and existing
document vectors, generate synthetic queries, quantise (llama-quantize, the GPTQ exporter, or the vector quantiser + <code>.vqw</code> compiler),
evaluate the client against the unchanged index with a paired interval. Everything in the report is reproducible from it; only the raw
per-run results, the research log and a customer's data stay out.</p>
<h2>Licences</h2>
<p>Code Apache-2.0 (NOTICE lists wllama MIT, @huggingface/tokenizers Apache-2.0, a Q2_K decoder ported from llama.cpp MIT). Weights under
their base model's licence (MIT / Apache-2.0). Demo corpora: SciDocs (BEIR, CC BY 4.0), WebFAQ-cs (CC BY 4.0). Files derived from
non-commercial bases (jina-embeddings-v5, CC BY-NC) are reported as numbers only.</p>
</main></body></html>
|