Qwen3-Embedding-0.6B vector-quantised query client (Czech, 3.6 bits per weight, WebGPU)
One file, 237 MiB, that encodes Czech queries for an unchanged Qwen3-Embedding-0.6B index and runs in the browser on our open
WebGPU runtime. On the public Czech benchmark WebFAQ-cs (71 529 passages, 7 231 human-judged questions) it keeps 96.8 % of the
full-precision nDCG@10 (0.7072 vs 0.7308), cosine to the fp32 query vector 0.946, top-10 overlap 0.755. The released scalar client of
the same model (llama-quantize Q4_K_M + 4-bit token table, 340 MiB) keeps 98.7 % on a Czech legal index; this file is 70 % of its size.
| file | MiB | bits/weight (blocks) | WebFAQ-cs nDCG@10 (% of fp32 0.7308) / cosine / top-10 overlap | Czech legal index (% of fp32 0.3186) | browser check |
|---|---|---|---|---|---|
qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw |
237.0 | 3.57 (2-d codebooks, K = 128, 7-bit indices) + Q2_K token table (full vocabulary, 49 MiB) | 0.7072 = 96.8 % / 0.946 / 0.755 (torch simulation of the file, test split) | 0.2965 = 93.1 % / 0.930 / 0.653 (1 137 synthetic queries) | WebGPU, integrated Intel GPU: 0.6393 vs 0.6393 in simulation on the first 50 test queries (Δ 0.0000), p50 112 ms / p95 130 ms per query, 2.5 s load, 412 MiB tab memory |
Calibration: 3 000 real questions from the WebFAQ-cs train split (disjoint from the test split) in the model's official query prompt,
one rotation seed. The Czech legal number is a saturated synthetic test with calibration from another domain; the public WebFAQ-cs
number decides. Everything is one seed: differences under 0.01 nDCG@10 are inside the calibration-draw variance. sha256 of the file: a8de0ba4ae92… (full value in the repository's results/raw/vqcz_box/exports_table.jsonl).
Why two-dimensional codebooks
The released harrier containers use 4-d codebooks (256 entries, 2.1 bits per weight). Qwen3-Embedding-0.6B is an order of magnitude
more fragile under compression: at 2.1 bits it collapses on Czech (65–81 % of fp32, the same run on two machines), and every
llama.cpp file at or below 3 bits fails too. The probe of 2026-09-15/16 (docs/release/vq_results.md §7 in the repository) found that
2-d codebooks beat 4-d ones per bit: 2-d K = 64 at 3.07 bpw ≥ 4-d K = 2048 at 3.15 bpw, +0.024 nDCG@10 / +0.06 cosine over 4-d K = 1024.
2-d K = 128 at 3.57 bpw keeps 98.0 % of fp32 with generic Czech calibration and 98.8 % with question-shaped calibration on the blocks
alone; the Q2_K token table of the runtime's current format costs 2.0 points and 0.025 cosine (a trimmed 35k-token table costs 0.8
points more and fails the cosine criterion, so the full table ships). A 4-bit token table is the next format step.
Quickstart in the browser (@thinletterio/vqweb)
import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js'; // or: npm install @thinletterio/vqweb
const client = await loadVqwClient('https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw');
const q = await client.embed('Kolik stojí parkování v centru Prahy?'); // Float32Array(1024), L2-normalised: dot it with your Qwen/Qwen3-Embedding-0.6B index
The package (Apache-2.0, npm, source)
downloads the file with progress, keeps it in the browser's private file system for the next visit, uploads it to the GPU, takes the
tokenizer and the prompt from the container header and compiles the WebGPU pipelines; client.embed(text) takes the raw query.
Needs WebGPU with shader-f16 (Chrome / Edge; probeWebGPU() says why not), no CPU fallback. About 100 ms per query on an
integrated GPU, 2–3 s to load after the download.
How to run it
- In the browser: https://thinletter.io/demo, corpus "WebFAQ CS", client
vq3.5(WebGPU with shader-f16; no CPU fallback), or the mirror at https://huggingface.co/spaces/thinletter/demo. - Locally:
client/browser/serve.py --mount vqw=<dir>andclient/vqweb/bench.html?vqw=/vqw/<file>.vqw&dataset=webfaq-cs&split=test, orscripts/browser_run.py --runtime vqweb --vqw <file>; the torch referencescripts/vq_reference.py <file> --embed "text". - The runtime (
client/vqweb/), the quantiser and the container compiler are Apache-2.0 in https://github.com/rosecky/embedding-quantization-public; the file was produced withscripts/vq_export_web.py --teacher qwen3-0.6b --dim 2 --bits 3.5 --calib webfaq-cs_train_q --n_seq 3000 --vocab full.
Compatibility
Query encoder for an index built with Qwen/Qwen3-Embedding-0.6B: documents without instruction, queries with the official
Sentence-Transformers prompt Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery: (no trailing
space), last-token pooling, L2-normalised 1024-d vectors. tokenizer-qwen3-0.6b.json / tokenizer_config-qwen3-0.6b.json in this repository are the files the
container header references.
Limits
One model, one seed; the browser reproduces the simulation on 50 queries, the full 7 231-query number is the simulation of the file; the legal-index number is below the release criterion and is reported as such. Discrete-GPU and mobile latencies are not measured. The file is 3.6 bits per weight, not the 2 bits of the harrier containers: this model does not survive 2 bits.
Licence
Weights derived from Qwen/Qwen3-Embedding-0.6B (Apache-2.0), calibration on WebFAQ (PaDaS-Lab/webfaq-retrieval, CC BY 4.0) — the questions are not redistributed. Container format, compiler and runtime: Apache-2.0 (repository LICENSE / NOTICE). Questions: info@thinletter.io, or the issues of the public repository.
- Downloads last month
- 5