Qwen3-Embedding-0.6B vector-quantised query client (Czech, 3.6 bits per weight, WebGPU)

One file, 237 MiB, that encodes Czech queries for an unchanged Qwen3-Embedding-0.6B index and runs in the browser on our open WebGPU runtime. On the public Czech benchmark WebFAQ-cs (71 529 passages, 7 231 human-judged questions) it keeps 96.8 % of the full-precision nDCG@10 (0.7072 vs 0.7308), cosine to the fp32 query vector 0.946, top-10 overlap 0.755. The released scalar client of the same model (llama-quantize Q4_K_M + 4-bit token table, 340 MiB) keeps 98.7 % on a Czech legal index; this file is 70 % of its size.

file MiB bits/weight (blocks) WebFAQ-cs nDCG@10 (% of fp32 0.7308) / cosine / top-10 overlap Czech legal index (% of fp32 0.3186) browser check
qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw 237.0 3.57 (2-d codebooks, K = 128, 7-bit indices) + Q2_K token table (full vocabulary, 49 MiB) 0.7072 = 96.8 % / 0.946 / 0.755 (torch simulation of the file, test split) 0.2965 = 93.1 % / 0.930 / 0.653 (1 137 synthetic queries) WebGPU, integrated Intel GPU: 0.6393 vs 0.6393 in simulation on the first 50 test queries (Δ 0.0000), p50 112 ms / p95 130 ms per query, 2.5 s load, 412 MiB tab memory

Calibration: 3 000 real questions from the WebFAQ-cs train split (disjoint from the test split) in the model's official query prompt, one rotation seed. The Czech legal number is a saturated synthetic test with calibration from another domain; the public WebFAQ-cs number decides. Everything is one seed: differences under 0.01 nDCG@10 are inside the calibration-draw variance. sha256 of the file: a8de0ba4ae92… (full value in the repository's results/raw/vqcz_box/exports_table.jsonl).

Why two-dimensional codebooks

The released harrier containers use 4-d codebooks (256 entries, 2.1 bits per weight). Qwen3-Embedding-0.6B is an order of magnitude more fragile under compression: at 2.1 bits it collapses on Czech (65–81 % of fp32, the same run on two machines), and every llama.cpp file at or below 3 bits fails too. The probe of 2026-09-15/16 (docs/release/vq_results.md §7 in the repository) found that 2-d codebooks beat 4-d ones per bit: 2-d K = 64 at 3.07 bpw ≥ 4-d K = 2048 at 3.15 bpw, +0.024 nDCG@10 / +0.06 cosine over 4-d K = 1024. 2-d K = 128 at 3.57 bpw keeps 98.0 % of fp32 with generic Czech calibration and 98.8 % with question-shaped calibration on the blocks alone; the Q2_K token table of the runtime's current format costs 2.0 points and 0.025 cosine (a trimmed 35k-token table costs 0.8 points more and fails the cosine criterion, so the full table ships). A 4-bit token table is the next format step.

Quickstart in the browser (@thinletterio/vqweb)

import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js';   // or: npm install @thinletterio/vqweb

const client = await loadVqwClient('https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw');
const q = await client.embed('Kolik stojí parkování v centru Prahy?');   // Float32Array(1024), L2-normalised: dot it with your Qwen/Qwen3-Embedding-0.6B index

The package (Apache-2.0, npm, source) downloads the file with progress, keeps it in the browser's private file system for the next visit, uploads it to the GPU, takes the tokenizer and the prompt from the container header and compiles the WebGPU pipelines; client.embed(text) takes the raw query. Needs WebGPU with shader-f16 (Chrome / Edge; probeWebGPU() says why not), no CPU fallback. About 100 ms per query on an integrated GPU, 2–3 s to load after the download.

How to run it

  • In the browser: https://thinletter.io/demo, corpus "WebFAQ CS", client vq3.5 (WebGPU with shader-f16; no CPU fallback), or the mirror at https://huggingface.co/spaces/thinletter/demo.
  • Locally: client/browser/serve.py --mount vqw=<dir> and client/vqweb/bench.html?vqw=/vqw/<file>.vqw&dataset=webfaq-cs&split=test, or scripts/browser_run.py --runtime vqweb --vqw <file>; the torch reference scripts/vq_reference.py <file> --embed "text".
  • The runtime (client/vqweb/), the quantiser and the container compiler are Apache-2.0 in https://github.com/rosecky/embedding-quantization-public; the file was produced with scripts/vq_export_web.py --teacher qwen3-0.6b --dim 2 --bits 3.5 --calib webfaq-cs_train_q --n_seq 3000 --vocab full.

Compatibility

Query encoder for an index built with Qwen/Qwen3-Embedding-0.6B: documents without instruction, queries with the official Sentence-Transformers prompt Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery: (no trailing space), last-token pooling, L2-normalised 1024-d vectors. tokenizer-qwen3-0.6b.json / tokenizer_config-qwen3-0.6b.json in this repository are the files the container header references.

Limits

One model, one seed; the browser reproduces the simulation on 50 queries, the full 7 231-query number is the simulation of the file; the legal-index number is below the release criterion and is reported as such. Discrete-GPU and mobile latencies are not measured. The file is 3.6 bits per weight, not the 2 bits of the harrier containers: this model does not survive 2 bits.

Licence

Weights derived from Qwen/Qwen3-Embedding-0.6B (Apache-2.0), calibration on WebFAQ (PaDaS-Lab/webfaq-retrieval, CC BY 4.0) — the questions are not redistributed. Container format, compiler and runtime: Apache-2.0 (repository LICENSE / NOTICE). Questions: info@thinletter.io, or the issues of the public repository.

Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thinletter/qwen3-embedding-0.6b-vq-clients

Finetuned
(285)
this model