honza-rosecky commited on
Commit
0b8e31b
·
verified ·
1 Parent(s): 31afe74

org card: quickstart with @thinletterio/vqweb

Browse files
Files changed (1) hide show
  1. README.md +3 -3
README.md CHANGED
@@ -18,13 +18,13 @@ the WebGPU runtime, the verification recipe and every released file. Web: [thinl
18
  ## Quickstart: a released client in your page
19
 
20
  ```js
21
- import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js'; // or: npm install @thinletter/vqweb
22
 
23
  const client = await loadVqwClient('https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw');
24
  const q = await client.embed('Kolik stojí parkování v centru Prahy?'); // Float32Array(1024), L2-normalised: dot it with your Qwen3-Embedding-0.6B index
25
  ```
26
 
27
- `@thinletter/vqweb` ([npm](https://www.npmjs.com/package/@thinletter/vqweb), Apache-2.0) is the open WebGPU runtime as a package: download with
28
  progress, a cache in the browser's private file system, tokenizer and prompt from the container header, pipelines, warm-up. Needs WebGPU
29
  with `shader-f16` (Chrome / Edge), no CPU fallback; about 100 ms per query on an integrated GPU. The scalar GGUF files run with
30
  [wllama](https://github.com/ngxson/wllama) or llama.cpp.
@@ -54,7 +54,7 @@ The index does not change.
54
  | [demo](https://huggingface.co/spaces/thinletter/demo) (Space) | — | the browser demo, static | search SciDocs (25 656 abstracts) or WebFAQ-cs (71 529 Czech passages) locally with a scalar or a vector-quantised client; models and indexes download from thinletter.io |
55
  | [llm-weight-compression-evidence](https://huggingface.co/datasets/thinletter/llm-weight-compression-evidence) | dataset · MIT | per-window evidence | which objective a post-training quantiser should optimise (companion project) |
56
 
57
- Each model repository has a summary post in its Community tab with the numbers, what was found along the way, and the limits. The `.vqw` files load with `@thinletter/vqweb` (above); the GGUF files with llama.cpp / wllama.
58
 
59
  ## How we measure
60
 
 
18
  ## Quickstart: a released client in your page
19
 
20
  ```js
21
+ import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js'; // or: npm install @thinletterio/vqweb
22
 
23
  const client = await loadVqwClient('https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw');
24
  const q = await client.embed('Kolik stojí parkování v centru Prahy?'); // Float32Array(1024), L2-normalised: dot it with your Qwen3-Embedding-0.6B index
25
  ```
26
 
27
+ `@thinletterio/vqweb` ([npm](https://www.npmjs.com/package/@thinletterio/vqweb), Apache-2.0) is the open WebGPU runtime as a package: download with
28
  progress, a cache in the browser's private file system, tokenizer and prompt from the container header, pipelines, warm-up. Needs WebGPU
29
  with `shader-f16` (Chrome / Edge), no CPU fallback; about 100 ms per query on an integrated GPU. The scalar GGUF files run with
30
  [wllama](https://github.com/ngxson/wllama) or llama.cpp.
 
54
  | [demo](https://huggingface.co/spaces/thinletter/demo) (Space) | — | the browser demo, static | search SciDocs (25 656 abstracts) or WebFAQ-cs (71 529 Czech passages) locally with a scalar or a vector-quantised client; models and indexes download from thinletter.io |
55
  | [llm-weight-compression-evidence](https://huggingface.co/datasets/thinletter/llm-weight-compression-evidence) | dataset · MIT | per-window evidence | which objective a post-training quantiser should optimise (companion project) |
56
 
57
+ Each model repository has a summary post in its Community tab with the numbers, what was found along the way, and the limits. The `.vqw` files load with `@thinletterio/vqweb` (above); the GGUF files with llama.cpp / wllama.
58
 
59
  ## How we measure
60