Spaces:
Running
Running
org card: quickstart with @thinletterio/vqweb
Browse files
README.md
CHANGED
|
@@ -18,13 +18,13 @@ the WebGPU runtime, the verification recipe and every released file. Web: [thinl
|
|
| 18 |
## Quickstart: a released client in your page
|
| 19 |
|
| 20 |
```js
|
| 21 |
-
import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js'; // or: npm install @
|
| 22 |
|
| 23 |
const client = await loadVqwClient('https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw');
|
| 24 |
const q = await client.embed('Kolik stojí parkování v centru Prahy?'); // Float32Array(1024), L2-normalised: dot it with your Qwen3-Embedding-0.6B index
|
| 25 |
```
|
| 26 |
|
| 27 |
-
`@
|
| 28 |
progress, a cache in the browser's private file system, tokenizer and prompt from the container header, pipelines, warm-up. Needs WebGPU
|
| 29 |
with `shader-f16` (Chrome / Edge), no CPU fallback; about 100 ms per query on an integrated GPU. The scalar GGUF files run with
|
| 30 |
[wllama](https://github.com/ngxson/wllama) or llama.cpp.
|
|
@@ -54,7 +54,7 @@ The index does not change.
|
|
| 54 |
| [demo](https://huggingface.co/spaces/thinletter/demo) (Space) | — | the browser demo, static | search SciDocs (25 656 abstracts) or WebFAQ-cs (71 529 Czech passages) locally with a scalar or a vector-quantised client; models and indexes download from thinletter.io |
|
| 55 |
| [llm-weight-compression-evidence](https://huggingface.co/datasets/thinletter/llm-weight-compression-evidence) | dataset · MIT | per-window evidence | which objective a post-training quantiser should optimise (companion project) |
|
| 56 |
|
| 57 |
-
Each model repository has a summary post in its Community tab with the numbers, what was found along the way, and the limits. The `.vqw` files load with `@
|
| 58 |
|
| 59 |
## How we measure
|
| 60 |
|
|
|
|
| 18 |
## Quickstart: a released client in your page
|
| 19 |
|
| 20 |
```js
|
| 21 |
+
import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js'; // or: npm install @thinletterio/vqweb
|
| 22 |
|
| 23 |
const client = await loadVqwClient('https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw');
|
| 24 |
const q = await client.embed('Kolik stojí parkování v centru Prahy?'); // Float32Array(1024), L2-normalised: dot it with your Qwen3-Embedding-0.6B index
|
| 25 |
```
|
| 26 |
|
| 27 |
+
`@thinletterio/vqweb` ([npm](https://www.npmjs.com/package/@thinletterio/vqweb), Apache-2.0) is the open WebGPU runtime as a package: download with
|
| 28 |
progress, a cache in the browser's private file system, tokenizer and prompt from the container header, pipelines, warm-up. Needs WebGPU
|
| 29 |
with `shader-f16` (Chrome / Edge), no CPU fallback; about 100 ms per query on an integrated GPU. The scalar GGUF files run with
|
| 30 |
[wllama](https://github.com/ngxson/wllama) or llama.cpp.
|
|
|
|
| 54 |
| [demo](https://huggingface.co/spaces/thinletter/demo) (Space) | — | the browser demo, static | search SciDocs (25 656 abstracts) or WebFAQ-cs (71 529 Czech passages) locally with a scalar or a vector-quantised client; models and indexes download from thinletter.io |
|
| 55 |
| [llm-weight-compression-evidence](https://huggingface.co/datasets/thinletter/llm-weight-compression-evidence) | dataset · MIT | per-window evidence | which objective a post-training quantiser should optimise (companion project) |
|
| 56 |
|
| 57 |
+
Each model repository has a summary post in its Community tab with the numbers, what was found along the way, and the limits. The `.vqw` files load with `@thinletterio/vqweb` (above); the GGUF files with llama.cpp / wllama.
|
| 58 |
|
| 59 |
## How we measure
|
| 60 |
|