honza-rosecky commited on
Commit
15159ad
路
verified 路
1 Parent(s): 0b8e31b

org card: quickstart with @thinletterio/vqweb

Browse files
Files changed (1) hide show
  1. index.html +3 -3
index.html CHANGED
@@ -8,12 +8,12 @@ the WebGPU runtime, the verification recipe and every released file. Web: <a hre
8
  <a href="https://github.com/rosecky/embedding-quantization-public">github.com/rosecky/embedding-quantization-public</a> 路 demo:
9
  <a href="https://thinletter.io/demo">thinletter.io/demo</a> or the Space <a href="https://huggingface.co/spaces/thinletter/demo">thinletter/demo</a> 路 contact: info@thinletter.io</p>
10
  <h2>Quickstart: a released client in your page</h2>
11
- <pre><code class="language-js">import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js'; // or: npm install @thinletter/vqweb
12
 
13
  const client = await loadVqwClient('https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw');
14
  const q = await client.embed('Kolik stoj铆 parkov谩n铆 v centru Prahy?'); // Float32Array(1024), L2-normalised: dot it with your Qwen3-Embedding-0.6B index
15
  </code></pre>
16
- <p><code>@thinletter/vqweb</code> (<a href="https://www.npmjs.com/package/@thinletter/vqweb">npm</a>, Apache-2.0) is the open WebGPU runtime as a package: download with
17
  progress, a cache in the browser's private file system, tokenizer and prompt from the container header, pipelines, warm-up. Needs WebGPU
18
  with <code>shader-f16</code> (Chrome / Edge), no CPU fallback; about 100 ms per query on an integrated GPU. The scalar GGUF files run with
19
  <a href="https://github.com/ngxson/wllama">wllama</a> or llama.cpp.</p>
@@ -111,7 +111,7 @@ The index does not change.</p>
111
  </tr>
112
  </tbody>
113
  </table>
114
- <p>Each model repository has a summary post in its Community tab with the numbers, what was found along the way, and the limits. The <code>.vqw</code> files load with <code>@thinletter/vqweb</code> (above); the GGUF files with llama.cpp / wllama.</p>
115
  <h2>How we measure</h2>
116
  <p>Test split, the model's own fp32 index, native runtime (llama.cpp) or the browser. Every run is pre-registered with a prediction and a
117
  kill rule; comparisons change one thing at a time at the same calibration budget; differences are paired bootstrap intervals over queries
 
8
  <a href="https://github.com/rosecky/embedding-quantization-public">github.com/rosecky/embedding-quantization-public</a> 路 demo:
9
  <a href="https://thinletter.io/demo">thinletter.io/demo</a> or the Space <a href="https://huggingface.co/spaces/thinletter/demo">thinletter/demo</a> 路 contact: info@thinletter.io</p>
10
  <h2>Quickstart: a released client in your page</h2>
11
+ <pre><code class="language-js">import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js'; // or: npm install @thinletterio/vqweb
12
 
13
  const client = await loadVqwClient('https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw');
14
  const q = await client.embed('Kolik stoj铆 parkov谩n铆 v centru Prahy?'); // Float32Array(1024), L2-normalised: dot it with your Qwen3-Embedding-0.6B index
15
  </code></pre>
16
+ <p><code>@thinletterio/vqweb</code> (<a href="https://www.npmjs.com/package/@thinletterio/vqweb">npm</a>, Apache-2.0) is the open WebGPU runtime as a package: download with
17
  progress, a cache in the browser's private file system, tokenizer and prompt from the container header, pipelines, warm-up. Needs WebGPU
18
  with <code>shader-f16</code> (Chrome / Edge), no CPU fallback; about 100 ms per query on an integrated GPU. The scalar GGUF files run with
19
  <a href="https://github.com/ngxson/wllama">wllama</a> or llama.cpp.</p>
 
111
  </tr>
112
  </tbody>
113
  </table>
114
+ <p>Each model repository has a summary post in its Community tab with the numbers, what was found along the way, and the limits. The <code>.vqw</code> files load with <code>@thinletterio/vqweb</code> (above); the GGUF files with llama.cpp / wllama.</p>
115
  <h2>How we measure</h2>
116
  <p>Test split, the model's own fp32 index, native runtime (llama.cpp) or the browser. Every run is pre-registered with a prediction and a
117
  kill rule; comparisons change one thing at a time at the same calibration budget; differences are paired bootstrap intervals over queries