hotchpotch commited on
Commit
cf92c2f
·
verified ·
1 Parent(s): 2c3f04e

Add browser ONNX export with INT8 embeddings and usage guide

Browse files
onnx_browser/README.md ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Bekko System One v0 17M — browser ONNX
2
+
3
+ This directory contains an ONNX conversion for the dedicated **Bekko System One browser runtime**. It supports Choice, Noul, and Score typed decisions. The companion Node.js runtime can also load this directory.
4
+
5
+ This is not a Sentence Transformers ONNX backend export or a general-purpose sentence embedding model. Use the model repository's root files and `inference_v0.py` for Python / Sentence Transformers inference.
6
+
7
+ ## Files and precision
8
+
9
+ | File | Purpose |
10
+ | --- | --- |
11
+ | `model.onnx` | Shared-prefix encoder and three prediction heads |
12
+ | `manifest.json` | Model filename, hashes, task order, token budgets, quantization policy, and source revision |
13
+ | `tokenizer.json` | Tokenizer vocabulary and processing rules |
14
+ | `tokenizer_config.json` | Special tokens used by the dedicated runtime |
15
+
16
+ Only the token embedding table is stored as row-wise symmetric INT8. Transformer blocks and prediction heads remain FP32; this is not fully INT8 inference. The tokenizer is unchanged. The ONNX file is 29.0 MB (decimal), excluding tokenizer and runtime downloads.
17
+
18
+ ## Loading
19
+
20
+ Use the [Bekko browser runtime and usage guide](https://github.com/hotchpotch/bekko-system-one/tree/main/browser). Set the model base URL to:
21
+
22
+ ```text
23
+ https://huggingface.co/hotchpotch/bekko-system-one-v0-17m/resolve/<COMMIT_REVISION>/onnx_browser/
24
+ ```
25
+
26
+ Replace `<COMMIT_REVISION>` with a commit containing this directory and pin that same revision for all files. The runtime reads `manifest.json`, both tokenizer files, and the model named in the manifest. It handles typed-decision input rendering, candidate grouping, and output probabilities. Loading `model.onnx` alone does not provide those application-level operations.
27
+
28
+ Browser execution uses ONNX Runtime Web. Available execution providers, memory use, and practical input sizes depend on the browser and device. Download size is not peak runtime memory; larger models may exceed a device's memory budget. Node.js usage is documented in the same guide.
29
+
30
+ ## Conversion and verification
31
+
32
+ - [FP32 ONNX export script](https://github.com/hotchpotch/bekko-system-one/blob/main/browser/scripts/export_onnx.py)
33
+ - [Embedding INT8 conversion script](https://github.com/hotchpotch/bekko-system-one/blob/main/browser/scripts/quantize_onnx.py) (`--embedding-only`)
34
+ - [ONNX verification script](https://github.com/hotchpotch/bekko-system-one/blob/main/browser/scripts/verify-export.js)
35
+
36
+ The INT8 output is distributed here as `model.onnx`; the manifest records its quantization policy. FP32 reference files and parity fixtures are retained separately for verification.
37
+
38
+ Source model revision: `2c3f04e55e8d0764620ee46fc86be33e408d1ebd`. The export checkpoint weights and tokenizer were checked against this revision. On 5 synthetic fixtures, Node.js / ONNX Runtime verification measured a maximum FP32 logit error of 3.3378601e-06 against the Python reference, and a maximum absolute probability difference of 0.0043458041 between FP32 and embedding-INT8 ONNX. These checks are not a full quality evaluation or a guarantee of browser/device compatibility.
onnx_browser/manifest.json ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "format_version": 1,
3
+ "tasks": [
4
+ "choice",
5
+ "noul",
6
+ "score"
7
+ ],
8
+ "query_length": 4096,
9
+ "document_length": 2048,
10
+ "cls_token_id": 50281,
11
+ "sep_token_id": 50282,
12
+ "pad_token_id": 50283,
13
+ "model_sha256": "1a27d9ce0f44ddb8e3247aec245e3c05e2425996d5ff0def114a093c58c464de",
14
+ "heads_sha256": "03aaa7864b6fd5c08c5ca57f233b65f907226ad43579c91f0998f5eaf51f7aac",
15
+ "weights_sha256": "e4c2587853fa9073b022c18801857733abf2950bb765fe37e434b6baaae18ac7",
16
+ "model_file": "model.onnx",
17
+ "model_bytes": 29023607,
18
+ "source_model_sha256": "b3b55bba127f6d3eca5c3f4762e04b77b1bef301e3242132a58fca5e93813b52",
19
+ "quantization": {
20
+ "embedding": "rowwise_symmetric_int8",
21
+ "blocks": "fp32",
22
+ "heads": "fp32",
23
+ "integer_matmuls": 0
24
+ },
25
+ "source_repository": "hotchpotch/bekko-system-one-v0-17m",
26
+ "source_revision": "2c3f04e55e8d0764620ee46fc86be33e408d1ebd"
27
+ }
onnx_browser/model.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1a27d9ce0f44ddb8e3247aec245e3c05e2425996d5ff0def114a093c58c464de
3
+ size 29023607
onnx_browser/tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
onnx_browser/tokenizer_config.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "cls_token": "[CLS]",
3
+ "sep_token": "[SEP]",
4
+ "pad_token": "[PAD]",
5
+ "unk_token": "[UNK]",
6
+ "mask_token": "[MASK]"
7
+ }