sentence-transformers
ONNX
Safetensors
English
modernbert
typed-decisions
classification
scoring
custom-code
Instructions to use hotchpotch/bekko-system-one-v0-17m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use hotchpotch/bekko-system-one-v0-17m with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("hotchpotch/bekko-system-one-v0-17m") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
Add browser ONNX export with INT8 embeddings and usage guide
Browse files- onnx_browser/README.md +38 -0
- onnx_browser/manifest.json +27 -0
- onnx_browser/model.onnx +3 -0
- onnx_browser/tokenizer.json +0 -0
- onnx_browser/tokenizer_config.json +7 -0
onnx_browser/README.md
ADDED
|
@@ -0,0 +1,38 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Bekko System One v0 17M — browser ONNX
|
| 2 |
+
|
| 3 |
+
This directory contains an ONNX conversion for the dedicated **Bekko System One browser runtime**. It supports Choice, Noul, and Score typed decisions. The companion Node.js runtime can also load this directory.
|
| 4 |
+
|
| 5 |
+
This is not a Sentence Transformers ONNX backend export or a general-purpose sentence embedding model. Use the model repository's root files and `inference_v0.py` for Python / Sentence Transformers inference.
|
| 6 |
+
|
| 7 |
+
## Files and precision
|
| 8 |
+
|
| 9 |
+
| File | Purpose |
|
| 10 |
+
| --- | --- |
|
| 11 |
+
| `model.onnx` | Shared-prefix encoder and three prediction heads |
|
| 12 |
+
| `manifest.json` | Model filename, hashes, task order, token budgets, quantization policy, and source revision |
|
| 13 |
+
| `tokenizer.json` | Tokenizer vocabulary and processing rules |
|
| 14 |
+
| `tokenizer_config.json` | Special tokens used by the dedicated runtime |
|
| 15 |
+
|
| 16 |
+
Only the token embedding table is stored as row-wise symmetric INT8. Transformer blocks and prediction heads remain FP32; this is not fully INT8 inference. The tokenizer is unchanged. The ONNX file is 29.0 MB (decimal), excluding tokenizer and runtime downloads.
|
| 17 |
+
|
| 18 |
+
## Loading
|
| 19 |
+
|
| 20 |
+
Use the [Bekko browser runtime and usage guide](https://github.com/hotchpotch/bekko-system-one/tree/main/browser). Set the model base URL to:
|
| 21 |
+
|
| 22 |
+
```text
|
| 23 |
+
https://huggingface.co/hotchpotch/bekko-system-one-v0-17m/resolve/<COMMIT_REVISION>/onnx_browser/
|
| 24 |
+
```
|
| 25 |
+
|
| 26 |
+
Replace `<COMMIT_REVISION>` with a commit containing this directory and pin that same revision for all files. The runtime reads `manifest.json`, both tokenizer files, and the model named in the manifest. It handles typed-decision input rendering, candidate grouping, and output probabilities. Loading `model.onnx` alone does not provide those application-level operations.
|
| 27 |
+
|
| 28 |
+
Browser execution uses ONNX Runtime Web. Available execution providers, memory use, and practical input sizes depend on the browser and device. Download size is not peak runtime memory; larger models may exceed a device's memory budget. Node.js usage is documented in the same guide.
|
| 29 |
+
|
| 30 |
+
## Conversion and verification
|
| 31 |
+
|
| 32 |
+
- [FP32 ONNX export script](https://github.com/hotchpotch/bekko-system-one/blob/main/browser/scripts/export_onnx.py)
|
| 33 |
+
- [Embedding INT8 conversion script](https://github.com/hotchpotch/bekko-system-one/blob/main/browser/scripts/quantize_onnx.py) (`--embedding-only`)
|
| 34 |
+
- [ONNX verification script](https://github.com/hotchpotch/bekko-system-one/blob/main/browser/scripts/verify-export.js)
|
| 35 |
+
|
| 36 |
+
The INT8 output is distributed here as `model.onnx`; the manifest records its quantization policy. FP32 reference files and parity fixtures are retained separately for verification.
|
| 37 |
+
|
| 38 |
+
Source model revision: `2c3f04e55e8d0764620ee46fc86be33e408d1ebd`. The export checkpoint weights and tokenizer were checked against this revision. On 5 synthetic fixtures, Node.js / ONNX Runtime verification measured a maximum FP32 logit error of 3.3378601e-06 against the Python reference, and a maximum absolute probability difference of 0.0043458041 between FP32 and embedding-INT8 ONNX. These checks are not a full quality evaluation or a guarantee of browser/device compatibility.
|
onnx_browser/manifest.json
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"format_version": 1,
|
| 3 |
+
"tasks": [
|
| 4 |
+
"choice",
|
| 5 |
+
"noul",
|
| 6 |
+
"score"
|
| 7 |
+
],
|
| 8 |
+
"query_length": 4096,
|
| 9 |
+
"document_length": 2048,
|
| 10 |
+
"cls_token_id": 50281,
|
| 11 |
+
"sep_token_id": 50282,
|
| 12 |
+
"pad_token_id": 50283,
|
| 13 |
+
"model_sha256": "1a27d9ce0f44ddb8e3247aec245e3c05e2425996d5ff0def114a093c58c464de",
|
| 14 |
+
"heads_sha256": "03aaa7864b6fd5c08c5ca57f233b65f907226ad43579c91f0998f5eaf51f7aac",
|
| 15 |
+
"weights_sha256": "e4c2587853fa9073b022c18801857733abf2950bb765fe37e434b6baaae18ac7",
|
| 16 |
+
"model_file": "model.onnx",
|
| 17 |
+
"model_bytes": 29023607,
|
| 18 |
+
"source_model_sha256": "b3b55bba127f6d3eca5c3f4762e04b77b1bef301e3242132a58fca5e93813b52",
|
| 19 |
+
"quantization": {
|
| 20 |
+
"embedding": "rowwise_symmetric_int8",
|
| 21 |
+
"blocks": "fp32",
|
| 22 |
+
"heads": "fp32",
|
| 23 |
+
"integer_matmuls": 0
|
| 24 |
+
},
|
| 25 |
+
"source_repository": "hotchpotch/bekko-system-one-v0-17m",
|
| 26 |
+
"source_revision": "2c3f04e55e8d0764620ee46fc86be33e408d1ebd"
|
| 27 |
+
}
|
onnx_browser/model.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:1a27d9ce0f44ddb8e3247aec245e3c05e2425996d5ff0def114a093c58c464de
|
| 3 |
+
size 29023607
|
onnx_browser/tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
onnx_browser/tokenizer_config.json
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"cls_token": "[CLS]",
|
| 3 |
+
"sep_token": "[SEP]",
|
| 4 |
+
"pad_token": "[PAD]",
|
| 5 |
+
"unk_token": "[UNK]",
|
| 6 |
+
"mask_token": "[MASK]"
|
| 7 |
+
}
|