sat-3l-sm-onnx
Weight-only 8-bit ONNX export of SaT 3L-SM (segment-any-text/sat-3l-sm), wtpsplit's
3-layer XLM-RoBERTa sentence-boundary model, with an 8-bit GatherBlockQuantized word-embedding
table (the "q8w-gather" build), for use by
Sokuji's local sentence-segmentation stage.
Upstream
- Weights: segment-any-text/sat-3l-sm β MIT.
- Code: segment-any-text/wtpsplit 2.2.1 β MIT
(
LICENSEin this repo is the verbatim upstream text). - Architecture: a 3-layer XLM-RoBERTa encoder with a single output label ("a sentence ends after this subword"), trained to work without punctuation or casing. It inserts no punctuation marks itself β it only predicts sentence boundaries. Covers 85 languages upstream.
Files
| file | bytes | sha256 |
|---|---|---|
model.onnx |
241,945,842 | 1132546a16fa1aa9da218602049c6e7aeba0ec5b76409131446f40ebda2be99f |
tokenizer.json |
9,096,718 | a898ea75433890f6610f4e470b8ebeb0c21dce5c8dd61f892eb09eb5919d2e2c |
tokenizer.json is FacebookAI/xlm-roberta-base's tokenizer, unmodified.
Conversion
The published model is fp16 and fails to load on WebGPU adapters without shader-f16 support
("Program Gather requires f16 but the device does not support it"). This build instead:
- Rewrites the published fp16 graph to fp32 (every
FLOAT16initializer,Constant,ConstantOfShapevalue, I/O type andCasttarget converted toFLOAT, and a missingcom.microsoftopset import added soonnx.checkeraccepts the graph β the graph usesSkipLayerNormalizationandBiasGelufrom that domain without declaring it). - Quantizes every MatMul with a constant weight to
com.microsoft:MatMulNBits, 8 bits, block size 32, symmetric (onnxruntime.quantization.matmul_nbits_quantizer.MatMulNBitsQuantizer) β the same recipe asfireredpunc-onnx's weight-only build. - Replaces the 250,002 Γ 768 word-embedding
Gatherwithcom.microsoft:GatherBlockQuantized, 8 bits, block size 32 along the hidden axis: uint8 data with the implicit zero point 128 and one float32 scale per 32-value block (max|w| / 127), built by hand because onnxruntime's own quantizer only emits a 4-bitGatherBlockQuantized.
No tensor in the result is float16, so it needs no shader-f16 support and runs on both the
WebGPU and WASM execution providers of onnxruntime-web 1.26.
Exact command, run against Sokuji's punctuation benchmark harness
(benchmark/punctuation-restoration/):
python parity/sat-3l-sm-q8w.py --force
This writes both sat/sat-3l-sm-q8w/model.onnx (the MatMulNBits-only build, 793,953,242 bytes,
word embedding left float32) and sat/sat-3l-sm-q8w-gather/model.onnx (the same plus the 8-bit
GatherBlockQuantized embedding, 241,945,842 bytes) plus a copied tokenizer.json in each
directory. model.onnx here is the q8w-gather output, copied unmodified β the smaller and
lower-memory of the two, and the one recommended for shipping.
Measured numbers
From docs/superpowers/notes/2026-09-14-asr-punctuation-benchmark.md (Sokuji repo, 2026-09-14):
- Parity vs. the published fp16 file (WASM, same module, 270 rows: 244 corpus rows, 12 long
multi-window rows, 14 edge-case rows):
q8w-gatherreproduces 261/270 (96.7%) identical splits; every flip has both boundary probabilities within 0.015 of the 0.25 decision threshold. The fp16 file itself, run on WASM, upcasts to float32 at load and matches a true fp32 rewrite on all 270 rows. - Quality (boundaries only, no marks inserted): ja sentence-end F1 94.1 offline, streaming P/R 93.0 / 95.2; ko sentence-end F1 90.9 offline, streaming P/R 100 / 87.5 (8-character right-context commit). Best-scoring model in the benchmark for both languages, and the only ported model covering languages beyond zh/en/ja/ko/es/fr/de/pt (85 languages upstream). English sentence-end quality depends heavily on input casing: F1 98.6 on cased input vs. 64.6 lowercased.
- Renderer cost (Electron 40.8.5, DGX Spark GB10 aarch64, WebGPU): 251 MB on disk; ja/zh/en/ko 120-character inputs run in 11β16 ms, 480-character inputs in 13β16 ms β output identical to 4-thread WASM on all 101 corpus inputs tested. Resident memory on WASM: 1,251 MB, versus 2,046 MB for the unmodified published fp16 file (which onnxruntime-web upcasts to float32 on load).
License
MIT, carried over from segment-any-text/sat-3l-sm and segment-any-text/wtpsplit. See
LICENSE in this repository (verbatim copy of the upstream wtpsplit LICENSE file).