sat-3l-sm-onnx

Weight-only 8-bit ONNX export of SaT 3L-SM (segment-any-text/sat-3l-sm), wtpsplit's 3-layer XLM-RoBERTa sentence-boundary model, with an 8-bit GatherBlockQuantized word-embedding table (the "q8w-gather" build), for use by Sokuji's local sentence-segmentation stage.

Upstream

  • Weights: segment-any-text/sat-3l-sm β€” MIT.
  • Code: segment-any-text/wtpsplit 2.2.1 β€” MIT (LICENSE in this repo is the verbatim upstream text).
  • Architecture: a 3-layer XLM-RoBERTa encoder with a single output label ("a sentence ends after this subword"), trained to work without punctuation or casing. It inserts no punctuation marks itself β€” it only predicts sentence boundaries. Covers 85 languages upstream.

Files

file bytes sha256
model.onnx 241,945,842 1132546a16fa1aa9da218602049c6e7aeba0ec5b76409131446f40ebda2be99f
tokenizer.json 9,096,718 a898ea75433890f6610f4e470b8ebeb0c21dce5c8dd61f892eb09eb5919d2e2c

tokenizer.json is FacebookAI/xlm-roberta-base's tokenizer, unmodified.

Conversion

The published model is fp16 and fails to load on WebGPU adapters without shader-f16 support ("Program Gather requires f16 but the device does not support it"). This build instead:

  1. Rewrites the published fp16 graph to fp32 (every FLOAT16 initializer, Constant, ConstantOfShape value, I/O type and Cast target converted to FLOAT, and a missing com.microsoft opset import added so onnx.checker accepts the graph β€” the graph uses SkipLayerNormalization and BiasGelu from that domain without declaring it).
  2. Quantizes every MatMul with a constant weight to com.microsoft:MatMulNBits, 8 bits, block size 32, symmetric (onnxruntime.quantization.matmul_nbits_quantizer.MatMulNBitsQuantizer) β€” the same recipe as fireredpunc-onnx's weight-only build.
  3. Replaces the 250,002 Γ— 768 word-embedding Gather with com.microsoft:GatherBlockQuantized, 8 bits, block size 32 along the hidden axis: uint8 data with the implicit zero point 128 and one float32 scale per 32-value block (max|w| / 127), built by hand because onnxruntime's own quantizer only emits a 4-bit GatherBlockQuantized.

No tensor in the result is float16, so it needs no shader-f16 support and runs on both the WebGPU and WASM execution providers of onnxruntime-web 1.26.

Exact command, run against Sokuji's punctuation benchmark harness (benchmark/punctuation-restoration/):

python parity/sat-3l-sm-q8w.py --force

This writes both sat/sat-3l-sm-q8w/model.onnx (the MatMulNBits-only build, 793,953,242 bytes, word embedding left float32) and sat/sat-3l-sm-q8w-gather/model.onnx (the same plus the 8-bit GatherBlockQuantized embedding, 241,945,842 bytes) plus a copied tokenizer.json in each directory. model.onnx here is the q8w-gather output, copied unmodified β€” the smaller and lower-memory of the two, and the one recommended for shipping.

Measured numbers

From docs/superpowers/notes/2026-09-14-asr-punctuation-benchmark.md (Sokuji repo, 2026-09-14):

  • Parity vs. the published fp16 file (WASM, same module, 270 rows: 244 corpus rows, 12 long multi-window rows, 14 edge-case rows): q8w-gather reproduces 261/270 (96.7%) identical splits; every flip has both boundary probabilities within 0.015 of the 0.25 decision threshold. The fp16 file itself, run on WASM, upcasts to float32 at load and matches a true fp32 rewrite on all 270 rows.
  • Quality (boundaries only, no marks inserted): ja sentence-end F1 94.1 offline, streaming P/R 93.0 / 95.2; ko sentence-end F1 90.9 offline, streaming P/R 100 / 87.5 (8-character right-context commit). Best-scoring model in the benchmark for both languages, and the only ported model covering languages beyond zh/en/ja/ko/es/fr/de/pt (85 languages upstream). English sentence-end quality depends heavily on input casing: F1 98.6 on cased input vs. 64.6 lowercased.
  • Renderer cost (Electron 40.8.5, DGX Spark GB10 aarch64, WebGPU): 251 MB on disk; ja/zh/en/ko 120-character inputs run in 11–16 ms, 480-character inputs in 13–16 ms β€” output identical to 4-thread WASM on all 101 corpus inputs tested. Resident memory on WASM: 1,251 MB, versus 2,046 MB for the unmodified published fp16 file (which onnxruntime-web upcasts to float32 on load).

License

MIT, carried over from segment-any-text/sat-3l-sm and segment-any-text/wtpsplit. See LICENSE in this repository (verbatim copy of the upstream wtpsplit LICENSE file).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support