Instructions to use chagatai-project/chagatai-sentence-segmentation with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Stanza
How to use chagatai-project/chagatai-sentence-segmentation with Stanza:
import stanza stanza.download("chagatai-sentence-segmentation") nlp = stanza.Pipeline("chagatai-sentence-segmentation") - Notebooks
- Google Colab
- Kaggle
Chagatai Sentence Boundary Detection (SBD) [Stanza Models]
Neural Sentence Boundary Detection (SBD) models for Chagatai (Chaghatay / جغتای), a classical Turkic literary language written in Perso-Arabic script, based on the Stanford Stanza tokenization and Character Language Model (CharLM) architecture.
Classical Chagatai texts (e.g., Babur's Baburnama, Navoi's works) were written without modern punctuation (no periods, question marks, or exclamation marks). Robust sentence boundary detection is the primary prerequisite for subsequent NLP tasks: dependency parsing, machine translation, corpus analysis, and LLM pre-training.
All models are compiled into single, self-contained ONNX files with in-graph soft-voting ensembles and dynamic quantization, requiring zero PyTorch or Stanza dependencies at inference time.
📊 Benchmark & Evaluation Results
Evaluated on the out-of-sample Chagatai test set (295 paragraphs, 71,673 character tokens, 868 true sentence boundaries):
| Model | Type | Training Dataset | Format | Size | F1 | Precision | Recall | Exact Match | Key Benefit |
|---|---|---|---|---|---|---|---|---|---|
| Tri-Hybrid [Stanza] | Ensemble (3 models) | Chagatai + Uyghur + S. Uzbek (balanced + only) |
FP32 | 53.3 MB | 74.80% | 79.95% | 70.28% | 27.46% (81/295) | Best F1 & Precision (lowest false positives) |
| Tri-Hybrid [Stanza] | Ensemble (3 models) | Chagatai + Uyghur + S. Uzbek (balanced + only) |
INT8 | 13.5 MB | 74.71% | 79.74% | 70.28% | 27.12% (80/295) | ~4x compressed, loss of only 0.09% F1 |
| BiCharLM [Stanza] | Ensemble (3 models) | Chagatai + Uyghur + S. Uzbek (balanced + CharLM) |
FP32 | 82.8 MB | 74.59% | 79.17% | 70.51% | 27.80% (82/295) | Highest Exact Match record |
| BiCharLM [Stanza] | Ensemble (3 models) | Chagatai + Uyghur + S. Uzbek (balanced + CharLM) |
INT8 | 26.0 MB | 74.42% | 78.94% | 70.39% | 27.80% (82/295) | 3.2x compressed, 100% EM preserved |
| Standalone [Stanza] | Single Model | Chagatai + Uyghur + S. Uzbek (balanced + CharLM) |
FP32 | 29.2 MB | 71.94% | 71.09% | 72.81% | 22.37% (66/295) | Fastest single model (no ensemble) |
- Exact Match (EM): percentage of full paragraphs (2 to 5 sentences each) segmented with 100% precision and recall (zero errors).
- Zero Leakage: All hyperparameter tuning and threshold selection (tau) were performed strictly on the Dev set (145 samples); Test set (295 samples) was strictly held out.
🧩 Ensemble Composition, Architectures & Datasets
Each ONNX artifact bundles its entire multi-model pipeline into a single execution graph:
1. Tri-Hybrid Ensemble (chagatai_sbd_tri_hybrid_stanza.onnx / INT8)
- Ensemble Strategy: Soft-voting probability averaging across 3 complementary inductive biases with equal weights
[1/3, 1/3, 1/3]and calibrated thresholdtau = 0.42. - Sub-model 1:
FwdCharLM_Wide3- Dataset:
chagatai_uzs_uyghur_balanced(train split: Chagatai + South Uzbek + Uyghur) + unannotated Chagatai text for CharLM. - Architecture: 1-layer Forward CharLM (512-dim LSTM) + 3-layer BiLSTM (hidden_dim=256, effective 512-dim) + 5 character features + Hierarchical BiLSTM second-pass.
- Role: Contextual character representations to minimize false positives and maximize Precision.
- Dataset:
- Sub-model 2:
Only_Deep2- Dataset:
chagatai_only(pure historical Chagatai corpus without modern cross-lingual mixing). - Architecture: 2-layer BiLSTM (hidden_dim=128) + 5 character features + Hierarchical BiLSTM.
- Role: In-domain grammatical purity, capturing Chagatai-specific syntactic structures.
- Dataset:
- Sub-model 3:
Weighted_Balanced- Dataset:
chagatai_uzs_uyghur_balanced. - Architecture: 3-layer BiLSTM (hidden_dim=256) + Weighted Cross-Entropy Loss (
pos_weight = 4.0) + Hierarchical BiLSTM. - Role: High-recall sensitivity to prevent missed sentence boundaries.
- Dataset:
2. BiCharLM Super-Ensemble (chagatai_sbd_bicharlm_em_stanza.onnx / INT8)
- Ensemble Strategy: Soft-voting with bidirectional character language modeling, equal weights
[1/3, 1/3, 1/3], calibrated thresholdtau = 0.43. - Sub-model 1:
FwdCharLM_Wide3- Dataset:
chagatai_uzs_uyghur_balanced+ unannotated Chagatai CharLM text. - Architecture: Forward CharLM (512-dim) + 3-layer BiLSTM (hidden_dim=256) + Hierarchical BiLSTM.
- Dataset:
- Sub-model 2:
Weighted_Balanced- Dataset:
chagatai_uzs_uyghur_balanced. - Architecture: 3-layer BiLSTM (hidden_dim=256) + Weighted Loss (
pos_weight = 4.0).
- Dataset:
- Sub-model 3:
BiCharLM_Weighted- Dataset:
chagatai_uzs_uyghur_balanced+ unannotated Chagatai CharLM text. - Architecture: Forward CharLM (512-dim) + Backward CharLM (512-dim) = 1024-dim bidirectional contextual character representation + 3-layer BiLSTM (hidden_dim=256) + Weighted Loss (
pos_weight = 4.0) + Hierarchical BiLSTM. - Role: Backward CharLM resolves trailing ambiguities at paragraph boundaries, achieving the all-time project record for Exact Match: 27.80% (82/295 full paragraphs segmented with zero errors).
- Dataset:
3. Standalone Single Model (chagatai_sbd_fwd_charlm_stanza.onnx)
- Type: Single model (No ensemble overhead).
- Dataset:
chagatai_uzs_uyghur_balanced+ unannotated Chagatai CharLM text. - Architecture: 1-layer Forward CharLM (512-dim LSTM) + 3-layer BiLSTM (hidden_dim=256) + 5 character features + Hierarchical BiLSTM, threshold
tau = 0.24. - Role: Lightweight, fast inference (29.2 MB) with high Recall (72.81%).
📚 Training Datasets Overview
The models were trained on the unified dataset builds from the Chagatai Sentence Segmentation project:
chagatai_only:- Cleaned, deduplicated historical Chagatai texts.
- Preserves classical literary syntax and archaic morphology.
- Strict 70/10/20 train/dev/test split with zero out-of-sample data leakage.
chagatai_uzs_uyghur_balanced:- Cross-lingual transfer learning dataset combining Chagatai training sentences with equal parts South Uzbek (
uzs) and Uyghur (uyg). - Balances rare Arabic-script characters and sub-word n-grams without skewing the target Chagatai evaluation distribution (Dev and Test sets contain 100% Chagatai text).
- Cross-lingual transfer learning dataset combining Chagatai training sentences with equal parts South Uzbek (
CharLM Unsupervised Corpus:- Historical Chagatai corpus used to pre-train Forward (512-dim) and Backward (512-dim) Character Language Models.
- Enables sub-word morphological understanding without requiring fixed tokenization dictionaries.
📦 Model Files & Recommended Thresholds
| File | Type | Format | Size | Threshold | Recommended For |
|---|---|---|---|---|---|
chagatai_sbd_tri_hybrid_stanza.onnx |
Ensemble (3 models) | FP32 | 53.3 MB | tau = 0.42 | Highest F1 score and precision |
chagatai_sbd_tri_hybrid_stanza_int8.onnx |
Ensemble (3 models) | INT8 | 13.5 MB | tau = 0.42 | Ultra-lightweight production deployment (~4x smaller) |
chagatai_sbd_bicharlm_em_stanza.onnx |
Ensemble (3 models) | FP32 | 82.8 MB | tau = 0.43 | Highest paragraph-level exact match (27.80%) |
chagatai_sbd_bicharlm_em_stanza_int8.onnx |
Ensemble (3 models) | INT8 | 26.0 MB | tau = 0.43 | Compact exact-match ensemble (3.2x smaller) |
chagatai_sbd_fwd_charlm_stanza.onnx |
Single Model | FP32 | 29.2 MB | tau = 0.24 | Maximum throughput (single model) |
vocab.json |
Vocabulary | JSON | 2.4 KB | — | 80-unit Perso-Arabic character mapping |
inference.py |
Engine | Python | 5.2 KB | — | Zero-dependency standalone inference script |
🚀 Quickstart (Zero-Dependency Python Inference)
No PyTorch, Transformers, or Stanza required. All you need is onnxruntime and numpy:
pip install onnxruntime numpy huggingface_hub
Segmentation Example
from huggingface_hub import hf_hub_download
import importlib.util
REPO = "chagatai-project/chagatai-sentence-segmentation"
# 1. Download model, vocabulary, and inference script (using lightweight 13.5 MB INT8 model)
model_path = hf_hub_download(repo_id=REPO, filename="chagatai_sbd_tri_hybrid_stanza_int8.onnx")
vocab_path = hf_hub_download(repo_id=REPO, filename="vocab.json")
code_path = hf_hub_download(repo_id=REPO, filename="inference.py")
# 2. Load inference engine
spec = importlib.util.spec_from_file_location("inference", code_path)
inf = importlib.util.module_from_spec(spec)
spec.loader.exec_module(inf)
# 3. Initialize detector
sbd = inf.ChagataiSBD(model_path=model_path, vocab_path=vocab_path)
# 4. Segment historical Chagatai text (Baburnama)
text = "سلطان ابوسعید میرزا شهادت تاپقاندین سونگ سمرقند ولایتینی میرزا احمد آلدی و هرات ولایتینی سلطان حسین میرزا ضبط قیلدی و اندیجان ولایتینی عمرشیخ میرزا توتتی"
sentences = sbd.segment(text)
for idx, sent in enumerate(sentences, 1):
print(f"[{idx}] {sent}")
Output:
[1] سلطان ابوسعید میرزا شهادت تاپقاندین سونگ سمرقند ولایتینی میرزا احمد آلدی
[2] و هرات ولایتینی سلطان حسین میرزا ضبط قیلدی
[3] و اندیجان ولایتینی عمرشیخ میرزا توتتی
📖 Sources and Citation
If you use these models or datasets in your research, please cite the underlying sources:
- Chagatai (
chg): Original historical texts sourced from KazCorpus / QazCorpora (National Corpus of the Kazakh Language), containing the manuscript Shajara-i Turk (Genealogy of the Turks) by Abu al-Ghazi Bahadur Khan (17th c.). - South Uzbek (
uzs): Auxiliary data sourced from the Hugging Face datasettahrirchi/lutfiy(Tahrirchi, "Filling the Gap for Uzbek: Creating Translation Resources for Southern Uzbek", arXiv:2508.14586).
- Downloads last month
- -
Dataset used to train chagatai-project/chagatai-sentence-segmentation
Paper for chagatai-project/chagatai-sentence-segmentation
Evaluation results
- Sentence Boundary F1 on Chagatai Historical SBD Test Settest set self-reported74.800
- Sentence Precision on Chagatai Historical SBD Test Settest set self-reported79.950
- Sentence Recall on Chagatai Historical SBD Test Settest set self-reported70.280
- Paragraph Exact Match on Chagatai Historical SBD Test Settest set self-reported27.800