Chagatai Sentence Boundary Detection (SBD) [Stanza Models]

Neural Sentence Boundary Detection (SBD) models for Chagatai (Chaghatay / جغتای), a classical Turkic literary language written in Perso-Arabic script, based on the Stanford Stanza tokenization and Character Language Model (CharLM) architecture.

Classical Chagatai texts (e.g., Babur's Baburnama, Navoi's works) were written without modern punctuation (no periods, question marks, or exclamation marks). Robust sentence boundary detection is the primary prerequisite for subsequent NLP tasks: dependency parsing, machine translation, corpus analysis, and LLM pre-training.

All models are compiled into single, self-contained ONNX files with in-graph soft-voting ensembles and dynamic quantization, requiring zero PyTorch or Stanza dependencies at inference time.


📊 Benchmark & Evaluation Results

Evaluated on the out-of-sample Chagatai test set (295 paragraphs, 71,673 character tokens, 868 true sentence boundaries):

Model Type Training Dataset Format Size F1 Precision Recall Exact Match Key Benefit
Tri-Hybrid [Stanza] Ensemble (3 models) Chagatai + Uyghur + S. Uzbek (balanced + only) FP32 53.3 MB 74.80% 79.95% 70.28% 27.46% (81/295) Best F1 & Precision (lowest false positives)
Tri-Hybrid [Stanza] Ensemble (3 models) Chagatai + Uyghur + S. Uzbek (balanced + only) INT8 13.5 MB 74.71% 79.74% 70.28% 27.12% (80/295) ~4x compressed, loss of only 0.09% F1
BiCharLM [Stanza] Ensemble (3 models) Chagatai + Uyghur + S. Uzbek (balanced + CharLM) FP32 82.8 MB 74.59% 79.17% 70.51% 27.80% (82/295) Highest Exact Match record
BiCharLM [Stanza] Ensemble (3 models) Chagatai + Uyghur + S. Uzbek (balanced + CharLM) INT8 26.0 MB 74.42% 78.94% 70.39% 27.80% (82/295) 3.2x compressed, 100% EM preserved
Standalone [Stanza] Single Model Chagatai + Uyghur + S. Uzbek (balanced + CharLM) FP32 29.2 MB 71.94% 71.09% 72.81% 22.37% (66/295) Fastest single model (no ensemble)
  • Exact Match (EM): percentage of full paragraphs (2 to 5 sentences each) segmented with 100% precision and recall (zero errors).
  • Zero Leakage: All hyperparameter tuning and threshold selection (tau) were performed strictly on the Dev set (145 samples); Test set (295 samples) was strictly held out.

🧩 Ensemble Composition, Architectures & Datasets

Each ONNX artifact bundles its entire multi-model pipeline into a single execution graph:

1. Tri-Hybrid Ensemble (chagatai_sbd_tri_hybrid_stanza.onnx / INT8)

  • Ensemble Strategy: Soft-voting probability averaging across 3 complementary inductive biases with equal weights [1/3, 1/3, 1/3] and calibrated threshold tau = 0.42.
  • Sub-model 1: FwdCharLM_Wide3
    • Dataset: chagatai_uzs_uyghur_balanced (train split: Chagatai + South Uzbek + Uyghur) + unannotated Chagatai text for CharLM.
    • Architecture: 1-layer Forward CharLM (512-dim LSTM) + 3-layer BiLSTM (hidden_dim=256, effective 512-dim) + 5 character features + Hierarchical BiLSTM second-pass.
    • Role: Contextual character representations to minimize false positives and maximize Precision.
  • Sub-model 2: Only_Deep2
    • Dataset: chagatai_only (pure historical Chagatai corpus without modern cross-lingual mixing).
    • Architecture: 2-layer BiLSTM (hidden_dim=128) + 5 character features + Hierarchical BiLSTM.
    • Role: In-domain grammatical purity, capturing Chagatai-specific syntactic structures.
  • Sub-model 3: Weighted_Balanced
    • Dataset: chagatai_uzs_uyghur_balanced.
    • Architecture: 3-layer BiLSTM (hidden_dim=256) + Weighted Cross-Entropy Loss (pos_weight = 4.0) + Hierarchical BiLSTM.
    • Role: High-recall sensitivity to prevent missed sentence boundaries.

2. BiCharLM Super-Ensemble (chagatai_sbd_bicharlm_em_stanza.onnx / INT8)

  • Ensemble Strategy: Soft-voting with bidirectional character language modeling, equal weights [1/3, 1/3, 1/3], calibrated threshold tau = 0.43.
  • Sub-model 1: FwdCharLM_Wide3
    • Dataset: chagatai_uzs_uyghur_balanced + unannotated Chagatai CharLM text.
    • Architecture: Forward CharLM (512-dim) + 3-layer BiLSTM (hidden_dim=256) + Hierarchical BiLSTM.
  • Sub-model 2: Weighted_Balanced
    • Dataset: chagatai_uzs_uyghur_balanced.
    • Architecture: 3-layer BiLSTM (hidden_dim=256) + Weighted Loss (pos_weight = 4.0).
  • Sub-model 3: BiCharLM_Weighted
    • Dataset: chagatai_uzs_uyghur_balanced + unannotated Chagatai CharLM text.
    • Architecture: Forward CharLM (512-dim) + Backward CharLM (512-dim) = 1024-dim bidirectional contextual character representation + 3-layer BiLSTM (hidden_dim=256) + Weighted Loss (pos_weight = 4.0) + Hierarchical BiLSTM.
    • Role: Backward CharLM resolves trailing ambiguities at paragraph boundaries, achieving the all-time project record for Exact Match: 27.80% (82/295 full paragraphs segmented with zero errors).

3. Standalone Single Model (chagatai_sbd_fwd_charlm_stanza.onnx)

  • Type: Single model (No ensemble overhead).
  • Dataset: chagatai_uzs_uyghur_balanced + unannotated Chagatai CharLM text.
  • Architecture: 1-layer Forward CharLM (512-dim LSTM) + 3-layer BiLSTM (hidden_dim=256) + 5 character features + Hierarchical BiLSTM, threshold tau = 0.24.
  • Role: Lightweight, fast inference (29.2 MB) with high Recall (72.81%).

📚 Training Datasets Overview

The models were trained on the unified dataset builds from the Chagatai Sentence Segmentation project:

  1. chagatai_only:
    • Cleaned, deduplicated historical Chagatai texts.
    • Preserves classical literary syntax and archaic morphology.
    • Strict 70/10/20 train/dev/test split with zero out-of-sample data leakage.
  2. chagatai_uzs_uyghur_balanced:
    • Cross-lingual transfer learning dataset combining Chagatai training sentences with equal parts South Uzbek (uzs) and Uyghur (uyg).
    • Balances rare Arabic-script characters and sub-word n-grams without skewing the target Chagatai evaluation distribution (Dev and Test sets contain 100% Chagatai text).
  3. CharLM Unsupervised Corpus:
    • Historical Chagatai corpus used to pre-train Forward (512-dim) and Backward (512-dim) Character Language Models.
    • Enables sub-word morphological understanding without requiring fixed tokenization dictionaries.

📦 Model Files & Recommended Thresholds

File Type Format Size Threshold Recommended For
chagatai_sbd_tri_hybrid_stanza.onnx Ensemble (3 models) FP32 53.3 MB tau = 0.42 Highest F1 score and precision
chagatai_sbd_tri_hybrid_stanza_int8.onnx Ensemble (3 models) INT8 13.5 MB tau = 0.42 Ultra-lightweight production deployment (~4x smaller)
chagatai_sbd_bicharlm_em_stanza.onnx Ensemble (3 models) FP32 82.8 MB tau = 0.43 Highest paragraph-level exact match (27.80%)
chagatai_sbd_bicharlm_em_stanza_int8.onnx Ensemble (3 models) INT8 26.0 MB tau = 0.43 Compact exact-match ensemble (3.2x smaller)
chagatai_sbd_fwd_charlm_stanza.onnx Single Model FP32 29.2 MB tau = 0.24 Maximum throughput (single model)
vocab.json Vocabulary JSON 2.4 KB — 80-unit Perso-Arabic character mapping
inference.py Engine Python 5.2 KB — Zero-dependency standalone inference script

🚀 Quickstart (Zero-Dependency Python Inference)

No PyTorch, Transformers, or Stanza required. All you need is onnxruntime and numpy:

pip install onnxruntime numpy huggingface_hub

Segmentation Example

from huggingface_hub import hf_hub_download
import importlib.util

REPO = "chagatai-project/chagatai-sentence-segmentation"

# 1. Download model, vocabulary, and inference script (using lightweight 13.5 MB INT8 model)
model_path = hf_hub_download(repo_id=REPO, filename="chagatai_sbd_tri_hybrid_stanza_int8.onnx")
vocab_path = hf_hub_download(repo_id=REPO, filename="vocab.json")
code_path = hf_hub_download(repo_id=REPO, filename="inference.py")

# 2. Load inference engine
spec = importlib.util.spec_from_file_location("inference", code_path)
inf = importlib.util.module_from_spec(spec)
spec.loader.exec_module(inf)

# 3. Initialize detector
sbd = inf.ChagataiSBD(model_path=model_path, vocab_path=vocab_path)

# 4. Segment historical Chagatai text (Baburnama)
text = "سلطان ابوسعید میرزا شهادت تاپقاندین سونگ سمرقند ولایتینی میرزا احمد آلدی و هرات ولایتینی سلطان حسین میرزا ضبط قیلدی و اندیجان ولایتینی عمرشیخ میرزا توتتی"

sentences = sbd.segment(text)
for idx, sent in enumerate(sentences, 1):
    print(f"[{idx}] {sent}")

Output:

[1] سلطان ابوسعید میرزا شهادت تاپقاندین سونگ سمرقند ولایتینی میرزا احمد آلدی
[2] و هرات ولایتینی سلطان حسین میرزا ضبط قیلدی
[3] و اندیجان ولایتینی عمرشیخ میرزا توتتی

📖 Sources and Citation

If you use these models or datasets in your research, please cite the underlying sources:

  • Chagatai (chg): Original historical texts sourced from KazCorpus / QazCorpora (National Corpus of the Kazakh Language), containing the manuscript Shajara-i Turk (Genealogy of the Turks) by Abu al-Ghazi Bahadur Khan (17th c.).
  • South Uzbek (uzs): Auxiliary data sourced from the Hugging Face dataset tahrirchi/lutfiy (Tahrirchi, "Filling the Gap for Uzbek: Creating Translation Resources for Southern Uzbek", arXiv:2508.14586).
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train chagatai-project/chagatai-sentence-segmentation

Paper for chagatai-project/chagatai-sentence-segmentation

Evaluation results