--- language: - chg license: mit tags: - sentence-boundary-detection - sentence-segmentation - stanza - onnx - charlm - bilstm - low-resource-nlp - historical-linguistics - chagatai - turkic pipeline_tag: token-classification datasets: - chagatai-project/chagatai-sbd metrics: - f1 - precision - recall - accuracy model-index: - name: chagatai-sentence-segmentation results: - task: type: token-classification name: Sentence Boundary Detection dataset: name: Chagatai Historical SBD Test Set type: chagatai-project/chagatai-sbd split: test metrics: - name: Sentence Boundary F1 type: f1 value: 74.80 - name: Sentence Precision type: precision value: 79.95 - name: Sentence Recall type: recall value: 70.28 - name: Paragraph Exact Match type: accuracy value: 27.80 --- # Chagatai Sentence Boundary Detection (SBD) [Stanza Models] Neural Sentence Boundary Detection (SBD) models for **Chagatai** (*Chaghatay / جغتای*), a classical Turkic literary language written in Perso-Arabic script, based on the **Stanford Stanza** tokenization and Character Language Model (CharLM) architecture. Classical Chagatai texts (e.g., Babur's *Baburnama*, Navoi's works) were written without modern punctuation (no periods, question marks, or exclamation marks). Robust sentence boundary detection is the primary prerequisite for subsequent NLP tasks: dependency parsing, machine translation, corpus analysis, and LLM pre-training. All models are compiled into **single, self-contained ONNX files** with in-graph soft-voting ensembles and dynamic quantization, requiring **zero PyTorch or Stanza dependencies** at inference time. --- ## 📊 Benchmark & Evaluation Results Evaluated on the out-of-sample Chagatai test set (**295 paragraphs**, **71,673 character tokens**, **868 true sentence boundaries**): | Model | Type | Training Dataset | Format | Size | F1 | Precision | Recall | Exact Match | Key Benefit | |:---|:---:|:---|:---:|:---:|:---:|:---:|:---:|:---:|:---| | **Tri-Hybrid [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + `only`) | FP32 | 53.3 MB | **74.80%** | **79.95%** | 70.28% | 27.46% (81/295) | Best F1 & Precision (lowest false positives) | | **Tri-Hybrid [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + `only`) | INT8 | **13.5 MB** | **74.71%** | 79.74% | 70.28% | 27.12% (80/295) | **~4x compressed**, loss of only 0.09% F1 | | **BiCharLM [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + CharLM) | FP32 | 82.8 MB | 74.59% | 79.17% | 70.51% | **27.80% (82/295)** | **Highest Exact Match record** | | **BiCharLM [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + CharLM) | INT8 | **26.0 MB** | 74.42% | 78.94% | 70.39% | **27.80% (82/295)** | **3.2x compressed, 100% EM preserved** | | **Standalone [Stanza]** | **Single Model** | Chagatai + Uyghur + S. Uzbek (`balanced` + CharLM) | FP32 | 29.2 MB | 71.94% | 71.09% | **72.81%** | 22.37% (66/295) | Fastest single model (no ensemble) | * **Exact Match (EM)**: percentage of full paragraphs (2 to 5 sentences each) segmented with 100% precision and recall (zero errors). * **Zero Leakage**: All hyperparameter tuning and threshold selection (tau) were performed strictly on the Dev set (145 samples); Test set (295 samples) was strictly held out. --- ## 🧩 Ensemble Composition, Architectures & Datasets Each ONNX artifact bundles its entire multi-model pipeline into a single execution graph: ### 1. Tri-Hybrid Ensemble (`chagatai_sbd_tri_hybrid_stanza.onnx` / INT8) * **Ensemble Strategy**: Soft-voting probability averaging across 3 complementary inductive biases with equal weights `[1/3, 1/3, 1/3]` and calibrated threshold `tau = 0.42`. * **Sub-model 1: `FwdCharLM_Wide3`** * **Dataset**: `chagatai_uzs_uyghur_balanced` (train split: Chagatai + South Uzbek + Uyghur) + unannotated Chagatai text for CharLM. * **Architecture**: 1-layer Forward CharLM (512-dim LSTM) + 3-layer BiLSTM (hidden_dim=256, effective 512-dim) + 5 character features + Hierarchical BiLSTM second-pass. * **Role**: Contextual character representations to minimize false positives and maximize Precision. * **Sub-model 2: `Only_Deep2`** * **Dataset**: `chagatai_only` (pure historical Chagatai corpus without modern cross-lingual mixing). * **Architecture**: 2-layer BiLSTM (hidden_dim=128) + 5 character features + Hierarchical BiLSTM. * **Role**: In-domain grammatical purity, capturing Chagatai-specific syntactic structures. * **Sub-model 3: `Weighted_Balanced`** * **Dataset**: `chagatai_uzs_uyghur_balanced`. * **Architecture**: 3-layer BiLSTM (hidden_dim=256) + Weighted Cross-Entropy Loss (`pos_weight = 4.0`) + Hierarchical BiLSTM. * **Role**: High-recall sensitivity to prevent missed sentence boundaries. ### 2. BiCharLM Super-Ensemble (`chagatai_sbd_bicharlm_em_stanza.onnx` / INT8) * **Ensemble Strategy**: Soft-voting with bidirectional character language modeling, equal weights `[1/3, 1/3, 1/3]`, calibrated threshold `tau = 0.43`. * **Sub-model 1: `FwdCharLM_Wide3`** * **Dataset**: `chagatai_uzs_uyghur_balanced` + unannotated Chagatai CharLM text. * **Architecture**: Forward CharLM (512-dim) + 3-layer BiLSTM (hidden_dim=256) + Hierarchical BiLSTM. * **Sub-model 2: `Weighted_Balanced`** * **Dataset**: `chagatai_uzs_uyghur_balanced`. * **Architecture**: 3-layer BiLSTM (hidden_dim=256) + Weighted Loss (`pos_weight = 4.0`). * **Sub-model 3: `BiCharLM_Weighted`** * **Dataset**: `chagatai_uzs_uyghur_balanced` + unannotated Chagatai CharLM text. * **Architecture**: Forward CharLM (512-dim) + Backward CharLM (512-dim) = **1024-dim bidirectional contextual character representation** + 3-layer BiLSTM (hidden_dim=256) + Weighted Loss (`pos_weight = 4.0`) + Hierarchical BiLSTM. * **Role**: Backward CharLM resolves trailing ambiguities at paragraph boundaries, achieving the **all-time project record for Exact Match: 27.80% (82/295 full paragraphs segmented with zero errors)**. ### 3. Standalone Single Model (`chagatai_sbd_fwd_charlm_stanza.onnx`) * **Type**: Single model (No ensemble overhead). * **Dataset**: `chagatai_uzs_uyghur_balanced` + unannotated Chagatai CharLM text. * **Architecture**: 1-layer Forward CharLM (512-dim LSTM) + 3-layer BiLSTM (hidden_dim=256) + 5 character features + Hierarchical BiLSTM, threshold `tau = 0.24`. * **Role**: Lightweight, fast inference (29.2 MB) with high Recall (72.81%). --- ## 📚 Training Datasets Overview The models were trained on the unified dataset builds from the Chagatai Sentence Segmentation project: 1. **`chagatai_only`**: - Cleaned, deduplicated historical Chagatai texts. - Preserves classical literary syntax and archaic morphology. - Strict 70/10/20 train/dev/test split with zero out-of-sample data leakage. 2. **`chagatai_uzs_uyghur_balanced`**: - Cross-lingual transfer learning dataset combining Chagatai training sentences with equal parts South Uzbek (`uzs`) and Uyghur (`uyg`). - Balances rare Arabic-script characters and sub-word n-grams without skewing the target Chagatai evaluation distribution (Dev and Test sets contain 100% Chagatai text). 3. **`CharLM Unsupervised Corpus`**: - Historical Chagatai corpus used to pre-train Forward (512-dim) and Backward (512-dim) Character Language Models. - Enables sub-word morphological understanding without requiring fixed tokenization dictionaries. --- ## 📦 Model Files & Recommended Thresholds | File | Type | Format | Size | Threshold | Recommended For | |:---|:---:|:---:|:---:|:---:|:---| | [`chagatai_sbd_tri_hybrid_stanza.onnx`](./chagatai_sbd_tri_hybrid_stanza.onnx) | Ensemble (3 models) | FP32 | 53.3 MB | tau = 0.42 | Highest F1 score and precision | | [`chagatai_sbd_tri_hybrid_stanza_int8.onnx`](./chagatai_sbd_tri_hybrid_stanza_int8.onnx) | Ensemble (3 models) | INT8 | 13.5 MB | tau = 0.42 | Ultra-lightweight production deployment (~4x smaller) | | [`chagatai_sbd_bicharlm_em_stanza.onnx`](./chagatai_sbd_bicharlm_em_stanza.onnx) | Ensemble (3 models) | FP32 | 82.8 MB | tau = 0.43 | Highest paragraph-level exact match (27.80%) | | [`chagatai_sbd_bicharlm_em_stanza_int8.onnx`](./chagatai_sbd_bicharlm_em_stanza_int8.onnx) | Ensemble (3 models) | INT8 | 26.0 MB | tau = 0.43 | Compact exact-match ensemble (3.2x smaller) | | [`chagatai_sbd_fwd_charlm_stanza.onnx`](./chagatai_sbd_fwd_charlm_stanza.onnx) | Single Model | FP32 | 29.2 MB | tau = 0.24 | Maximum throughput (single model) | | [`vocab.json`](./vocab.json) | Vocabulary | JSON | 2.4 KB | — | 80-unit Perso-Arabic character mapping | | [`inference.py`](./inference.py) | Engine | Python | 5.2 KB | — | Zero-dependency standalone inference script | --- ## 🚀 Quickstart (Zero-Dependency Python Inference) No PyTorch, Transformers, or Stanza required. All you need is `onnxruntime` and `numpy`: ```bash pip install onnxruntime numpy huggingface_hub ``` ### Segmentation Example ```python from huggingface_hub import hf_hub_download import importlib.util REPO = "chagatai-project/chagatai-sentence-segmentation" # 1. Download model, vocabulary, and inference script (using lightweight 13.5 MB INT8 model) model_path = hf_hub_download(repo_id=REPO, filename="chagatai_sbd_tri_hybrid_stanza_int8.onnx") vocab_path = hf_hub_download(repo_id=REPO, filename="vocab.json") code_path = hf_hub_download(repo_id=REPO, filename="inference.py") # 2. Load inference engine spec = importlib.util.spec_from_file_location("inference", code_path) inf = importlib.util.module_from_spec(spec) spec.loader.exec_module(inf) # 3. Initialize detector sbd = inf.ChagataiSBD(model_path=model_path, vocab_path=vocab_path) # 4. Segment historical Chagatai text (Baburnama) text = "سلطان ابوسعید میرزا شهادت تاپقاندین سونگ سمرقند ولایتینی میرزا احمد آلدی و هرات ولایتینی سلطان حسین میرزا ضبط قیلدی و اندیجان ولایتینی عمرشیخ میرزا توتتی" sentences = sbd.segment(text) for idx, sent in enumerate(sentences, 1): print(f"[{idx}] {sent}") ``` **Output:** ```text [1] سلطان ابوسعید میرزا شهادت تاپقاندین سونگ سمرقند ولایتینی میرزا احمد آلدی [2] و هرات ولایتینی سلطان حسین میرزا ضبط قیلدی [3] و اندیجان ولایتینی عمرشیخ میرزا توتتی ``` ## 📖 Sources and Citation If you use these models or datasets in your research, please cite the underlying sources: - **Chagatai (`chg`)**: Original historical texts sourced from [KazCorpus / QazCorpora](https://qazcorpora.kz/) (National Corpus of the Kazakh Language), containing the manuscript *Shajara-i Turk* (*Genealogy of the Turks*) by Abu al-Ghazi Bahadur Khan (17th c.). - **South Uzbek (`uzs`)**: Auxiliary data sourced from the Hugging Face dataset [`tahrirchi/lutfiy`](https://huggingface.co/datasets/tahrirchi/lutfiy) (Tahrirchi, *"Filling the Gap for Uzbek: Creating Translation Resources for Southern Uzbek"*, arXiv:2508.14586).