Token Classification
Stanza
ONNX
Chagatai
sentence-boundary-detection
sentence-segmentation
charlm
bilstm
low-resource-nlp
historical-linguistics
chagatai
turkic
Eval Results (legacy)
Instructions to use chagatai-project/chagatai-sentence-segmentation with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Stanza
How to use chagatai-project/chagatai-sentence-segmentation with Stanza:
import stanza stanza.download("chagatai-sentence-segmentation") nlp = stanza.Pipeline("chagatai-sentence-segmentation") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from chagatai-project/chagatai-sentence-segmentation: direct link, hf CLI and curl.
- Browser
- Download file 11.2 kB
-
https://huggingface.co/chagatai-project/chagatai-sentence-segmentation/resolve/main/README.md
- Command line
-
hf download hf://chagatai-project/chagatai-sentence-segmentation/README.md
-
curl -L -o README.md https://huggingface.co/chagatai-project/chagatai-sentence-segmentation/resolve/main/README.md
11.2 kB
| language: | |
| - chg | |
| license: mit | |
| tags: | |
| - sentence-boundary-detection | |
| - sentence-segmentation | |
| - stanza | |
| - onnx | |
| - charlm | |
| - bilstm | |
| - low-resource-nlp | |
| - historical-linguistics | |
| - chagatai | |
| - turkic | |
| pipeline_tag: token-classification | |
| datasets: | |
| - chagatai-project/chagatai-sbd | |
| metrics: | |
| - f1 | |
| - precision | |
| - recall | |
| - accuracy | |
| model-index: | |
| - name: chagatai-sentence-segmentation | |
| results: | |
| - task: | |
| type: token-classification | |
| name: Sentence Boundary Detection | |
| dataset: | |
| name: Chagatai Historical SBD Test Set | |
| type: chagatai-project/chagatai-sbd | |
| split: test | |
| metrics: | |
| - name: Sentence Boundary F1 | |
| type: f1 | |
| value: 74.80 | |
| - name: Sentence Precision | |
| type: precision | |
| value: 79.95 | |
| - name: Sentence Recall | |
| type: recall | |
| value: 70.28 | |
| - name: Paragraph Exact Match | |
| type: accuracy | |
| value: 27.80 | |
| # Chagatai Sentence Boundary Detection (SBD) [Stanza Models] | |
| Neural Sentence Boundary Detection (SBD) models for **Chagatai** (*Chaghatay / جغتای*), a classical Turkic literary language written in Perso-Arabic script, based on the **Stanford Stanza** tokenization and Character Language Model (CharLM) architecture. | |
| Classical Chagatai texts (e.g., Babur's *Baburnama*, Navoi's works) were written without modern punctuation (no periods, question marks, or exclamation marks). Robust sentence boundary detection is the primary prerequisite for subsequent NLP tasks: dependency parsing, machine translation, corpus analysis, and LLM pre-training. | |
| All models are compiled into **single, self-contained ONNX files** with in-graph soft-voting ensembles and dynamic quantization, requiring **zero PyTorch or Stanza dependencies** at inference time. | |
| --- | |
| ## 📊 Benchmark & Evaluation Results | |
| Evaluated on the out-of-sample Chagatai test set (**295 paragraphs**, **71,673 character tokens**, **868 true sentence boundaries**): | |
| | Model | Type | Training Dataset | Format | Size | F1 | Precision | Recall | Exact Match | Key Benefit | | |
| |:---|:---:|:---|:---:|:---:|:---:|:---:|:---:|:---:|:---| | |
| | **Tri-Hybrid [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + `only`) | FP32 | 53.3 MB | **74.80%** | **79.95%** | 70.28% | 27.46% (81/295) | Best F1 & Precision (lowest false positives) | | |
| | **Tri-Hybrid [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + `only`) | INT8 | **13.5 MB** | **74.71%** | 79.74% | 70.28% | 27.12% (80/295) | **~4x compressed**, loss of only 0.09% F1 | | |
| | **BiCharLM [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + CharLM) | FP32 | 82.8 MB | 74.59% | 79.17% | 70.51% | **27.80% (82/295)** | **Highest Exact Match record** | | |
| | **BiCharLM [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + CharLM) | INT8 | **26.0 MB** | 74.42% | 78.94% | 70.39% | **27.80% (82/295)** | **3.2x compressed, 100% EM preserved** | | |
| | **Standalone [Stanza]** | **Single Model** | Chagatai + Uyghur + S. Uzbek (`balanced` + CharLM) | FP32 | 29.2 MB | 71.94% | 71.09% | **72.81%** | 22.37% (66/295) | Fastest single model (no ensemble) | | |
| * **Exact Match (EM)**: percentage of full paragraphs (2 to 5 sentences each) segmented with 100% precision and recall (zero errors). | |
| * **Zero Leakage**: All hyperparameter tuning and threshold selection (tau) were performed strictly on the Dev set (145 samples); Test set (295 samples) was strictly held out. | |
| --- | |
| ## 🧩 Ensemble Composition, Architectures & Datasets | |
| Each ONNX artifact bundles its entire multi-model pipeline into a single execution graph: | |
| ### 1. Tri-Hybrid Ensemble (`chagatai_sbd_tri_hybrid_stanza.onnx` / INT8) | |
| * **Ensemble Strategy**: Soft-voting probability averaging across 3 complementary inductive biases with equal weights `[1/3, 1/3, 1/3]` and calibrated threshold `tau = 0.42`. | |
| * **Sub-model 1: `FwdCharLM_Wide3`** | |
| * **Dataset**: `chagatai_uzs_uyghur_balanced` (train split: Chagatai + South Uzbek + Uyghur) + unannotated Chagatai text for CharLM. | |
| * **Architecture**: 1-layer Forward CharLM (512-dim LSTM) + 3-layer BiLSTM (hidden_dim=256, effective 512-dim) + 5 character features + Hierarchical BiLSTM second-pass. | |
| * **Role**: Contextual character representations to minimize false positives and maximize Precision. | |
| * **Sub-model 2: `Only_Deep2`** | |
| * **Dataset**: `chagatai_only` (pure historical Chagatai corpus without modern cross-lingual mixing). | |
| * **Architecture**: 2-layer BiLSTM (hidden_dim=128) + 5 character features + Hierarchical BiLSTM. | |
| * **Role**: In-domain grammatical purity, capturing Chagatai-specific syntactic structures. | |
| * **Sub-model 3: `Weighted_Balanced`** | |
| * **Dataset**: `chagatai_uzs_uyghur_balanced`. | |
| * **Architecture**: 3-layer BiLSTM (hidden_dim=256) + Weighted Cross-Entropy Loss (`pos_weight = 4.0`) + Hierarchical BiLSTM. | |
| * **Role**: High-recall sensitivity to prevent missed sentence boundaries. | |
| ### 2. BiCharLM Super-Ensemble (`chagatai_sbd_bicharlm_em_stanza.onnx` / INT8) | |
| * **Ensemble Strategy**: Soft-voting with bidirectional character language modeling, equal weights `[1/3, 1/3, 1/3]`, calibrated threshold `tau = 0.43`. | |
| * **Sub-model 1: `FwdCharLM_Wide3`** | |
| * **Dataset**: `chagatai_uzs_uyghur_balanced` + unannotated Chagatai CharLM text. | |
| * **Architecture**: Forward CharLM (512-dim) + 3-layer BiLSTM (hidden_dim=256) + Hierarchical BiLSTM. | |
| * **Sub-model 2: `Weighted_Balanced`** | |
| * **Dataset**: `chagatai_uzs_uyghur_balanced`. | |
| * **Architecture**: 3-layer BiLSTM (hidden_dim=256) + Weighted Loss (`pos_weight = 4.0`). | |
| * **Sub-model 3: `BiCharLM_Weighted`** | |
| * **Dataset**: `chagatai_uzs_uyghur_balanced` + unannotated Chagatai CharLM text. | |
| * **Architecture**: Forward CharLM (512-dim) + Backward CharLM (512-dim) = **1024-dim bidirectional contextual character representation** + 3-layer BiLSTM (hidden_dim=256) + Weighted Loss (`pos_weight = 4.0`) + Hierarchical BiLSTM. | |
| * **Role**: Backward CharLM resolves trailing ambiguities at paragraph boundaries, achieving the **all-time project record for Exact Match: 27.80% (82/295 full paragraphs segmented with zero errors)**. | |
| ### 3. Standalone Single Model (`chagatai_sbd_fwd_charlm_stanza.onnx`) | |
| * **Type**: Single model (No ensemble overhead). | |
| * **Dataset**: `chagatai_uzs_uyghur_balanced` + unannotated Chagatai CharLM text. | |
| * **Architecture**: 1-layer Forward CharLM (512-dim LSTM) + 3-layer BiLSTM (hidden_dim=256) + 5 character features + Hierarchical BiLSTM, threshold `tau = 0.24`. | |
| * **Role**: Lightweight, fast inference (29.2 MB) with high Recall (72.81%). | |
| --- | |
| ## 📚 Training Datasets Overview | |
| The models were trained on the unified dataset builds from the Chagatai Sentence Segmentation project: | |
| 1. **`chagatai_only`**: | |
| - Cleaned, deduplicated historical Chagatai texts. | |
| - Preserves classical literary syntax and archaic morphology. | |
| - Strict 70/10/20 train/dev/test split with zero out-of-sample data leakage. | |
| 2. **`chagatai_uzs_uyghur_balanced`**: | |
| - Cross-lingual transfer learning dataset combining Chagatai training sentences with equal parts South Uzbek (`uzs`) and Uyghur (`uyg`). | |
| - Balances rare Arabic-script characters and sub-word n-grams without skewing the target Chagatai evaluation distribution (Dev and Test sets contain 100% Chagatai text). | |
| 3. **`CharLM Unsupervised Corpus`**: | |
| - Historical Chagatai corpus used to pre-train Forward (512-dim) and Backward (512-dim) Character Language Models. | |
| - Enables sub-word morphological understanding without requiring fixed tokenization dictionaries. | |
| --- | |
| ## 📦 Model Files & Recommended Thresholds | |
| | File | Type | Format | Size | Threshold | Recommended For | | |
| |:---|:---:|:---:|:---:|:---:|:---| | |
| | [`chagatai_sbd_tri_hybrid_stanza.onnx`](./chagatai_sbd_tri_hybrid_stanza.onnx) | Ensemble (3 models) | FP32 | 53.3 MB | tau = 0.42 | Highest F1 score and precision | | |
| | [`chagatai_sbd_tri_hybrid_stanza_int8.onnx`](./chagatai_sbd_tri_hybrid_stanza_int8.onnx) | Ensemble (3 models) | INT8 | 13.5 MB | tau = 0.42 | Ultra-lightweight production deployment (~4x smaller) | | |
| | [`chagatai_sbd_bicharlm_em_stanza.onnx`](./chagatai_sbd_bicharlm_em_stanza.onnx) | Ensemble (3 models) | FP32 | 82.8 MB | tau = 0.43 | Highest paragraph-level exact match (27.80%) | | |
| | [`chagatai_sbd_bicharlm_em_stanza_int8.onnx`](./chagatai_sbd_bicharlm_em_stanza_int8.onnx) | Ensemble (3 models) | INT8 | 26.0 MB | tau = 0.43 | Compact exact-match ensemble (3.2x smaller) | | |
| | [`chagatai_sbd_fwd_charlm_stanza.onnx`](./chagatai_sbd_fwd_charlm_stanza.onnx) | Single Model | FP32 | 29.2 MB | tau = 0.24 | Maximum throughput (single model) | | |
| | [`vocab.json`](./vocab.json) | Vocabulary | JSON | 2.4 KB | — | 80-unit Perso-Arabic character mapping | | |
| | [`inference.py`](./inference.py) | Engine | Python | 5.2 KB | — | Zero-dependency standalone inference script | | |
| --- | |
| ## 🚀 Quickstart (Zero-Dependency Python Inference) | |
| No PyTorch, Transformers, or Stanza required. All you need is `onnxruntime` and `numpy`: | |
| ```bash | |
| pip install onnxruntime numpy huggingface_hub | |
| ``` | |
| ### Segmentation Example | |
| ```python | |
| from huggingface_hub import hf_hub_download | |
| import importlib.util | |
| REPO = "chagatai-project/chagatai-sentence-segmentation" | |
| # 1. Download model, vocabulary, and inference script (using lightweight 13.5 MB INT8 model) | |
| model_path = hf_hub_download(repo_id=REPO, filename="chagatai_sbd_tri_hybrid_stanza_int8.onnx") | |
| vocab_path = hf_hub_download(repo_id=REPO, filename="vocab.json") | |
| code_path = hf_hub_download(repo_id=REPO, filename="inference.py") | |
| # 2. Load inference engine | |
| spec = importlib.util.spec_from_file_location("inference", code_path) | |
| inf = importlib.util.module_from_spec(spec) | |
| spec.loader.exec_module(inf) | |
| # 3. Initialize detector | |
| sbd = inf.ChagataiSBD(model_path=model_path, vocab_path=vocab_path) | |
| # 4. Segment historical Chagatai text (Baburnama) | |
| text = "سلطان ابوسعید میرزا شهادت تاپقاندین سونگ سمرقند ولایتینی میرزا احمد آلدی و هرات ولایتینی سلطان حسین میرزا ضبط قیلدی و اندیجان ولایتینی عمرشیخ میرزا توتتی" | |
| sentences = sbd.segment(text) | |
| for idx, sent in enumerate(sentences, 1): | |
| print(f"[{idx}] {sent}") | |
| ``` | |
| **Output:** | |
| ```text | |
| [1] سلطان ابوسعید میرزا شهادت تاپقاندین سونگ سمرقند ولایتینی میرزا احمد آلدی | |
| [2] و هرات ولایتینی سلطان حسین میرزا ضبط قیلدی | |
| [3] و اندیجان ولایتینی عمرشیخ میرزا توتتی | |
| ``` | |
| ## 📖 Sources and Citation | |
| If you use these models or datasets in your research, please cite the underlying sources: | |
| - **Chagatai (`chg`)**: Original historical texts sourced from [KazCorpus / QazCorpora](https://qazcorpora.kz/) (National Corpus of the Kazakh Language), containing the manuscript *Shajara-i Turk* (*Genealogy of the Turks*) by Abu al-Ghazi Bahadur Khan (17th c.). | |
| - **South Uzbek (`uzs`)**: Auxiliary data sourced from the Hugging Face dataset [`tahrirchi/lutfiy`](https://huggingface.co/datasets/tahrirchi/lutfiy) (Tahrirchi, *"Filling the Gap for Uzbek: Creating Translation Resources for Southern Uzbek"*, arXiv:2508.14586). | |