chagatai-project's picture
docs: remove Uyghur from sources citation section
aaba3ec verified
|
Raw History Blame Contribute Delete
11.2 kB
---
language:
- chg
license: mit
tags:
- sentence-boundary-detection
- sentence-segmentation
- stanza
- onnx
- charlm
- bilstm
- low-resource-nlp
- historical-linguistics
- chagatai
- turkic
pipeline_tag: token-classification
datasets:
- chagatai-project/chagatai-sbd
metrics:
- f1
- precision
- recall
- accuracy
model-index:
- name: chagatai-sentence-segmentation
results:
- task:
type: token-classification
name: Sentence Boundary Detection
dataset:
name: Chagatai Historical SBD Test Set
type: chagatai-project/chagatai-sbd
split: test
metrics:
- name: Sentence Boundary F1
type: f1
value: 74.80
- name: Sentence Precision
type: precision
value: 79.95
- name: Sentence Recall
type: recall
value: 70.28
- name: Paragraph Exact Match
type: accuracy
value: 27.80
---
# Chagatai Sentence Boundary Detection (SBD) [Stanza Models]
Neural Sentence Boundary Detection (SBD) models for **Chagatai** (*Chaghatay / جغتای*), a classical Turkic literary language written in Perso-Arabic script, based on the **Stanford Stanza** tokenization and Character Language Model (CharLM) architecture.
Classical Chagatai texts (e.g., Babur's *Baburnama*, Navoi's works) were written without modern punctuation (no periods, question marks, or exclamation marks). Robust sentence boundary detection is the primary prerequisite for subsequent NLP tasks: dependency parsing, machine translation, corpus analysis, and LLM pre-training.
All models are compiled into **single, self-contained ONNX files** with in-graph soft-voting ensembles and dynamic quantization, requiring **zero PyTorch or Stanza dependencies** at inference time.
---
## 📊 Benchmark & Evaluation Results
Evaluated on the out-of-sample Chagatai test set (**295 paragraphs**, **71,673 character tokens**, **868 true sentence boundaries**):
| Model | Type | Training Dataset | Format | Size | F1 | Precision | Recall | Exact Match | Key Benefit |
|:---|:---:|:---|:---:|:---:|:---:|:---:|:---:|:---:|:---|
| **Tri-Hybrid [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + `only`) | FP32 | 53.3 MB | **74.80%** | **79.95%** | 70.28% | 27.46% (81/295) | Best F1 & Precision (lowest false positives) |
| **Tri-Hybrid [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + `only`) | INT8 | **13.5 MB** | **74.71%** | 79.74% | 70.28% | 27.12% (80/295) | **~4x compressed**, loss of only 0.09% F1 |
| **BiCharLM [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + CharLM) | FP32 | 82.8 MB | 74.59% | 79.17% | 70.51% | **27.80% (82/295)** | **Highest Exact Match record** |
| **BiCharLM [Stanza]** | **Ensemble (3 models)** | Chagatai + Uyghur + S. Uzbek (`balanced` + CharLM) | INT8 | **26.0 MB** | 74.42% | 78.94% | 70.39% | **27.80% (82/295)** | **3.2x compressed, 100% EM preserved** |
| **Standalone [Stanza]** | **Single Model** | Chagatai + Uyghur + S. Uzbek (`balanced` + CharLM) | FP32 | 29.2 MB | 71.94% | 71.09% | **72.81%** | 22.37% (66/295) | Fastest single model (no ensemble) |
* **Exact Match (EM)**: percentage of full paragraphs (2 to 5 sentences each) segmented with 100% precision and recall (zero errors).
* **Zero Leakage**: All hyperparameter tuning and threshold selection (tau) were performed strictly on the Dev set (145 samples); Test set (295 samples) was strictly held out.
---
## 🧩 Ensemble Composition, Architectures & Datasets
Each ONNX artifact bundles its entire multi-model pipeline into a single execution graph:
### 1. Tri-Hybrid Ensemble (`chagatai_sbd_tri_hybrid_stanza.onnx` / INT8)
* **Ensemble Strategy**: Soft-voting probability averaging across 3 complementary inductive biases with equal weights `[1/3, 1/3, 1/3]` and calibrated threshold `tau = 0.42`.
* **Sub-model 1: `FwdCharLM_Wide3`**
* **Dataset**: `chagatai_uzs_uyghur_balanced` (train split: Chagatai + South Uzbek + Uyghur) + unannotated Chagatai text for CharLM.
* **Architecture**: 1-layer Forward CharLM (512-dim LSTM) + 3-layer BiLSTM (hidden_dim=256, effective 512-dim) + 5 character features + Hierarchical BiLSTM second-pass.
* **Role**: Contextual character representations to minimize false positives and maximize Precision.
* **Sub-model 2: `Only_Deep2`**
* **Dataset**: `chagatai_only` (pure historical Chagatai corpus without modern cross-lingual mixing).
* **Architecture**: 2-layer BiLSTM (hidden_dim=128) + 5 character features + Hierarchical BiLSTM.
* **Role**: In-domain grammatical purity, capturing Chagatai-specific syntactic structures.
* **Sub-model 3: `Weighted_Balanced`**
* **Dataset**: `chagatai_uzs_uyghur_balanced`.
* **Architecture**: 3-layer BiLSTM (hidden_dim=256) + Weighted Cross-Entropy Loss (`pos_weight = 4.0`) + Hierarchical BiLSTM.
* **Role**: High-recall sensitivity to prevent missed sentence boundaries.
### 2. BiCharLM Super-Ensemble (`chagatai_sbd_bicharlm_em_stanza.onnx` / INT8)
* **Ensemble Strategy**: Soft-voting with bidirectional character language modeling, equal weights `[1/3, 1/3, 1/3]`, calibrated threshold `tau = 0.43`.
* **Sub-model 1: `FwdCharLM_Wide3`**
* **Dataset**: `chagatai_uzs_uyghur_balanced` + unannotated Chagatai CharLM text.
* **Architecture**: Forward CharLM (512-dim) + 3-layer BiLSTM (hidden_dim=256) + Hierarchical BiLSTM.
* **Sub-model 2: `Weighted_Balanced`**
* **Dataset**: `chagatai_uzs_uyghur_balanced`.
* **Architecture**: 3-layer BiLSTM (hidden_dim=256) + Weighted Loss (`pos_weight = 4.0`).
* **Sub-model 3: `BiCharLM_Weighted`**
* **Dataset**: `chagatai_uzs_uyghur_balanced` + unannotated Chagatai CharLM text.
* **Architecture**: Forward CharLM (512-dim) + Backward CharLM (512-dim) = **1024-dim bidirectional contextual character representation** + 3-layer BiLSTM (hidden_dim=256) + Weighted Loss (`pos_weight = 4.0`) + Hierarchical BiLSTM.
* **Role**: Backward CharLM resolves trailing ambiguities at paragraph boundaries, achieving the **all-time project record for Exact Match: 27.80% (82/295 full paragraphs segmented with zero errors)**.
### 3. Standalone Single Model (`chagatai_sbd_fwd_charlm_stanza.onnx`)
* **Type**: Single model (No ensemble overhead).
* **Dataset**: `chagatai_uzs_uyghur_balanced` + unannotated Chagatai CharLM text.
* **Architecture**: 1-layer Forward CharLM (512-dim LSTM) + 3-layer BiLSTM (hidden_dim=256) + 5 character features + Hierarchical BiLSTM, threshold `tau = 0.24`.
* **Role**: Lightweight, fast inference (29.2 MB) with high Recall (72.81%).
---
## 📚 Training Datasets Overview
The models were trained on the unified dataset builds from the Chagatai Sentence Segmentation project:
1. **`chagatai_only`**:
- Cleaned, deduplicated historical Chagatai texts.
- Preserves classical literary syntax and archaic morphology.
- Strict 70/10/20 train/dev/test split with zero out-of-sample data leakage.
2. **`chagatai_uzs_uyghur_balanced`**:
- Cross-lingual transfer learning dataset combining Chagatai training sentences with equal parts South Uzbek (`uzs`) and Uyghur (`uyg`).
- Balances rare Arabic-script characters and sub-word n-grams without skewing the target Chagatai evaluation distribution (Dev and Test sets contain 100% Chagatai text).
3. **`CharLM Unsupervised Corpus`**:
- Historical Chagatai corpus used to pre-train Forward (512-dim) and Backward (512-dim) Character Language Models.
- Enables sub-word morphological understanding without requiring fixed tokenization dictionaries.
---
## 📦 Model Files & Recommended Thresholds
| File | Type | Format | Size | Threshold | Recommended For |
|:---|:---:|:---:|:---:|:---:|:---|
| [`chagatai_sbd_tri_hybrid_stanza.onnx`](./chagatai_sbd_tri_hybrid_stanza.onnx) | Ensemble (3 models) | FP32 | 53.3 MB | tau = 0.42 | Highest F1 score and precision |
| [`chagatai_sbd_tri_hybrid_stanza_int8.onnx`](./chagatai_sbd_tri_hybrid_stanza_int8.onnx) | Ensemble (3 models) | INT8 | 13.5 MB | tau = 0.42 | Ultra-lightweight production deployment (~4x smaller) |
| [`chagatai_sbd_bicharlm_em_stanza.onnx`](./chagatai_sbd_bicharlm_em_stanza.onnx) | Ensemble (3 models) | FP32 | 82.8 MB | tau = 0.43 | Highest paragraph-level exact match (27.80%) |
| [`chagatai_sbd_bicharlm_em_stanza_int8.onnx`](./chagatai_sbd_bicharlm_em_stanza_int8.onnx) | Ensemble (3 models) | INT8 | 26.0 MB | tau = 0.43 | Compact exact-match ensemble (3.2x smaller) |
| [`chagatai_sbd_fwd_charlm_stanza.onnx`](./chagatai_sbd_fwd_charlm_stanza.onnx) | Single Model | FP32 | 29.2 MB | tau = 0.24 | Maximum throughput (single model) |
| [`vocab.json`](./vocab.json) | Vocabulary | JSON | 2.4 KB | — | 80-unit Perso-Arabic character mapping |
| [`inference.py`](./inference.py) | Engine | Python | 5.2 KB | — | Zero-dependency standalone inference script |
---
## 🚀 Quickstart (Zero-Dependency Python Inference)
No PyTorch, Transformers, or Stanza required. All you need is `onnxruntime` and `numpy`:
```bash
pip install onnxruntime numpy huggingface_hub
```
### Segmentation Example
```python
from huggingface_hub import hf_hub_download
import importlib.util
REPO = "chagatai-project/chagatai-sentence-segmentation"
# 1. Download model, vocabulary, and inference script (using lightweight 13.5 MB INT8 model)
model_path = hf_hub_download(repo_id=REPO, filename="chagatai_sbd_tri_hybrid_stanza_int8.onnx")
vocab_path = hf_hub_download(repo_id=REPO, filename="vocab.json")
code_path = hf_hub_download(repo_id=REPO, filename="inference.py")
# 2. Load inference engine
spec = importlib.util.spec_from_file_location("inference", code_path)
inf = importlib.util.module_from_spec(spec)
spec.loader.exec_module(inf)
# 3. Initialize detector
sbd = inf.ChagataiSBD(model_path=model_path, vocab_path=vocab_path)
# 4. Segment historical Chagatai text (Baburnama)
text = "سلطان ابوسعید میرزا شهادت تاپقاندین سونگ سمرقند ولایتینی میرزا احمد آلدی و هرات ولایتینی سلطان حسین میرزا ضبط قیلدی و اندیجان ولایتینی عمرشیخ میرزا توتتی"
sentences = sbd.segment(text)
for idx, sent in enumerate(sentences, 1):
print(f"[{idx}] {sent}")
```
**Output:**
```text
[1] سلطان ابوسعید میرزا شهادت تاپقاندین سونگ سمرقند ولایتینی میرزا احمد آلدی
[2] و هرات ولایتینی سلطان حسین میرزا ضبط قیلدی
[3] و اندیجان ولایتینی عمرشیخ میرزا توتتی
```
## 📖 Sources and Citation
If you use these models or datasets in your research, please cite the underlying sources:
- **Chagatai (`chg`)**: Original historical texts sourced from [KazCorpus / QazCorpora](https://qazcorpora.kz/) (National Corpus of the Kazakh Language), containing the manuscript *Shajara-i Turk* (*Genealogy of the Turks*) by Abu al-Ghazi Bahadur Khan (17th c.).
- **South Uzbek (`uzs`)**: Auxiliary data sourced from the Hugging Face dataset [`tahrirchi/lutfiy`](https://huggingface.co/datasets/tahrirchi/lutfiy) (Tahrirchi, *"Filling the Gap for Uzbek: Creating Translation Resources for Southern Uzbek"*, arXiv:2508.14586).