ansarzeinulla commited on
Commit
2eb2260
·
verified ·
1 Parent(s): 2797004

Card: real training settings, token count, TD against the human reference

Browse files
Files changed (1) hide show
  1. README.md +37 -101
README.md CHANGED
@@ -12,130 +12,66 @@ tags:
12
  - low-resource
13
  - nogai
14
  - turkic
15
- - kazakh
16
- - cross-lingual
17
- - continuous-pre-training
18
  datasets:
19
  - ansarzeinulla/Nogai-Unified-Corpus-v1
20
- metrics:
21
- - perplexity
22
  ---
23
 
24
- <div align="center">
25
 
26
- # ⚡️ Qwen2.5-1.5B-Nogai-LoRA (Phase 1 Baseline)
27
- ### Experimental Continuous Pre-Training (CPT) Adapter for Zero-Resource Turkic NLP
28
 
29
- [![Paper](https://img.shields.io/badge/Paper-ACM%20TALLIP-B31B1B.svg)](#)
30
- [![Framework](https://img.shields.io/badge/Framework-Apple%20MLX-000000.svg?logo=apple&logoColor=white)](https://github.com/ml-explore/mlx)
31
- [![Base Model](https://img.shields.io/badge/Base-Qwen2.5--1.5B--Instruct-blue.svg)](https://huggingface.co/Qwen/Qwen2.5-1.5B-Instruct)
32
- [![Dataset](https://img.shields.io/badge/Dataset-Nogai--Unified-green.svg)](https://huggingface.co/datasets/ansarzeinulla/Nogai-Unified-Corpus-v1)
33
 
34
- *A parameter-efficient low-rank adaptation establishing the mathematical baseline for Nogai morphological mapping.*
 
 
 
 
 
 
 
35
 
36
- </div>
37
 
38
- <br>
39
 
40
- ## 🔬 Model Overview
 
 
 
 
41
 
42
- This is a **Quantized LoRA (QLoRA) adapter** for `Qwen2.5-1.5B-Instruct`, engineered entirely on constrained consumer silicon via the Apple `mlx-lm` framework.
 
 
43
 
44
- This model represents **Phase 1** of the NogaiLLM adaptation pipeline. It was trained exclusively on the unstructured text of the `Nogai-Unified-Corpus-v1` (~4 million tokens) for 2,500 iterations. Its primary objective was to evaluate the resilience of dense Byte-Pair Encoding (BPE) architectures to zero-resource Cyrillic vocabulary expansion.
45
 
46
- ---
47
-
48
- ## ⚠️ Scientific Disclaimer: Catastrophic Forgetting
49
-
50
- **This adapter is released strictly for academic reproducibility and baseline measurement.** It is *not* a conversational agent.
51
-
52
- Because this adapter was trained via Continuous Pre-Training (CPT) on raw, unstructured journalistic text without an accompanying instruction-tuning phase, it exhibits severe **Catastrophic Forgetting**.
53
- * **The Success:** The multi-head attention layers successfully mapped the local Turkic morphology, dropping perplexity from a raw 191.13 to an optimized **12.14**. The model correctly generates Nogai-specific Cyrillic digraphs (e.g., `нъ`, `аь`, `оь`, `уь`).
54
- * **The Failure State:** The unstructured training completely overwrote the base model's latent instruction-following manifold. When presented with standard conversational templates (e.g., `ChatML`), the model's predictive routing collapses. It acts solely as an unconditional text generator, aggressively mimicking newspaper formatting rather than answering prompts.
55
-
56
- *(Note: To test the fully recovered, instruction-following model, please refer to the Phase 2 SFT model: [Nogai-SFT-Experimental](https://huggingface.co/ansarzeinulla/Qwen2.5-1.5B-Nogai-SFT-Experimental)).*
57
-
58
- ---
59
-
60
- ## 📊 Empirical Benchmarks & Mode Collapse
61
-
62
- This adapter was evaluated against Western-centric models in a rigorous 6-model architectural ablation study. The results prove that rigid, Latin-optimized tokenizers (Llama) cannot be force-mapped onto zero-resource Turkic languages via LoRA alone, whereas Eastern architectures (Qwen) achieve successful structural transfer.
63
-
64
- | Model Configuration | Scale | Test Loss ($\downarrow$) | Perplexity ($\downarrow$) | Nogai Typographic Density ($\uparrow$) |
65
- | :--- | :---: | :---: | :---: | :---: |
66
- | Qwen-2.5-0.5B (Raw Base) | 0.5B | 5.401 | 221.62 | 0.72% |
67
- | **Qwen-2.5-0.5B (Adapted)** | 0.5B | 3.379 | 29.33 | **22.95%** |
68
- | Qwen-2.5-1.5B (Raw Base) | 1.5B | 5.253 | 191.13 | 17.16% |
69
- | **Qwen-2.5-1.5B (Adapted)** | 1.5B | **2.497** | **12.14** | **15.05%*** |
70
- | Llama-3.2-1B (Raw Base) | 1.0B | 5.836 | 342.44 | 89.66%** |
71
- | Llama-3.2-1B (Adapted) | 1.0B | 5.505 | 245.82 | 0.00% *(Mode Collapse)* |
72
-
73
- *\*Represents morphologically accurate digraph placement, free of hallucination loops.*
74
- *\*\*Artifact: Llama base exhibited severe OOD fragmentation, falling into an autoregressive Cyrillic repetition loop.*
75
-
76
- ---
77
-
78
- ## 🚀 Usage (Apple Silicon / MLX)
79
 
80
- Because this model acts as an unconditional morphological generator rather than a chat model, **you must bypass the chat template** during generation.
81
-
82
- ### CLI Execution
83
- Install the MLX framework natively on macOS:
84
  ```bash
85
  pip install mlx-lm
86
- ```
87
-
88
- Run generation with the `--ignore-chat-template` flag to prevent format-induced hallucination:
89
- ```bash
90
- mlx_lm.generate \
91
- --model Qwen/Qwen2.5-1.5B-Instruct \
92
  --adapter-path ansarzeinulla/Qwen2.5-1.5B-Nogai-LoRA \
93
- --prompt "Буьгуьнги куьн" \
94
- --max-tokens 128 \
95
- --temp 0.5 \
96
- --ignore-chat-template
97
  ```
98
 
99
- ### Python API Integration
100
- ```python
101
- from mlx_lm import load, generate
102
-
103
- # Load the base model and dynamically apply the Phase 1 LoRA weights
104
- model, tokenizer = load(
105
- "Qwen/Qwen2.5-1.5B-Instruct",
106
- adapter_path="ansarzeinulla/Qwen2.5-1.5B-Nogai-LoRA"
107
- )
108
-
109
- # Unconditional text completion (Do NOT apply chat templates)
110
- prompt = "Ногай районынынъ орталыгы Терекли-Мектеб авлында"
111
-
112
- response = generate(
113
- model,
114
- tokenizer,
115
- prompt=prompt,
116
- max_tokens=100,
117
- verbose=True
118
- )
119
 
120
- print(response)
 
 
 
121
  ```
122
 
123
- ---
124
-
125
- ## 📚 Citation
126
-
127
- If you analyze this adapter's weights, utilize the ablation data, or study our mode collapse findings, please cite the core ACM TALLIP / arXiv paper:
128
 
129
  ```bibtex
130
- @article{zeinulla2026nogaillm,
131
- title={NogaiLLM: Parameter-Efficient Continuous Pre-Training and Architectures of Catastrophic Forgetting in Zero-Resource Turkic Languages},
132
- author={Zeinulla, Ansar},
133
- journal={arXiv preprint arXiv:2607.xxxxx},
134
- year={2026},
135
- publisher={Nazarbayev University / ACM TALLIP}
136
  }
137
  ```
138
-
139
- <div align="center">
140
- <i>Maintained by <a href="https://huggingface.co/ansarzeinulla">@ansarzeinulla</a> • Nazarbayev University</i>
141
- </div>
 
12
  - low-resource
13
  - nogai
14
  - turkic
15
+ - continued-pre-training
 
 
16
  datasets:
17
  - ansarzeinulla/Nogai-Unified-Corpus-v1
 
 
18
  ---
19
 
20
+ # Qwen2.5-1.5B-Nogai-LoRA (Phase 1: continued pre-training)
21
 
22
+ An `mlx-lm` LoRA adapter for Qwen2.5-1.5B-Instruct, trained on raw Nogai text ([Nogai-Unified-Corpus-v1](https://huggingface.co/datasets/ansarzeinulla/Nogai-Unified-Corpus-v1), 9.8 M Qwen2.5 tokens). This is Phase 1 of [NogaiLLM](https://github.com/ansarzeinulla/NogaiLLM-Apple-Silicon). It learns Nogai spelling and morphology, but after this phase the model **no longer follows chat instructions**: it continues text like a newspaper. For translation, use Phase 2 ([Qwen2.5-1.5B-Nogai-SFT-Experimental](https://huggingface.co/ansarzeinulla/Qwen2.5-1.5B-Nogai-SFT-Experimental)) on top of this adapter.
 
23
 
24
+ ## Training (from `adapter_config.json`)
 
 
 
25
 
26
+ | | |
27
+ |---|---|
28
+ | Base | MLX copy of Qwen2.5-1.5B-Instruct |
29
+ | LoRA | rank 8, scale 20, dropout 0 |
30
+ | Trained layers | q/k/v/o/gate/up/down projections of layers 12–27 (16 blocks), 5.28 M parameters |
31
+ | Batch / iterations / learning rate | 2 / 2,500 / 2e-4, seed 0 |
32
+ | Max sequence length | 512 tokens |
33
+ | Hardware | Apple M2 Pro, 16 GB, `mlx-lm` |
34
 
35
+ ## Results
36
 
37
+ Reported in the first version of the paper, on the corpus validation split:
38
 
39
+ | Model | Per-token perplexity | TD (digraphs / 100 words) |
40
+ |---|---|---|
41
+ | Qwen2.5-1.5B-Instruct (base) | 191.13 | 17.16 |
42
+ | + this adapter | 12.14 | 15.05 |
43
+ | Human Nogai text (reference) | — | 22.7 |
44
 
45
+ Notes:
46
+ - **TD** counts the Nogai digraphs аь/оь/уь/нъ per 100 words. It is only meaningful next to the human value (22.7). By that measure, this adapter's output is *further* from natural Nogai than the base model's.
47
+ - **Per-token perplexity can't be compared across tokenizers.** These numbers are being re-measured as bits per byte with [`eval_bits_per_byte.py`](https://github.com/ansarzeinulla/NogaiLLM-Apple-Silicon/blob/main/evaluation/eval_bits_per_byte.py).
48
 
49
+ ## Usage (Apple Silicon / MLX)
50
 
51
+ The model is a text continuer after this phase, so skip the chat template:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52
 
 
 
 
 
53
  ```bash
54
  pip install mlx-lm
55
+ mlx_lm.generate --model Qwen/Qwen2.5-1.5B-Instruct \
 
 
 
 
 
56
  --adapter-path ansarzeinulla/Qwen2.5-1.5B-Nogai-LoRA \
57
+ --prompt "Буьгуьнги куьн" --max-tokens 128 --temp 0.5 --ignore-chat-template
 
 
 
58
  ```
59
 
60
+ To build the base for Phase 2, fuse this adapter into the model:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
61
 
62
+ ```bash
63
+ mlx_lm.fuse --model Qwen/Qwen2.5-1.5B-Instruct \
64
+ --adapter-path ansarzeinulla/Qwen2.5-1.5B-Nogai-LoRA \
65
+ --save-path local_qwen_1.5B_Nogai_Base
66
  ```
67
 
68
+ ## Citation
 
 
 
 
69
 
70
  ```bibtex
71
+ @misc{zeinulla2026nogaillm,
72
+ title = {NogaiLLM: Parameter-Efficient Continued Pre-Training and Catastrophic Forgetting in Zero-Resource Turkic Languages},
73
+ author = {Zeinulla, Ansar},
74
+ year = {2026},
75
+ note = {Manuscript under revision. Code: https://github.com/ansarzeinulla/NogaiLLM-Apple-Silicon}
 
76
  }
77
  ```