File size: 3,009 Bytes
74e4753
 
c4c02aa
2797004
 
 
 
c4c02aa
 
 
 
 
 
 
2eb2260
c4c02aa
 
74e4753
c4c02aa
2eb2260
2797004
2eb2260
2797004
2eb2260
2797004
2eb2260
 
 
 
 
 
 
 
2797004
2eb2260
2797004
2eb2260
2797004
2eb2260
 
 
 
 
2797004
2eb2260
 
 
2797004
2eb2260
2797004
2eb2260
c4c02aa
 
 
2eb2260
c4c02aa
2eb2260
2797004
 
2eb2260
2797004
2eb2260
 
 
 
2797004
 
2eb2260
2797004
 
2eb2260
 
 
 
 
2797004
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
---
license: apache-2.0
base_model: Qwen/Qwen2.5-1.5B-Instruct
pipeline_tag: text-generation
library_name: mlx
language:
- nog
tags:
- mlx
- lora
- peft
- low-resource
- nogai
- turkic
- continued-pre-training
datasets:
- ansarzeinulla/Nogai-Unified-Corpus-v1
---

# Qwen2.5-1.5B-Nogai-LoRA (Phase 1: continued pre-training)

An `mlx-lm` LoRA adapter for Qwen2.5-1.5B-Instruct, trained on raw Nogai text ([Nogai-Unified-Corpus-v1](https://huggingface.co/datasets/ansarzeinulla/Nogai-Unified-Corpus-v1), 9.8 M Qwen2.5 tokens). This is Phase 1 of [NogaiLLM](https://github.com/ansarzeinulla/NogaiLLM-Apple-Silicon). It learns Nogai spelling and morphology, but after this phase the model **no longer follows chat instructions**: it continues text like a newspaper. For translation, use Phase 2 ([Qwen2.5-1.5B-Nogai-SFT-Experimental](https://huggingface.co/ansarzeinulla/Qwen2.5-1.5B-Nogai-SFT-Experimental)) on top of this adapter.

## Training (from `adapter_config.json`)

| | |
|---|---|
| Base | MLX copy of Qwen2.5-1.5B-Instruct |
| LoRA | rank 8, scale 20, dropout 0 |
| Trained layers | q/k/v/o/gate/up/down projections of layers 12–27 (16 blocks), 5.28 M parameters |
| Batch / iterations / learning rate | 2 / 2,500 / 2e-4, seed 0 |
| Max sequence length | 512 tokens |
| Hardware | Apple M2 Pro, 16 GB, `mlx-lm` |

## Results

Reported in the first version of the paper, on the corpus validation split:

| Model | Per-token perplexity | TD (digraphs / 100 words) |
|---|---|---|
| Qwen2.5-1.5B-Instruct (base) | 191.13 | 17.16 |
| + this adapter | 12.14 | 15.05 |
| Human Nogai text (reference) | — | 22.7 |

Notes:
- **TD** counts the Nogai digraphs аь/оь/уь/нъ per 100 words. It is only meaningful next to the human value (22.7). By that measure, this adapter's output is *further* from natural Nogai than the base model's.
- **Per-token perplexity can't be compared across tokenizers.** These numbers are being re-measured as bits per byte with [`eval_bits_per_byte.py`](https://github.com/ansarzeinulla/NogaiLLM-Apple-Silicon/blob/main/evaluation/eval_bits_per_byte.py).

## Usage (Apple Silicon / MLX)

The model is a text continuer after this phase, so skip the chat template:

```bash
pip install mlx-lm
mlx_lm.generate --model Qwen/Qwen2.5-1.5B-Instruct \
    --adapter-path ansarzeinulla/Qwen2.5-1.5B-Nogai-LoRA \
    --prompt "Буьгуьнги куьн" --max-tokens 128 --temp 0.5 --ignore-chat-template
```

To build the base for Phase 2, fuse this adapter into the model:

```bash
mlx_lm.fuse --model Qwen/Qwen2.5-1.5B-Instruct \
    --adapter-path ansarzeinulla/Qwen2.5-1.5B-Nogai-LoRA \
    --save-path local_qwen_1.5B_Nogai_Base
```

## Citation

```bibtex
@misc{zeinulla2026nogaillm,
  title  = {NogaiLLM: Parameter-Efficient Continued Pre-Training and Catastrophic Forgetting in Zero-Resource Turkic Languages},
  author = {Zeinulla, Ansar},
  year   = {2026},
  note   = {Manuscript under revision. Code: https://github.com/ansarzeinulla/NogaiLLM-Apple-Silicon}
}
```