kamaalg commited on
Commit
398d887
·
verified ·
1 Parent(s): e755b9b

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +77 -34
README.md CHANGED
@@ -15,29 +15,69 @@ tags:
15
  # az_unigram_32k — Azerbaijani SentencePiece tokenizer (32K, Unigram)
16
 
17
  A purpose-built **32,000-vocab SentencePiece Unigram** tokenizer for **Latin-script Azerbaijani**.
18
- On neutral, human-translated text it is **~2.4× more token-efficient than GPT-4** for Azerbaijani — i.e.
19
- GPT-4 needs ~141% more tokens for the same text, which means higher cost and effectively less context.
20
 
21
- ## Headline: fertility (tokens/word — lower is better)
22
-
23
- Measured on **FLORES-200 Azerbaijani dev** (997 human-translated sentences, *not* in training — neutral
24
- for every tokenizer compared):
25
-
26
- | tokenizer | vocab | fertility (tok/word) | bytes/token | vs ours |
27
- |---|---:|---:|---:|---:|
28
- | **az_unigram_32k (ours)** | **32,000** | **1.544** | **5.603** | — |
29
- | XLM-R | 250,002 | 1.815 | 4.764 | 1.18× |
30
- | NLLB-200 | 256,204 | 2.104 | 4.111 | 1.36× |
31
- | mT5 | 250,100 | 2.362 | 3.660 | 1.53× |
32
- | GPT-4o (o200k_base) | 200,019 | 2.383 | 3.629 | 1.54× |
33
- | Llama 3 | 128,000 | 3.406 | 2.539 | 2.21× |
34
- | Qwen2.5 | 151,643 | 3.562 | 2.428 | 2.31× |
35
- | GPT-4 (cl100k_base) | 100,277 | 3.716 | 2.327 | **2.41×** |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
 
37
  We beat the massively-multilingual SentencePiece tokenizers (XLM-R, NLLB, mT5) **with ~8× smaller vocab**,
38
  and the margin is *larger* on clean text than on web text — purpose-built morpheme segmentation shows most
39
- where general tokenizers fragment Azerbaijani words. (On in-domain web text: ours 1.881 vs GPT-4 3.622.)
40
- Reproduce with `tokenizer/compare_fertility.py`.
41
 
42
  ## Design
43
 
@@ -63,27 +103,30 @@ print(len(ids), sp.decode(ids))
63
 
64
  - **Latin Azerbaijani only** — Cyrillic (legacy) and Perso-Arabic (South Azerbaijani) are out of scope.
65
  - Trained on web-heavy text; rare technical/scientific vocabulary may segment less efficiently.
 
 
66
  - Fertility numbers are vs the listed tokenizer versions at eval time.
67
 
68
- ## FAQ — why is this better than GPT-4's tokenizer?
69
 
70
- **Short version:** it's purpose-built for Azerbaijani, so it encodes the same text in ~2.4× fewer tokens
71
- than GPT-4 → cheaper, more text per context window, faster.
 
 
72
 
73
- **Why (the mechanism):** general tokenizers (GPT-4, Llama, Qwen) are trained mostly on English/Chinese;
74
- Azerbaijani is a rounding error in their data, so they never learned Azerbaijani subwords and **shatter
75
- words into tiny fragments** — sometimes raw bytes for ə/ğ/ı/ş/ç/ö/ü. Two concrete reasons we win:
76
- 1. **Agglutinative morphology** — Azerbaijani stacks suffixes (*ev → evlər → evlərimizdə*); we learned those
77
- stems+suffixes as units, general tokenizers fragment every word.
78
- 2. **Vocabulary budget** — our 32K vocab is spent *entirely* on Azerbaijani, vs GPT-4's 100K / XLM-R's 250K
79
- split across 100+ languages. Focused-and-small beats huge-and-diluted: we beat the 250K multilingual
80
- tokenizers with **8× less vocab**.
81
 
82
- **Is the benchmark fair?** Yes — measured on **FLORES-200** (human-translated; *none* of the tokenizers
83
- trained on it specifically). Neutral ground, fully reproducible (`compare_fertility.py`).
 
84
 
85
- **Honest caveat:** better *for Azerbaijani specifically* — not for multilingual text or code, and
86
- Latin-script only.
87
 
88
  ## Citation
89
 
 
15
  # az_unigram_32k — Azerbaijani SentencePiece tokenizer (32K, Unigram)
16
 
17
  A purpose-built **32,000-vocab SentencePiece Unigram** tokenizer for **Latin-script Azerbaijani**.
 
 
18
 
19
+ **On par with the best purpose-built Azerbaijani tokenizers at equal or smaller vocab — and ~2× more
20
+ token-efficient than the Turkish tokenizers people actually reuse for Azerbaijani.** Against
21
+ general-purpose tokenizers the gap is larger still: GPT-4 needs **~141% more tokens** for the same
22
+ Azerbaijani text (2.4×) — higher cost and effectively less usable context.
23
+
24
+ ![Azerbaijani tokenizer comparison](az_tokenizer_comparison.png)
25
+
26
+ ## Compared to other Azerbaijani & Turkic tokenizers (the real benchmark)
27
+
28
+ Beating GPT-4 on Azerbaijani is table stakes — *any* Azerbaijani-specific tokenizer does. The honest
29
+ question is how we stack up against tokenizers actually **built for**, or **reused on**, Azerbaijani.
30
+ Measured on **FLORES-200 Azerbaijani dev** (997 human-translated sentences — neutral for every tokenizer,
31
+ none trained on it; lower fertility is better):
32
+
33
+ | tokenizer | category | vocab | fertility (tok/word) | bytes/token |
34
+ |---|---|---:|---:|---:|
35
+ | LocalDoc az-en (Unigram) | Azerbaijani | 50,000 | 1.450 | 5.97 |
36
+ | aLLMA-2 (allmalab) | Azerbaijani | 32,000 | 1.522 | 5.68 |
37
+ | **az_unigram_32k (ours)** | **Azerbaijani** | **32,000** | **1.544** | **5.60** |
38
+ | XLM-R | Multilingual | 250,002 | 1.815 | 4.76 |
39
+ | NLLB-200 | Multilingual | 256,204 | 2.104 | 4.11 |
40
+ | mT5 | Multilingual | 250,100 | 2.362 | 3.66 |
41
+ | GPT-4o (o200k_base) | General-purpose | 200,019 | 2.383 | 3.63 |
42
+ | Turkish GPT-2 (ytu-cosmos) | Turkic / Az-tuned | 50,257 | 3.214 | 2.69 |
43
+ | BERTurk (Turkish, cased) | Turkic / Az-tuned | 32,000 | 3.223 | 2.68 |
44
+ | Llama 3 | General-purpose | 128,000 | 3.406 | 2.54 |
45
+ | mGPT-az (ai-forever) | Turkic / Az-tuned | 100,000 | 3.522 | 2.46 |
46
+ | Qwen2.5 | General-purpose | 151,643 | 3.562 | 2.43 |
47
+ | GPT-4 (cl100k_base) | General-purpose | 100,277 | 3.716 | 2.33 |
48
+
49
+ **Reading it honestly:**
50
+
51
+ - **Statistical tie with aLLMA-2 at equal 32K vocab** (1.544 vs 1.522 — a 1.4% gap). aLLMA-2 is the
52
+ tokenizer from the *"Open foundation models for Azerbaijani"* line of work (WordPiece/BPE, 32K).
53
+ - **LocalDoc's small edge is mostly its 56%-larger vocab** (50K vs 32K). Fertility drops with vocab; a
54
+ controlled sweep on this tokenizer showed 32K→48K ≈ −3.9%, so a 48–50K build of ours lands on top of it.
55
+ - **~2× ahead of both Turkish tokenizers** (BERTurk, Turkish GPT-2) and **2.3× ahead of Azerbaijani-adapted
56
+ mGPT**. Takeaway: you **cannot** reuse a Turkish tokenizer for Azerbaijani, and shipping an "Azerbaijani
57
+ model" on a stock multilingual tokenizer (like mGPT-az) leaves ~2× efficiency on the table.
58
+ - **A deliberate tradeoff explains the rest of the gap.** We set `split_digits=True` — every digit is its
59
+ own token, for cleaner numeric/arithmetic behavior — which aLLMA-2 and LocalDoc do **not**. On text with
60
+ numbers this costs us fertility *on purpose* (e.g. `2025` → 4 tokens, not 1). Back it out and ours **leads
61
+ the Azerbaijani field** (~1.67 vs aLLMA-2 1.77 / LocalDoc 1.73 on in-domain web text). Our subword modeling
62
+ is competitive-to-better; the fertility we spend is bought numeric behavior. All three preserve casing and
63
+ the dotted/dotless-i distinction — no one wins by lowercasing.
64
+
65
+ ## vs. general-purpose tokenizers (the cost hook)
66
+
67
+ Same FLORES eval, expressed as "how many more tokens the general tokenizers need vs ours":
68
+
69
+ | tokenizer | vocab | fertility (tok/word) | vs ours |
70
+ |---|---:|---:|---:|
71
+ | **az_unigram_32k (ours)** | **32,000** | **1.544** | — |
72
+ | XLM-R | 250,002 | 1.815 | 1.18× |
73
+ | GPT-4o (o200k_base) | 200,019 | 2.383 | 1.54× |
74
+ | Llama 3 | 128,000 | 3.406 | 2.21× |
75
+ | GPT-4 (cl100k_base) | 100,277 | 3.716 | **2.41×** |
76
 
77
  We beat the massively-multilingual SentencePiece tokenizers (XLM-R, NLLB, mT5) **with ~8× smaller vocab**,
78
  and the margin is *larger* on clean text than on web text — purpose-built morpheme segmentation shows most
79
+ where general tokenizers fragment Azerbaijani words. Reproduce everything with
80
+ `tokenizer/compare_fertility.py` (add `--include-az` for the Azerbaijani/Turkic set — on by default).
81
 
82
  ## Design
83
 
 
103
 
104
  - **Latin Azerbaijani only** — Cyrillic (legacy) and Perso-Arabic (South Azerbaijani) are out of scope.
105
  - Trained on web-heavy text; rare technical/scientific vocabulary may segment less efficiently.
106
+ - `split_digits=True` trades raw fertility on number-heavy text for cleaner numeric behavior (a deliberate
107
+ choice — see the comparison above).
108
  - Fertility numbers are vs the listed tokenizer versions at eval time.
109
 
110
+ ## FAQ
111
 
112
+ **How does this compare to other Azerbaijani tokenizers (not just GPT-4)?** It's a statistical tie with
113
+ aLLMA-2 at the same 32K vocab, within reach of LocalDoc's 50K (the difference is mostly its larger vocab),
114
+ and ~2× more efficient than Turkish tokenizers (BERTurk, Turkish GPT-2) or an Azerbaijani-tuned model that
115
+ kept a multilingual tokenizer (mGPT-az). See the table above.
116
 
117
+ **Why is it better than GPT-4's tokenizer?** It's purpose-built for Azerbaijani, so it encodes the same
118
+ text in ~2.4× fewer tokens → cheaper, more text per context window, faster. General tokenizers (GPT-4,
119
+ Llama, Qwen) are trained mostly on English/Chinese; Azerbaijani is a rounding error in their data, so they
120
+ **shatter words into tiny fragments** — sometimes raw bytes for ə/ğ/ı/ş/ç/ö/ü. Two mechanisms: (1)
121
+ Azerbaijani is agglutinative (*ev → evlər → evlərimizdə*) and we learned those stems+suffixes as units;
122
+ (2) our 32K vocab is spent *entirely* on Azerbaijani vs GPT-4's 100K split across 100+ languages.
 
 
123
 
124
+ **Is the benchmark fair?** Yes — measured on **FLORES-200** (human-translated; none of the tokenizers
125
+ trained on it specifically). Neutral ground, fully reproducible (`compare_fertility.py`). Every tokenizer
126
+ is run with `add_special_tokens=False`, and all Azerbaijani comparators preserve casing.
127
 
128
+ **Honest caveat:** better *for Azerbaijani specifically* — not for multilingual text or code, Latin-script
129
+ only.
130
 
131
  ## Citation
132