Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -15,29 +15,69 @@ tags:
|
|
| 15 |
# az_unigram_32k — Azerbaijani SentencePiece tokenizer (32K, Unigram)
|
| 16 |
|
| 17 |
A purpose-built **32,000-vocab SentencePiece Unigram** tokenizer for **Latin-script Azerbaijani**.
|
| 18 |
-
On neutral, human-translated text it is **~2.4× more token-efficient than GPT-4** for Azerbaijani — i.e.
|
| 19 |
-
GPT-4 needs ~141% more tokens for the same text, which means higher cost and effectively less context.
|
| 20 |
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
|
| 37 |
We beat the massively-multilingual SentencePiece tokenizers (XLM-R, NLLB, mT5) **with ~8× smaller vocab**,
|
| 38 |
and the margin is *larger* on clean text than on web text — purpose-built morpheme segmentation shows most
|
| 39 |
-
where general tokenizers fragment Azerbaijani words.
|
| 40 |
-
|
| 41 |
|
| 42 |
## Design
|
| 43 |
|
|
@@ -63,27 +103,30 @@ print(len(ids), sp.decode(ids))
|
|
| 63 |
|
| 64 |
- **Latin Azerbaijani only** — Cyrillic (legacy) and Perso-Arabic (South Azerbaijani) are out of scope.
|
| 65 |
- Trained on web-heavy text; rare technical/scientific vocabulary may segment less efficiently.
|
|
|
|
|
|
|
| 66 |
- Fertility numbers are vs the listed tokenizer versions at eval time.
|
| 67 |
|
| 68 |
-
## FAQ
|
| 69 |
|
| 70 |
-
**
|
| 71 |
-
|
|
|
|
|
|
|
| 72 |
|
| 73 |
-
**Why
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
2
|
| 79 |
-
split across 100+ languages. Focused-and-small beats huge-and-diluted: we beat the 250K multilingual
|
| 80 |
-
tokenizers with **8× less vocab**.
|
| 81 |
|
| 82 |
-
**Is the benchmark fair?** Yes — measured on **FLORES-200** (human-translated;
|
| 83 |
-
trained on it specifically). Neutral ground, fully reproducible (`compare_fertility.py`).
|
|
|
|
| 84 |
|
| 85 |
-
**Honest caveat:** better *for Azerbaijani specifically* — not for multilingual text or code,
|
| 86 |
-
|
| 87 |
|
| 88 |
## Citation
|
| 89 |
|
|
|
|
| 15 |
# az_unigram_32k — Azerbaijani SentencePiece tokenizer (32K, Unigram)
|
| 16 |
|
| 17 |
A purpose-built **32,000-vocab SentencePiece Unigram** tokenizer for **Latin-script Azerbaijani**.
|
|
|
|
|
|
|
| 18 |
|
| 19 |
+
**On par with the best purpose-built Azerbaijani tokenizers at equal or smaller vocab — and ~2× more
|
| 20 |
+
token-efficient than the Turkish tokenizers people actually reuse for Azerbaijani.** Against
|
| 21 |
+
general-purpose tokenizers the gap is larger still: GPT-4 needs **~141% more tokens** for the same
|
| 22 |
+
Azerbaijani text (2.4×) — higher cost and effectively less usable context.
|
| 23 |
+
|
| 24 |
+

|
| 25 |
+
|
| 26 |
+
## Compared to other Azerbaijani & Turkic tokenizers (the real benchmark)
|
| 27 |
+
|
| 28 |
+
Beating GPT-4 on Azerbaijani is table stakes — *any* Azerbaijani-specific tokenizer does. The honest
|
| 29 |
+
question is how we stack up against tokenizers actually **built for**, or **reused on**, Azerbaijani.
|
| 30 |
+
Measured on **FLORES-200 Azerbaijani dev** (997 human-translated sentences — neutral for every tokenizer,
|
| 31 |
+
none trained on it; lower fertility is better):
|
| 32 |
+
|
| 33 |
+
| tokenizer | category | vocab | fertility (tok/word) | bytes/token |
|
| 34 |
+
|---|---|---:|---:|---:|
|
| 35 |
+
| LocalDoc az-en (Unigram) | Azerbaijani | 50,000 | 1.450 | 5.97 |
|
| 36 |
+
| aLLMA-2 (allmalab) | Azerbaijani | 32,000 | 1.522 | 5.68 |
|
| 37 |
+
| **az_unigram_32k (ours)** | **Azerbaijani** | **32,000** | **1.544** | **5.60** |
|
| 38 |
+
| XLM-R | Multilingual | 250,002 | 1.815 | 4.76 |
|
| 39 |
+
| NLLB-200 | Multilingual | 256,204 | 2.104 | 4.11 |
|
| 40 |
+
| mT5 | Multilingual | 250,100 | 2.362 | 3.66 |
|
| 41 |
+
| GPT-4o (o200k_base) | General-purpose | 200,019 | 2.383 | 3.63 |
|
| 42 |
+
| Turkish GPT-2 (ytu-cosmos) | Turkic / Az-tuned | 50,257 | 3.214 | 2.69 |
|
| 43 |
+
| BERTurk (Turkish, cased) | Turkic / Az-tuned | 32,000 | 3.223 | 2.68 |
|
| 44 |
+
| Llama 3 | General-purpose | 128,000 | 3.406 | 2.54 |
|
| 45 |
+
| mGPT-az (ai-forever) | Turkic / Az-tuned | 100,000 | 3.522 | 2.46 |
|
| 46 |
+
| Qwen2.5 | General-purpose | 151,643 | 3.562 | 2.43 |
|
| 47 |
+
| GPT-4 (cl100k_base) | General-purpose | 100,277 | 3.716 | 2.33 |
|
| 48 |
+
|
| 49 |
+
**Reading it honestly:**
|
| 50 |
+
|
| 51 |
+
- **Statistical tie with aLLMA-2 at equal 32K vocab** (1.544 vs 1.522 — a 1.4% gap). aLLMA-2 is the
|
| 52 |
+
tokenizer from the *"Open foundation models for Azerbaijani"* line of work (WordPiece/BPE, 32K).
|
| 53 |
+
- **LocalDoc's small edge is mostly its 56%-larger vocab** (50K vs 32K). Fertility drops with vocab; a
|
| 54 |
+
controlled sweep on this tokenizer showed 32K→48K ≈ −3.9%, so a 48–50K build of ours lands on top of it.
|
| 55 |
+
- **~2× ahead of both Turkish tokenizers** (BERTurk, Turkish GPT-2) and **2.3× ahead of Azerbaijani-adapted
|
| 56 |
+
mGPT**. Takeaway: you **cannot** reuse a Turkish tokenizer for Azerbaijani, and shipping an "Azerbaijani
|
| 57 |
+
model" on a stock multilingual tokenizer (like mGPT-az) leaves ~2× efficiency on the table.
|
| 58 |
+
- **A deliberate tradeoff explains the rest of the gap.** We set `split_digits=True` — every digit is its
|
| 59 |
+
own token, for cleaner numeric/arithmetic behavior — which aLLMA-2 and LocalDoc do **not**. On text with
|
| 60 |
+
numbers this costs us fertility *on purpose* (e.g. `2025` → 4 tokens, not 1). Back it out and ours **leads
|
| 61 |
+
the Azerbaijani field** (~1.67 vs aLLMA-2 1.77 / LocalDoc 1.73 on in-domain web text). Our subword modeling
|
| 62 |
+
is competitive-to-better; the fertility we spend is bought numeric behavior. All three preserve casing and
|
| 63 |
+
the dotted/dotless-i distinction — no one wins by lowercasing.
|
| 64 |
+
|
| 65 |
+
## vs. general-purpose tokenizers (the cost hook)
|
| 66 |
+
|
| 67 |
+
Same FLORES eval, expressed as "how many more tokens the general tokenizers need vs ours":
|
| 68 |
+
|
| 69 |
+
| tokenizer | vocab | fertility (tok/word) | vs ours |
|
| 70 |
+
|---|---:|---:|---:|
|
| 71 |
+
| **az_unigram_32k (ours)** | **32,000** | **1.544** | — |
|
| 72 |
+
| XLM-R | 250,002 | 1.815 | 1.18× |
|
| 73 |
+
| GPT-4o (o200k_base) | 200,019 | 2.383 | 1.54× |
|
| 74 |
+
| Llama 3 | 128,000 | 3.406 | 2.21× |
|
| 75 |
+
| GPT-4 (cl100k_base) | 100,277 | 3.716 | **2.41×** |
|
| 76 |
|
| 77 |
We beat the massively-multilingual SentencePiece tokenizers (XLM-R, NLLB, mT5) **with ~8× smaller vocab**,
|
| 78 |
and the margin is *larger* on clean text than on web text — purpose-built morpheme segmentation shows most
|
| 79 |
+
where general tokenizers fragment Azerbaijani words. Reproduce everything with
|
| 80 |
+
`tokenizer/compare_fertility.py` (add `--include-az` for the Azerbaijani/Turkic set — on by default).
|
| 81 |
|
| 82 |
## Design
|
| 83 |
|
|
|
|
| 103 |
|
| 104 |
- **Latin Azerbaijani only** — Cyrillic (legacy) and Perso-Arabic (South Azerbaijani) are out of scope.
|
| 105 |
- Trained on web-heavy text; rare technical/scientific vocabulary may segment less efficiently.
|
| 106 |
+
- `split_digits=True` trades raw fertility on number-heavy text for cleaner numeric behavior (a deliberate
|
| 107 |
+
choice — see the comparison above).
|
| 108 |
- Fertility numbers are vs the listed tokenizer versions at eval time.
|
| 109 |
|
| 110 |
+
## FAQ
|
| 111 |
|
| 112 |
+
**How does this compare to other Azerbaijani tokenizers (not just GPT-4)?** It's a statistical tie with
|
| 113 |
+
aLLMA-2 at the same 32K vocab, within reach of LocalDoc's 50K (the difference is mostly its larger vocab),
|
| 114 |
+
and ~2× more efficient than Turkish tokenizers (BERTurk, Turkish GPT-2) or an Azerbaijani-tuned model that
|
| 115 |
+
kept a multilingual tokenizer (mGPT-az). See the table above.
|
| 116 |
|
| 117 |
+
**Why is it better than GPT-4's tokenizer?** It's purpose-built for Azerbaijani, so it encodes the same
|
| 118 |
+
text in ~2.4× fewer tokens → cheaper, more text per context window, faster. General tokenizers (GPT-4,
|
| 119 |
+
Llama, Qwen) are trained mostly on English/Chinese; Azerbaijani is a rounding error in their data, so they
|
| 120 |
+
**shatter words into tiny fragments** — sometimes raw bytes for ə/ğ/ı/ş/ç/ö/ü. Two mechanisms: (1)
|
| 121 |
+
Azerbaijani is agglutinative (*ev → evlər → evlərimizdə*) and we learned those stems+suffixes as units;
|
| 122 |
+
(2) our 32K vocab is spent *entirely* on Azerbaijani vs GPT-4's 100K split across 100+ languages.
|
|
|
|
|
|
|
| 123 |
|
| 124 |
+
**Is the benchmark fair?** Yes — measured on **FLORES-200** (human-translated; none of the tokenizers
|
| 125 |
+
trained on it specifically). Neutral ground, fully reproducible (`compare_fertility.py`). Every tokenizer
|
| 126 |
+
is run with `add_special_tokens=False`, and all Azerbaijani comparators preserve casing.
|
| 127 |
|
| 128 |
+
**Honest caveat:** better *for Azerbaijani specifically* — not for multilingual text or code, Latin-script
|
| 129 |
+
only.
|
| 130 |
|
| 131 |
## Citation
|
| 132 |
|