--- language: - uk license: mit library_name: transformers pipeline_tag: text2text-generation tags: - marian - text-normalization - verbalization - text-to-speech - ukrainian datasets: - skypro1111/uk-text-normalization metrics: - exact_match - wer - cer model-index: - name: marian-uk-verbalizer results: - task: type: text2text-generation name: Ukrainian text verbalization dataset: type: skypro1111/uk-text-normalization name: frozen 3,597-row benchmark slice metrics: - type: exact_match value: 85.21 name: Exact match - type: wer value: 2.98 name: Word error rate - type: cer value: 1.20 name: Character error rate --- # marian-uk-verbalizer A 52.7M-parameter Marian sequence-to-sequence model that converts written Ukrainian text into the form it should be **spoken** — the text-normalization front end for a Ukrainian TTS pipeline. It expands Roman numerals (with correct case/gender agreement), dates, currencies, coordinates, phone numbers, IBANs, card numbers, math expressions, symbols, abbreviations and acronyms into spoken words, while leaving ordinary prose untouched. ``` Потрібно прочитати XX розділ до понеділка. -> Потрібно прочитати двадцятий розділ до понеділка. Я пропрацював у ФБР двадцять років. -> Я пропрацював у еф-бе-ер двадцять років. ``` Full training/evaluation code, the router this model is meant to sit behind, and per-domain evaluation scripts: **[GitHub repository](https://github.com/vacuum-in/marian-uk-verbalizer)**. A quantized int8 CTranslate2 export of this same checkpoint (52 MB, ~99.1% output-identical) is published separately at [`aloudreader/marian-uk-verbalizer-ct2-int8`](https://huggingface.co/aloudreader/marian-uk-verbalizer-ct2-int8). ## Model description - Architecture: MarianMTModel, 6 encoder + 6 decoder layers, `d_model=512`, 8 attention heads, 2048 FFN dim — 52.7M parameters. - Tokenizer: shared 16k-token SentencePiece vocabulary (source and target). - Decoder start token is `` (token id 0), not the usual pad token — this matters if you re-export the model (see the GitHub repo's CTranslate2 conversion notes). - Trained with `input_preprocessing: model-routing-v1`: the model always receives the **original, unmodified sentence** and its output is used **raw** — no rule-based pre- or post-processing anywhere in the pipeline. This is a hard project constraint, not an implementation detail: see `AGENT.md` in the GitHub repository. ## Intended use Sits behind a lightweight shape-based router that sends plain Cyrillic prose straight through and everything else (digits, Latin letters, symbols, all-caps runs, …) to this model. Feeding the model plain prose it doesn't need to touch is safe — the training data includes plenty of identity (unchanged) targets — but the router avoids the extra inference cost. The router implementation (regex/shape only, no word lists) is in the GitHub repository at `src/verbalizer/text.py`. This model is Ukrainian-specific and was trained and evaluated on general prose plus template-generated numeric/structured text. It is not intended for other languages or for tasks other than pre-TTS text normalization. ## How to use ```python from transformers import MarianMTModel, MarianTokenizer tokenizer = MarianTokenizer.from_pretrained("aloudreader/marian-uk-verbalizer") model = MarianMTModel.from_pretrained("aloudreader/marian-uk-verbalizer") text = "Потрібно прочитати XX розділ до понеділка." batch = tokenizer([text], return_tensors="pt", padding=True) out = model.generate(**batch, max_new_tokens=128, num_beams=1) print(tokenizer.batch_decode(out, skip_special_tokens=True)[0]) ``` Greedy decoding (`num_beams=1`) is what this checkpoint was evaluated with; beam search is supported but was not found to improve the benchmark scores enough to justify the extra latency. ## Training data - The full [`skypro1111/uk-text-normalization`](https://huggingface.co/datasets/skypro1111/uk-text-normalization) dataset (a 3,597-row slice is frozen out as an evaluation-only benchmark and never trained on). - Synthetic rows generated offline with `num2words` and `pymorphy3` for domains and grammatical cases the base dataset covers thinly — most recently, Roman-numeral-plus-noun case/gender agreement across many governing verbs, prepositions and nouns per case, so the model learns to read the *noun's* ending rather than memorizing one preposition per case. Every generated target is cross-checked by parsing it back to a numeric value and comparing it against the source before it enters training. No text extracted from books is included in or was used to build this model's public evaluation data (see the GitHub repository's data policy). ## Evaluation Frozen 3,597-row benchmark (held out of training, no data leakage): | Metric | Score | |---|---| | Exact string match | 85.2% | | Word error rate | 3.0% | | Character error rate | 1.2% | | Numeric value fidelity (every digit sequence read back and compared to source) | 100.0% | | Digit sequence fidelity | 100.0% | | Identity accuracy (plain-prose rows left untouched) | 100.0% | Generated per-domain evaluation (lenient frame+slot scoring — the fixed part of each template sentence must match exactly, and each slot must be read as one of its accepted spoken forms): | Domain | Correct | Rows | |---|---|---| | Coordinates | 100.0% | 500 | | Phone numbers | 100.0% | 500 | | Roman numeral + noun case agreement | 99.0% | 301 | | Context (held-out book sentences, not published) | 97.4% | 500 | | English words embedded in Ukrainian | 88.4% | 500 | | Card numbers / IBAN | 86.8% | 500 | | Acronyms (letter-by-letter reading) | 85.5% | 207 | ## Limitations - Acronym and card/IBAN reading are the weakest domains (~85–87%); most failures are single-letter substitutions in acronym spelling or a single misread digit group in long card numbers. - The model has only ever seen Roman numerals I–XXXIX (by project convention, `L`/`C`/`D`/`M` are treated as ordinary letters — vitamin C, size L, flight D12 — since they collide with common non-numeral usage in Ukrainian text). - Like any seq2seq model, it can occasionally hallucinate or drop a token on out-of-distribution input; there is no rule-based safety net downstream — by design, its raw output is what gets spoken. ## License MIT. See the [GitHub repository](https://github.com/vacuum-in/marian-uk-verbalizer) for training/evaluation code under the same license. The `skypro1111/uk-text-normalization` training dataset has its own license — check it before redistributing data derived from this model.