YAML Metadata Warning:The pipeline tag "text2text-generation" is not in the official list: text-classification, token-classification, table-question-answering, question-answering, zero-shot-classification, translation, summarization, feature-extraction, text-generation, fill-mask, sentence-similarity, text-to-speech, text-to-audio, automatic-speech-recognition, audio-to-audio, audio-classification, audio-text-to-text, voice-activity-detection, depth-estimation, image-classification, object-detection, image-segmentation, text-to-image, image-to-text, image-to-image, image-to-video, unconditional-image-generation, video-classification, reinforcement-learning, robotics, tabular-classification, tabular-regression, tabular-to-text, table-to-text, multiple-choice, text-ranking, text-retrieval, time-series-forecasting, text-to-video, image-text-to-text, image-text-to-image, image-text-to-video, visual-question-answering, document-question-answering, zero-shot-image-classification, graph-ml, mask-generation, zero-shot-object-detection, text-to-3d, image-to-3d, image-feature-extraction, video-text-to-text, keypoint-detection, visual-document-retrieval, any-to-any, video-to-video, other
byt5-mitra-bo-tagger
A byte-level seq2seq tagger for Classical Tibetan in Wylie transliteration. Given a block of raw Wylie it returns, in one terse string:
- sentence segmentation (one sentence per line),
- dictionary-lexeme word segmentation
- a coarse part of speech for every word (noun / verb / verbal noun / adjective / particle …), and
- Sanskrit-unit brackets: runs of Tibetan words that render one Sanskrit word or compound
(
[chos/N kyi/C rgyal_po/N]= dharmarāja).
It is a fine-tune of buddhist-nlp/byt5-mitra-bo (ByT5 continued-pretrained on Tibetan Wylie) on ~5k sentences and short paragraphs from the Sanskrit–Tibetan parallel corpus of the Dharmamitra project.
Output scheme
| element | encoding | example |
|---|---|---|
| sentence | one per line | |
| word | syllables joined by _ |
byang_chub_sems_dpa'/N |
affix written attached in the input ('i, r, s, 'o, 'am …) |
own token with leading + |
dpa'i → dpa'/N +'i/C |
| part of speech | / + one letter |
see below |
| Sanskrit unit | [ … ] around the words rendering one Sanskrit word/compound |
[chos/N kyi/C rgyal_po/N] |
| shad | |/p (single), ||/p (double) |
POS letters: N noun · P proper noun · V verb (finite / predicate stem) · G verbal noun · J adjective · A adverb / fixed adverbial · M numeral · R pronoun · D determiner (demonstrative, plural, indefinite) · X negation · C case particle · K clause particle (after verbs) · L clitic / final particle · I interjection · p punctuation.
input : de nas tshe dang ldan pa kun dga' bos bcom ldan 'das la 'di skad ces gsol to/ /
output: de_nas/A [tshe_dang_ldan_pa/J] [kun_dga'_bo/P] +s/C [bcom_ldan_'das/N] la/C 'di_skad/A ces/L [gsol/V] to/L ||/p
Usage
from transformers import ByT5Tokenizer, AutoModelForSeq2SeqLM
tok = ByT5Tokenizer.from_pretrained("buddhist-nlp/byt5-mitra-bo-tagger")
model = AutoModelForSeq2SeqLM.from_pretrained("buddhist-nlp/byt5-mitra-bo-tagger")
x = tok(["de nas tshe dang ldan pa kun dga' bos bcom ldan 'das la 'di skad ces gsol to/ /"], return_tensors="pt")
print(tok.decode(model.generate(**x, max_new_tokens=1100)[0], skip_special_tokens=True))
A CTranslate2 int8 conversion is in ctranslate2-int8/ (same outputs within noise, ~30 sentences/s on one GPU):
import ctranslate2, transformers
tok = transformers.ByT5Tokenizer.from_pretrained("buddhist-nlp/byt5-mitra-bo-tagger")
tr = ctranslate2.Translator("<local path>/ctranslate2-int8", device="cuda", compute_type="int8_float16")
toks = [tok.convert_ids_to_tokens(tok(s)["input_ids"]) for s in texts]
out = tr.translate_batch(toks, max_decoding_length=1100, beam_size=1)
Inputs: Wylie (EWTS) as found in ACIP / DharmaNexus e-texts, up to ~700 characters per call (one sentence to a short paragraph). Outputs are at most ~1,000 characters.
Evaluation
Held-out 100 items:
| model | valid output | word boundary F1 | POS acc | verb/noun acc | sentence F1 | Sanskrit-unit F1 |
|---|---|---|---|---|---|---|
| byt5-mitra-bo-tagger (this model, 10 epochs) | .90 | .959 | .914 | .914 | .914 | .663 |
| Qwen3.5-9B mitra base, same data | 1.00 | .962 | .942 | .909 | .891 | .687 |
Limitations
- Trained on canonical translation literature (Kangyur / Tengyur texts with Sanskrit parallels); other registers are untested.
- The Sanskrit-unit layer is the noisiest: the training references bracket units inconsistently.
- The model can drop or alter a syllable in long or unusual spans; always validate that the output reproduces the input and fall back to a retry.
- Tense classes of verbs are not annotated (only the coarse V / G / N distinction).
- Downloads last month
- 295
Model tree for buddhist-nlp/byt5-mitra-bo-tagger
Base model
buddhist-nlp/byt5-mitra-bo