YAML Metadata Warning:The pipeline tag "text2text-generation" is not in the official list: text-classification, token-classification, table-question-answering, question-answering, zero-shot-classification, translation, summarization, feature-extraction, text-generation, fill-mask, sentence-similarity, text-to-speech, text-to-audio, automatic-speech-recognition, audio-to-audio, audio-classification, audio-text-to-text, voice-activity-detection, depth-estimation, image-classification, object-detection, image-segmentation, text-to-image, image-to-text, image-to-image, image-to-video, unconditional-image-generation, video-classification, reinforcement-learning, robotics, tabular-classification, tabular-regression, tabular-to-text, table-to-text, multiple-choice, text-ranking, text-retrieval, time-series-forecasting, text-to-video, image-text-to-text, image-text-to-image, image-text-to-video, visual-question-answering, document-question-answering, zero-shot-image-classification, graph-ml, mask-generation, zero-shot-object-detection, text-to-3d, image-to-3d, image-feature-extraction, video-text-to-text, keypoint-detection, visual-document-retrieval, any-to-any, video-to-video, other

byt5-mitra-bo-tagger

A byte-level seq2seq tagger for Classical Tibetan in Wylie transliteration. Given a block of raw Wylie it returns, in one terse string:

  1. sentence segmentation (one sentence per line),
  2. dictionary-lexeme word segmentation
  3. a coarse part of speech for every word (noun / verb / verbal noun / adjective / particle …), and
  4. Sanskrit-unit brackets: runs of Tibetan words that render one Sanskrit word or compound ([chos/N kyi/C rgyal_po/N] = dharmarāja).

It is a fine-tune of buddhist-nlp/byt5-mitra-bo (ByT5 continued-pretrained on Tibetan Wylie) on ~5k sentences and short paragraphs from the Sanskrit–Tibetan parallel corpus of the Dharmamitra project.

Output scheme

element encoding example
sentence one per line
word syllables joined by _ byang_chub_sems_dpa'/N
affix written attached in the input ('i, r, s, 'o, 'am …) own token with leading + dpa'i → dpa'/N +'i/C
part of speech / + one letter see below
Sanskrit unit [ … ] around the words rendering one Sanskrit word/compound [chos/N kyi/C rgyal_po/N]
shad |/p (single), ||/p (double)

POS letters: N noun · P proper noun · V verb (finite / predicate stem) · G verbal noun · J adjective · A adverb / fixed adverbial · M numeral · R pronoun · D determiner (demonstrative, plural, indefinite) · X negation · C case particle · K clause particle (after verbs) · L clitic / final particle · I interjection · p punctuation.

input : de nas tshe dang ldan pa kun dga' bos bcom ldan 'das la 'di skad ces gsol to/ /
output: de_nas/A [tshe_dang_ldan_pa/J] [kun_dga'_bo/P] +s/C [bcom_ldan_'das/N] la/C 'di_skad/A ces/L [gsol/V] to/L ||/p

Usage

from transformers import ByT5Tokenizer, AutoModelForSeq2SeqLM
tok = ByT5Tokenizer.from_pretrained("buddhist-nlp/byt5-mitra-bo-tagger")
model = AutoModelForSeq2SeqLM.from_pretrained("buddhist-nlp/byt5-mitra-bo-tagger")
x = tok(["de nas tshe dang ldan pa kun dga' bos bcom ldan 'das la 'di skad ces gsol to/ /"], return_tensors="pt")
print(tok.decode(model.generate(**x, max_new_tokens=1100)[0], skip_special_tokens=True))

A CTranslate2 int8 conversion is in ctranslate2-int8/ (same outputs within noise, ~30 sentences/s on one GPU):

import ctranslate2, transformers
tok = transformers.ByT5Tokenizer.from_pretrained("buddhist-nlp/byt5-mitra-bo-tagger")
tr = ctranslate2.Translator("<local path>/ctranslate2-int8", device="cuda", compute_type="int8_float16")
toks = [tok.convert_ids_to_tokens(tok(s)["input_ids"]) for s in texts]
out = tr.translate_batch(toks, max_decoding_length=1100, beam_size=1)

Inputs: Wylie (EWTS) as found in ACIP / DharmaNexus e-texts, up to ~700 characters per call (one sentence to a short paragraph). Outputs are at most ~1,000 characters.

Evaluation

Held-out 100 items:

model valid output word boundary F1 POS acc verb/noun acc sentence F1 Sanskrit-unit F1
byt5-mitra-bo-tagger (this model, 10 epochs) .90 .959 .914 .914 .914 .663
Qwen3.5-9B mitra base, same data 1.00 .962 .942 .909 .891 .687

Limitations

  • Trained on canonical translation literature (Kangyur / Tengyur texts with Sanskrit parallels); other registers are untested.
  • The Sanskrit-unit layer is the noisiest: the training references bracket units inconsistently.
  • The model can drop or alter a syllable in long or unusual spans; always validate that the output reproduces the input and fall back to a retry.
  • Tense classes of verbs are not annotated (only the coarse V / G / N distinction).
Downloads last month
295
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for buddhist-nlp/byt5-mitra-bo-tagger

Finetuned
(1)
this model