mrbesher's picture
Remove misplaced image badges
7e14ebe verified
|
Raw
History Blame Contribute Delete
1.93 kB
metadata
language:
  - tr
license: apache-2.0
library_name: transformers
base_model: ytu-ce-cosmos/modernbert-tr-base
base_model_relation: finetune
pipeline_tag: token-classification
datasets:
  - mrbesher/massive-tr
metrics:
  - f1
tags:
  - modernbert
  - massive
  - slot-filling
  - slu
  - onnx
  - encoderfile

ModernBERT-TR MASSIVE slots

ModernBERT-TR MASSIVE Slot Filling

A 150M-parameter Turkish token classifier with O plus BIO labels for the 55 slot types in MASSIVE 1.1.

Results

Our model scores 75.27 +/- 0.31% seqeval entity-level F1 on the MASSIVE 1.1 tr-TR test set. The released checkpoint scores 75.30%.

The quantized int8 version scores 75.42%.

Usage

from transformers import pipeline

fill_slots = pipeline(
    "token-classification",
    model="ytu-ce-cosmos/modernbert-tr-massive-slot",
    aggregation_strategy="first",
)
fill_slots("önümüzdeki cuma Ankara'ya bilet bul")

You can call the tokenizer with is_split_into_words=True to allow the tokenizer to split the sentence into words, and keep the first WordPiece label for each word.

Training

We finetune ytu-ce-cosmos/modernbert-tr-base jointly with a 60-way intent head and a 111-way slot head; this repository contains the exported slot head. The human-localized MASSIVE 1.1 Turkish split has 11,514 training, 2,033 validation, and 2,974 test utterances. We use encoder learning rate 5e-5, head learning rate 1e-4, batch size 64, 15 epochs, linear warmup and decay, weight decay 0.01, plain slot cross-entropy, first-subword alignment, and five seeds.

Standalone binaries are available in the encoderfile repo.

License

Apache-2.0. The MASSIVE dataset is distributed under CC BY 4.0.