Brazilian Portuguese: a native benchmark and a checkpoint fine-tuned with your script

#2
by ajgcvm - opened

Hi, and thanks for releasing Laya with the fine-tuning loop. I run a small company in Brazil (Felhen) and we use decision models inside our products, so the first thing I wanted to know was how Laya does on Brazilian Portuguese. We measured it, fine-tuned with your script, and put both artifacts out in the open.

The benchmark. Seven typed-decision tasks over native PT-BR text, no machine translation, from permissively licensed sources with human labels: legislative bills, court-of-accounts case law, scientific abstracts, social media comments, fact-checked claims and a FAQ answer-matching task. Five choice tasks (3 to 24 options) and two noul tasks, with train and test splits.
https://huggingface.co/datasets/felhen-ai/ptbr-typed-decisions-bench

The checkpoint. Fine-tuned from laya-multilingual with laya_finetune_mps.py, run on CUDA with a one-line device override and otherwise unchanged. It trains on the benchmark train split plus Portuguese decisions from our own operation and generated questions over real and synthetic texts.
https://huggingface.co/felhen-ai/saracura-ptbr-v0

Results (balanced accuracy, option order shuffled per item, up to 1,000 test items per task)

Model Mean over 7 tasks
Majority class 27.6%
laya-multilingual 39.2%
telepatia-ai/laya-pt-es-typed 44.6%
TF-IDF + logistic regression, one classifier per task 64.0%
felhen-ai/saracura-ptbr-v0, fine-tuned, single model for all tasks 68.5%

Zero-shot, the gap in Portuguese is large. Fine-tuning with your script closes most of it, and one model ends up covering all seven tasks. Details and limitations are in the model card.

Two things that may be useful to you: a --device cuda option in the MPS script would remove the need for a wrapper, and if you want other checkpoints run on this benchmark, I am happy to do it. Thanks again for publishing the loop; starting from it is what made this a few days of work instead of a few weeks.

Sign up or log in to comment