Results on a Brazilian Portuguese typed-decision benchmark

#1
by ajgcvm - opened

Hi! We ran jpt-4b on a PT-BR typed-decisions benchmark: 7 tasks over native Brazilian Portuguese text, plus a held-out Central Bank FAQ task. We served it with llm2jev (HF backend), unchanged. It scored 63.2% mean balanced accuracy, tied for second among the 13 untuned open decision models up to 4B we tested, and 98.4% on the FAQ task, the best of the 13. A 4B fine-tuned on Portuguese reached 72.9% on the benchmark.

Full table and per-task results: https://huggingface.co/datasets/felhen-ai/ptbr-typed-decisions-bench
Evaluator: https://github.com/felhen-ai/saracura/blob/main/research/eval_http.py

If anything looks off for your model, like settings or server flags, tell us and we will rerun it. Thanks for releasing the weights.

Sign up or log in to comment