Results on a Brazilian Portuguese typed-decision benchmark

#1
by ajgcvm - opened

Hi! We ran Jet v6.2 on a PT-BR typed-decisions benchmark: 7 tasks over native Brazilian Portuguese text, plus a held-out Central Bank FAQ task. We ran your Jet class unchanged, behind a minimal /v1/systemone wrapper. It scored 63.2% mean balanced accuracy, tied for second among the 13 untuned open decision models up to 4B we tested, and 97.6% on the FAQ task. A 4B fine-tuned on Portuguese reached 72.9% on the benchmark.

Full table and per-task results: https://huggingface.co/datasets/felhen-ai/ptbr-typed-decisions-bench
Evaluator: https://github.com/felhen-ai/saracura/blob/main/research/eval_http.py

If anything looks off for your model, like settings or server flags, tell us and we will rerun it. Thanks for releasing the weights.

michaljach changed discussion status to closed

Sign up or log in to comment