Laya (Fine-Tuned on Turkish MMLU Typed-Decisions)

This is Laya fine-tuned on canbingol/mmlu_typed_decision, a 10k-example dataset built by converting the Turkish MMLU dataset into Laya's typed-decision format (state / questions / gold triples, choice-type questions with per-option criteria).

Note: the accuracy/Brier/ECE numbers below are from the original LocalLLaMA/typed-decisions benchmark (1,200 training cases / 400-case test set across Agent Trace Observability, Customer Service, Invoice Processing, and Security Incidents) and reflect that benchmark, not this fine-tune's performance on Turkish MMLU. They're kept here for reference to the base checkpoint's reported numbers.

On the official 400-case test set (2,000 decisions), the base checkpoint achieves 0.365 Accuracy, trailing TypeSafe Jev 1.13.0 (0.727) and the benchmark's Teacher Self-Agreement ceiling (0.735).

Head-to-Head Benchmark Results (base checkpoint, LocalLLaMA/typed-decisions)

Model Kind Accuracy Soft Acc Brier Score ECE Score MAE Within 1 Level Latency (p50) Cost/Case
Turkish Laya fine-tuned 0.365 0.354 0.383 0.242 0.726 0.703 168.3 ms $0.00 (Self-Hosted)
TypeSafe Jev 1.13.0 general 0.727 0.580 0.148 0.144 0.391 0.952 710 ms $0.0004 (API)
ModernBERT-base (149M) specialist 0.646 0.542 0.119 0.179 0.444 0.931 349 ms $0.00
Teacher Self-Agreement ceiling 0.735 - - - - - - -

Training Data

Fine-tuned on canbingol/mmlu_typed_decision — ~10,000 examples derived from the Turkish MMLU dataset, reformatted as choice-type typed decisions (question → instructions, answer options → criteria, correct option → one-hot target).

Installation & Quickstart

pip install laya
import laya

# Load the fine-tuned model directly from Hugging Face
agent = laya.load("convaiinnovations/laya-typed-decisions")

# Evaluate any workflow state and typed questions in a single forward pass
result = agent.predict(state, questions)
print(result["answers"])

License

Apache 2.0. Developed by Convai Innovations.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train canbingol/laya-typed-decisions-turkish-mmlu-10k

Evaluation results