Statim Decide Multilingual Base

A decision model for Statim, the native C++ engine for typed decisions: ask any text a choice, a score or a yes/no question and get calibrated answers from one forward pass, on CPU or GPU, without Python at runtime. Version 0.4.0, fine-tuned from convaiinnovations/laya-multilingual (mmBERT-base encoder).

One support ticket, three typed answers, one forward pass: the 60-second film.

Quick start

# Statim release binary: https://github.com/BEKO2210/statim/releases
huggingface-cli download Beko2210/statim-decide-multilingual-base statim-decide-multilingual-base-q8_0.gguf --local-dir models
./statim serve -m multilingual=models/statim-decide-multilingual-base-q8_0.gguf --port 8080
curl -s localhost:8080/v1/systemone -d '{
  "state": {"subject": "Duplicate charge on invoice #4411",
            "body": "We were billed twice for March. Please refund the second charge."},
  "questions": {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
      "criteria": {"billing": "invoices, refunds", "technical": "bugs, outages", "sales": "pricing, contracts"}},
    "urgency": {"type": "score", "instructions": "How urgent is this request?",
      "criteria": ["not urgent", "soon", "critical"]},
    "refund": {"type": "noul", "instructions": "Does the user explicitly request a refund?"}}}'

Open http://127.0.0.1:8080/ for the playground. API reference: docs/API.md.

Files

File Size Use
statim-decide-multilingual-base-f32.gguf 0.91 GB reference precision, exact on GPU
statim-decide-multilingual-base-q8_0.gguf 0.36 GB recommended for CPU: 4x smaller
checkpoint/ 0.68 GB Laya-format checkpoint for fine-tuning and the Python reference

Checksums in SHA256SUMS. The f32 file reproduces the reference implementation within 1e-4 on Statim's parity tests; q8_0 is smaller and faster on CPU with slightly different logits.

Evaluation

Measured by Statim's no-harm gate (tools/finetune/gate.py) on held-out test data the model selection never looked at. Against the checkpoint it was trained from, on 54 held-out suites: 18 significant gains, 36 within noise, 0 regressions (two combined binomial standard errors).

Suite Role This model Base checkpoint Protocol
typed-decisions trained 0.7585 0.3510 test split, first 2,000 decisions; its train split is replay data
Banking77 trained 0.9035 0.5175 test split, first 2,000 rows, all 77 intents in one question
MASSIVE intents trained 0.7717 0.3400 mean over 12 languages, 150 seeded stratified test rows each
AG News held out 0.9315 0.9380 zero-shot (never trained on), first 2,000 test rows
DAIR Emotion held out 0.5265 0.5320 zero-shot, first 2,000 test rows
HWU64 intents held out 0.8200 0.5000 English, 150 rows; rows overlapping MASSIVE removed
SIB-200 topics held out 0.7217 0.7484 zero-shot, mean over 4 languages, 150 rows each
Sentiment held out 0.5933 0.5595 zero-shot, mean over 12 languages, 150 rows each
HateCheck held out 0.6467 0.6303 zero-shot, mean over 11 languages, 150 rows each
Belebele reading held out 0.3100 0.2467 zero-shot, mean over 4 languages, 150 rows each

Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768, laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881; Banking77 supervised MPNet 0.941; MASSIVE XLM-R base 0.857 (full train set). Sources: docs/ROADMAP.md.

All 54 held-out suites
Suite Accuracy Rows
amazon_massive_intent/ar 0.6600 150
amazon_massive_intent/de 0.7667 150
amazon_massive_intent/en 0.8133 150
amazon_massive_intent/es 0.7733 150
amazon_massive_intent/fr 0.7600 150
amazon_massive_intent/hi 0.7600 150
amazon_massive_intent/it 0.7667 150
amazon_massive_intent/ja 0.7800 150
amazon_massive_intent/pl 0.8000 150
amazon_massive_intent/ru 0.8200 150
amazon_massive_intent/tr 0.7600 150
amazon_massive_intent/zh-CN 0.8000 150
belebele/ar 0.2467 150
belebele/de 0.3400 150
belebele/en 0.3467 150
belebele/hi 0.3067 150
farstail/fa 0.6333 150
go_emotions/en 0.4667 150
hwu64/en 0.8200 150
indonli/id 0.7000 150
multi_hatecheck/ar 0.5800 150
multi_hatecheck/de 0.6667 150
multi_hatecheck/en 0.6333 150
multi_hatecheck/es 0.6733 150
multi_hatecheck/fr 0.6667 150
multi_hatecheck/hi 0.5600 150
multi_hatecheck/it 0.6867 150
multi_hatecheck/nl 0.6000 150
multi_hatecheck/pl 0.6400 150
multi_hatecheck/pt 0.7067 150
multi_hatecheck/zh 0.7000 150
multilingual_sentiments/ar 0.5533 150
multilingual_sentiments/de 0.5600 150
multilingual_sentiments/en 0.6400 150
multilingual_sentiments/es 0.5467 150
multilingual_sentiments/fr 0.5533 150
multilingual_sentiments/hi 0.4600 150
multilingual_sentiments/id 0.7133 150
multilingual_sentiments/it 0.6000 150
multilingual_sentiments/ja 0.6733 150
multilingual_sentiments/ms 0.5267 150
multilingual_sentiments/pt 0.6600 150
multilingual_sentiments/zh 0.6333 150
semrel/ar 0.2400 150
semrel/en 0.2200 150
semrel/hi 0.2200 150
sib200/ar 0.7067 150
sib200/de 0.7267 150
sib200/en 0.7467 150
sib200/hi 0.7067 150
test/ag_news 0.9315 2000
test/banking77 0.9035 2000
test/emotion 0.5265 2000
test/typed_decisions 0.7585 2000

Reproduce these numbers: REPRODUCE.md.

Training

Multi-task fine-tuning with train_multitask.py --clean from checkpoint laya-multilingual, best epoch 19/raw selected on validation data only. Training data: only sources whose licence permits commercial use and imposes no ShareAlike or copyleft terms (Banking77, MASSIVE, typed-decisions replay, a licence-audited tasksource mixture, Nemotron-Safety-Guard, IndicGuard, MINDS-14, SNIPS), every one listed with its licence in DATA_LICENSES.md. Evaluation test rows were removed from the training data.

Intended use and limits

  • Classification-style decisions over short texts and JSON: routing, triage, moderation, intent, yes/no checks, ordinal ratings. It does not generate text.
  • Accuracy varies by task and language (see the table). Reading comprehension (Belebele) and semantic similarity are weak; do not use it for them without your own evaluation.
  • Use the confidence: Statim's min_confidence option marks low-confidence answers with escalate: true so a person can review them. Do not automate high-stakes decisions about people without human review.

Licence

The weights may be used under any one of: PolyForm Noncommercial 1.0.0, PolyForm Small Business 1.0.0 (free commercial use below 100 people and 1 M USD revenue), PolyForm Free Trial 1.0.0 (any company, fewer than 32 days), or a Statim commercial licence (COMMERCIAL.md). Texts in LICENSE-MODEL.md. The Statim engine is Apache-2.0.

Built on Laya (Apache-2.0) and mmBERT-base (MIT). Training data attribution: Banking77 (Casanueva et al., 2020, PolyAI), MASSIVE (FitzGerald et al., 2022, Amazon), and the CC-BY sources in DATA_LICENSES.md. Statim is independent and not affiliated with the Laya authors.

Downloads last month
48
GGUF
Model size
0.3B params
Architecture
laya
Hardware compatibility
Log In to add your hardware

8-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Beko2210/statim-decide-multilingual-base

Quantized
(25)
this model

Datasets used to train Beko2210/statim-decide-multilingual-base

Space using Beko2210/statim-decide-multilingual-base 1

Evaluation results

  • accuracy on typed-decisions (test split, first 2,000 decisions; its train split is replay data)
    self-reported
    0.758
  • accuracy on Banking77 (test split, first 2,000 rows, all 77 intents in one question)
    self-reported
    0.903
  • accuracy on MASSIVE intents (mean over 12 languages, 150 seeded stratified test rows each)
    self-reported
    0.772
  • accuracy on AG News (zero-shot (never trained on), first 2,000 test rows)
    self-reported
    0.931
  • accuracy on DAIR Emotion (zero-shot, first 2,000 test rows)
    self-reported
    0.526
  • accuracy on HWU64 intents (English, 150 rows; rows overlapping MASSIVE removed)
    self-reported
    0.820
  • accuracy on SIB-200 topics (zero-shot, mean over 4 languages, 150 rows each)
    self-reported
    0.722
  • accuracy on Sentiment (zero-shot, mean over 12 languages, 150 rows each)
    self-reported
    0.593