Statim Decide Multilingual Base
A decision model for Statim, the native C++ engine for typed decisions: ask any text a
choice, a score or a yes/no question and get calibrated answers from one forward pass, on
CPU or GPU, without Python at runtime. Version 0.4.0, fine-tuned from
convaiinnovations/laya-multilingual (mmBERT-base encoder).
One support ticket, three typed answers, one forward pass: the 60-second film.
Quick start
# Statim release binary: https://github.com/BEKO2210/statim/releases
huggingface-cli download Beko2210/statim-decide-multilingual-base statim-decide-multilingual-base-q8_0.gguf --local-dir models
./statim serve -m multilingual=models/statim-decide-multilingual-base-q8_0.gguf --port 8080
curl -s localhost:8080/v1/systemone -d '{
"state": {"subject": "Duplicate charge on invoice #4411",
"body": "We were billed twice for March. Please refund the second charge."},
"questions": {
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, refunds", "technical": "bugs, outages", "sales": "pricing, contracts"}},
"urgency": {"type": "score", "instructions": "How urgent is this request?",
"criteria": ["not urgent", "soon", "critical"]},
"refund": {"type": "noul", "instructions": "Does the user explicitly request a refund?"}}}'
Open http://127.0.0.1:8080/ for the playground. API reference: docs/API.md.
Files
| File | Size | Use |
|---|---|---|
statim-decide-multilingual-base-f32.gguf |
0.91 GB | reference precision, exact on GPU |
statim-decide-multilingual-base-q8_0.gguf |
0.36 GB | recommended for CPU: 4x smaller |
checkpoint/ |
0.68 GB | Laya-format checkpoint for fine-tuning and the Python reference |
Checksums in SHA256SUMS. The f32 file reproduces the reference implementation within 1e-4 on
Statim's parity tests; q8_0 is smaller and faster on CPU with slightly different logits.
Evaluation
Measured by Statim's no-harm gate (tools/finetune/gate.py)
on held-out test data the model selection never looked at. Against the checkpoint it was trained from, on 54 held-out suites: 18 significant gains, 36 within noise, 0 regressions (two combined binomial standard errors).
| Suite | Role | This model | Base checkpoint | Protocol |
|---|---|---|---|---|
| typed-decisions | trained | 0.7585 | 0.3510 | test split, first 2,000 decisions; its train split is replay data |
| Banking77 | trained | 0.9035 | 0.5175 | test split, first 2,000 rows, all 77 intents in one question |
| MASSIVE intents | trained | 0.7717 | 0.3400 | mean over 12 languages, 150 seeded stratified test rows each |
| AG News | held out | 0.9315 | 0.9380 | zero-shot (never trained on), first 2,000 test rows |
| DAIR Emotion | held out | 0.5265 | 0.5320 | zero-shot, first 2,000 test rows |
| HWU64 intents | held out | 0.8200 | 0.5000 | English, 150 rows; rows overlapping MASSIVE removed |
| SIB-200 topics | held out | 0.7217 | 0.7484 | zero-shot, mean over 4 languages, 150 rows each |
| Sentiment | held out | 0.5933 | 0.5595 | zero-shot, mean over 12 languages, 150 rows each |
| HateCheck | held out | 0.6467 | 0.6303 | zero-shot, mean over 11 languages, 150 rows each |
| Belebele reading | held out | 0.3100 | 0.2467 | zero-shot, mean over 4 languages, 150 rows each |
Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768, laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881; Banking77 supervised MPNet 0.941; MASSIVE XLM-R base 0.857 (full train set). Sources: docs/ROADMAP.md.
All 54 held-out suites
| Suite | Accuracy | Rows |
|---|---|---|
amazon_massive_intent/ar |
0.6600 | 150 |
amazon_massive_intent/de |
0.7667 | 150 |
amazon_massive_intent/en |
0.8133 | 150 |
amazon_massive_intent/es |
0.7733 | 150 |
amazon_massive_intent/fr |
0.7600 | 150 |
amazon_massive_intent/hi |
0.7600 | 150 |
amazon_massive_intent/it |
0.7667 | 150 |
amazon_massive_intent/ja |
0.7800 | 150 |
amazon_massive_intent/pl |
0.8000 | 150 |
amazon_massive_intent/ru |
0.8200 | 150 |
amazon_massive_intent/tr |
0.7600 | 150 |
amazon_massive_intent/zh-CN |
0.8000 | 150 |
belebele/ar |
0.2467 | 150 |
belebele/de |
0.3400 | 150 |
belebele/en |
0.3467 | 150 |
belebele/hi |
0.3067 | 150 |
farstail/fa |
0.6333 | 150 |
go_emotions/en |
0.4667 | 150 |
hwu64/en |
0.8200 | 150 |
indonli/id |
0.7000 | 150 |
multi_hatecheck/ar |
0.5800 | 150 |
multi_hatecheck/de |
0.6667 | 150 |
multi_hatecheck/en |
0.6333 | 150 |
multi_hatecheck/es |
0.6733 | 150 |
multi_hatecheck/fr |
0.6667 | 150 |
multi_hatecheck/hi |
0.5600 | 150 |
multi_hatecheck/it |
0.6867 | 150 |
multi_hatecheck/nl |
0.6000 | 150 |
multi_hatecheck/pl |
0.6400 | 150 |
multi_hatecheck/pt |
0.7067 | 150 |
multi_hatecheck/zh |
0.7000 | 150 |
multilingual_sentiments/ar |
0.5533 | 150 |
multilingual_sentiments/de |
0.5600 | 150 |
multilingual_sentiments/en |
0.6400 | 150 |
multilingual_sentiments/es |
0.5467 | 150 |
multilingual_sentiments/fr |
0.5533 | 150 |
multilingual_sentiments/hi |
0.4600 | 150 |
multilingual_sentiments/id |
0.7133 | 150 |
multilingual_sentiments/it |
0.6000 | 150 |
multilingual_sentiments/ja |
0.6733 | 150 |
multilingual_sentiments/ms |
0.5267 | 150 |
multilingual_sentiments/pt |
0.6600 | 150 |
multilingual_sentiments/zh |
0.6333 | 150 |
semrel/ar |
0.2400 | 150 |
semrel/en |
0.2200 | 150 |
semrel/hi |
0.2200 | 150 |
sib200/ar |
0.7067 | 150 |
sib200/de |
0.7267 | 150 |
sib200/en |
0.7467 | 150 |
sib200/hi |
0.7067 | 150 |
test/ag_news |
0.9315 | 2000 |
test/banking77 |
0.9035 | 2000 |
test/emotion |
0.5265 | 2000 |
test/typed_decisions |
0.7585 | 2000 |
Reproduce these numbers: REPRODUCE.md.
Training
Multi-task fine-tuning with train_multitask.py --clean from checkpoint laya-multilingual, best
epoch 19/raw selected on validation data only. Training data: only sources whose
licence permits commercial use and imposes no ShareAlike or copyleft terms (Banking77, MASSIVE,
typed-decisions replay, a licence-audited tasksource mixture, Nemotron-Safety-Guard, IndicGuard,
MINDS-14, SNIPS), every one listed with its licence in
DATA_LICENSES.md. Evaluation test rows were removed from the training data.
Intended use and limits
- Classification-style decisions over short texts and JSON: routing, triage, moderation, intent, yes/no checks, ordinal ratings. It does not generate text.
- Accuracy varies by task and language (see the table). Reading comprehension (Belebele) and semantic similarity are weak; do not use it for them without your own evaluation.
- Use the confidence: Statim's
min_confidenceoption marks low-confidence answers withescalate: trueso a person can review them. Do not automate high-stakes decisions about people without human review.
Licence
The weights may be used under any one of: PolyForm Noncommercial 1.0.0, PolyForm Small Business 1.0.0 (free commercial use below 100 people and 1 M USD revenue), PolyForm Free Trial 1.0.0 (any company, fewer than 32 days), or a Statim commercial licence (COMMERCIAL.md). Texts in LICENSE-MODEL.md. The Statim engine is Apache-2.0.
Built on Laya (Apache-2.0) and mmBERT-base (MIT). Training data attribution: Banking77 (Casanueva et al., 2020, PolyAI), MASSIVE (FitzGerald et al., 2022, Amazon), and the CC-BY sources in DATA_LICENSES.md. Statim is independent and not affiliated with the Laya authors.
- Downloads last month
- 48
8-bit
32-bit
Model tree for Beko2210/statim-decide-multilingual-base
Base model
convaiinnovations/laya-multilingualDatasets used to train Beko2210/statim-decide-multilingual-base
AmazonScience/massive
PolyAI/minds14
Space using Beko2210/statim-decide-multilingual-base 1
Evaluation results
- accuracy on typed-decisions (test split, first 2,000 decisions; its train split is replay data)self-reported0.758
- accuracy on Banking77 (test split, first 2,000 rows, all 77 intents in one question)self-reported0.903
- accuracy on MASSIVE intents (mean over 12 languages, 150 seeded stratified test rows each)self-reported0.772
- accuracy on AG News (zero-shot (never trained on), first 2,000 test rows)self-reported0.931
- accuracy on DAIR Emotion (zero-shot, first 2,000 test rows)self-reported0.526
- accuracy on HWU64 intents (English, 150 rows; rows overlapping MASSIVE removed)self-reported0.820
- accuracy on SIB-200 topics (zero-shot, mean over 4 languages, 150 rows each)self-reported0.722
- accuracy on Sentiment (zero-shot, mean over 12 languages, 150 rows each)self-reported0.593