license: other
license_name: statim-weights
license_link: >-
https://huggingface.co/Beko2210/statim-decide-multilingual-base/blob/main/LICENSE-MODEL.md
language:
- ar
- de
- en
- es
- fr
- hi
- it
- ja
- pl
- ru
- tr
- zh
- id
- ms
- pt
- nl
- fa
library_name: gguf
pipeline_tag: zero-shot-classification
base_model: convaiinnovations/laya-multilingual
datasets:
- PolyAI/banking77
- AmazonScience/massive
- LocalLLaMA/typed-decisions
- tasksource/tasksource-jev-typed-decisions
- nvidia/Nemotron-Safety-Guard-Dataset-v3
- l3cube-pune/IndicGuard
- PolyAI/minds14
- benayas/snips
tags:
- statim
- gguf
- decision-making
- text-classification
- zero-shot-classification
- mmBERT-base
model-index:
- name: statim-decide-multilingual-base
results:
- task:
type: text-classification
dataset:
name: >-
typed-decisions (test split, first 2,000 decisions; its train split
is replay data)
type: LocalLLaMA/typed-decisions
metrics:
- type: accuracy
value: 0.763
- task:
type: text-classification
dataset:
name: >-
Banking77 (test split, first 2,000 rows, all 77 intents in one
question)
type: PolyAI/banking77
metrics:
- type: accuracy
value: 0.914
- task:
type: text-classification
dataset:
name: >-
MASSIVE intents (mean over 12 languages, 150 seeded stratified test
rows each)
type: AmazonScience/massive
metrics:
- type: accuracy
value: 0.7995
- task:
type: text-classification
dataset:
name: AG News (zero-shot (never trained on), first 2,000 test rows)
type: fancyzhx/ag_news
metrics:
- type: accuracy
value: 0.9295
- task:
type: text-classification
dataset:
name: DAIR Emotion (zero-shot, first 2,000 test rows)
type: dair-ai/emotion
metrics:
- type: accuracy
value: 0.504
- task:
type: text-classification
dataset:
name: HWU64 intents (English, 150 rows; rows overlapping MASSIVE removed)
type: hwu64
metrics:
- type: accuracy
value: 0.7867
- task:
type: text-classification
dataset:
name: SIB-200 topics (zero-shot, mean over 4 languages, 150 rows each)
type: Davlan/sib200
metrics:
- type: accuracy
value: 0.7817
- task:
type: text-classification
dataset:
name: Sentiment (zero-shot, mean over 12 languages, 150 rows each)
type: tyqiangz/multilingual-sentiments
metrics:
- type: accuracy
value: 0.605
- task:
type: text-classification
dataset:
name: HateCheck (zero-shot, mean over 11 languages, 150 rows each)
type: mteb/multi-hatecheck
metrics:
- type: accuracy
value: 0.6358
- task:
type: text-classification
dataset:
name: Belebele reading (zero-shot, mean over 4 languages, 150 rows each)
type: facebook/belebele
metrics:
- type: accuracy
value: 0.2583
Statim Decide Multilingual Base
A decision model for Statim, the native C++ engine for typed decisions: ask any text a
choice, a score or a yes/no question and get calibrated answers from one forward pass, on
CPU or GPU, without Python at runtime. Version 0.7.0, fine-tuned from
convaiinnovations/laya-multilingual (mmBERT-base encoder).
One support ticket, three typed answers, one forward pass: the 60-second film.
Quick start
# Statim release binary: https://github.com/BEKO2210/statim/releases
huggingface-cli download Beko2210/statim-decide-multilingual-base statim-decide-multilingual-base-q8_0.gguf --local-dir models
./statim serve -m multilingual=models/statim-decide-multilingual-base-q8_0.gguf --port 8080
curl -s localhost:8080/v1/systemone -d '{
"state": {"subject": "Duplicate charge on invoice #4411",
"body": "We were billed twice for March. Please refund the second charge."},
"questions": {
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, refunds", "technical": "bugs, outages", "sales": "pricing, contracts"}},
"urgency": {"type": "score", "instructions": "How urgent is this request?",
"criteria": ["not urgent", "soon", "critical"]},
"refund": {"type": "noul", "instructions": "Does the user explicitly request a refund?"}}}'
Open http://127.0.0.1:8080/ for the playground. API reference: docs/API.md.
Files
| File | Size | Use |
|---|---|---|
statim-decide-multilingual-base-f32.gguf |
0.91 GB | reference precision, exact on GPU |
statim-decide-multilingual-base-q8_0.gguf |
0.36 GB | recommended for CPU: 4x smaller |
checkpoint/ |
0.68 GB | Laya-format checkpoint for fine-tuning and the Python reference |
Checksums in SHA256SUMS. The f32 file reproduces the reference implementation within 1e-4 on
Statim's parity tests; q8_0 is smaller and faster on CPU with slightly different logits.
Evaluation
Measured by Statim's no-harm gate (tools/finetune/gate.py)
on held-out test data the model selection never looked at. Against the checkpoint it was trained from, on 89 held-out suites: 23 significant gains, 65 within noise, 0 regressions (gains: more than two combined binomial standard errors; regressions: significant after Holm-Bonferroni over all suites; 1 nominal drop beyond two standard errors did not stay significant).
| Suite | Role | This model | Base checkpoint | Protocol |
|---|---|---|---|---|
| typed-decisions | trained | 0.7630 | 0.7585 | test split, first 2,000 decisions; its train split is replay data |
| Banking77 | trained | 0.9140 | 0.9035 | test split, first 2,000 rows, all 77 intents in one question |
| MASSIVE intents | trained | 0.7995 | 0.7717 | mean over 12 languages, 150 seeded stratified test rows each |
| AG News | held out | 0.9295 | 0.9315 | zero-shot (never trained on), first 2,000 test rows |
| DAIR Emotion | held out | 0.5040 | 0.5265 | zero-shot, first 2,000 test rows |
| HWU64 intents | held out | 0.7867 | 0.8200 | English, 150 rows; rows overlapping MASSIVE removed |
| SIB-200 topics | held out | 0.7817 | 0.7217 | zero-shot, mean over 4 languages, 150 rows each |
| Sentiment | held out | 0.6050 | 0.5933 | zero-shot, mean over 12 languages, 150 rows each |
| HateCheck | held out | 0.6358 | 0.6467 | zero-shot, mean over 11 languages, 150 rows each |
| Belebele reading | held out | 0.2583 | 0.3100 | zero-shot, mean over 4 languages, 150 rows each |
Decision categories
One held-out suite per decision category, built from splits of the training sources that the mixture never loads; any text that also occurs in the training mixture is dropped. 150 items per language, macro over languages.
| Category | Languages | This model | Base checkpoint |
|---|---|---|---|
| complaint | en | 0.767 | 0.607 |
| emotion | de, en, es, fr, hi, zh | 0.586 | 0.520 |
| fact check | en | 0.313 | 0.260 |
| formality | ja, tr | 0.773 | 0.503 |
| intent | en, nl, tr | 0.753 | 0.516 |
| nli | en, ja, tr | 0.747 | 0.704 |
| pii | ar, de, en, es, fr, it, ja, nl, ru, sv, zh | 0.856 | 0.595 |
| reading | en | 0.927 | 0.560 |
| safety | en | 0.727 | 0.600 |
| sentiment | en, zh | 0.800 | 0.740 |
| similarity | pt | 0.833 | 0.627 |
| stance | en | 0.893 | 0.660 |
| topic | en | 0.607 | 0.267 |
| urgency | en | 0.893 | 0.660 |
Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768, laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881; Banking77 supervised MPNet 0.941; MASSIVE XLM-R base 0.857 over 12 languages (full train set). Sources: docs/ROADMAP.md.
All 89 held-out suites
| Suite | Accuracy | Rows |
|---|---|---|
amazon_massive_intent/ar |
0.7067 | 150 |
amazon_massive_intent/de |
0.7867 | 150 |
amazon_massive_intent/en |
0.8267 | 150 |
amazon_massive_intent/es |
0.8133 | 150 |
amazon_massive_intent/fr |
0.8200 | 150 |
amazon_massive_intent/hi |
0.7867 | 150 |
amazon_massive_intent/it |
0.8200 | 150 |
amazon_massive_intent/ja |
0.8267 | 150 |
amazon_massive_intent/pl |
0.8000 | 150 |
amazon_massive_intent/ru |
0.8333 | 150 |
amazon_massive_intent/tr |
0.7867 | 150 |
amazon_massive_intent/zh-CN |
0.7867 | 150 |
belebele/ar |
0.2933 | 150 |
belebele/de |
0.2133 | 150 |
belebele/en |
0.2467 | 150 |
belebele/hi |
0.2800 | 150 |
categories:complaint/en |
0.7667 | 150 |
categories:emotion/de |
0.4533 | 150 |
categories:emotion/en |
0.5733 | 150 |
categories:emotion/es |
0.5733 | 150 |
categories:emotion/fr |
0.6600 | 150 |
categories:emotion/hi |
0.7133 | 150 |
categories:emotion/zh |
0.5400 | 150 |
categories:fact_check/en |
0.3133 | 150 |
categories:formality/ja |
0.5467 | 150 |
categories:formality/tr |
1.0000 | 150 |
categories:intent/en |
0.8200 | 150 |
categories:intent/nl |
0.4400 | 150 |
categories:intent/tr |
1.0000 | 150 |
categories:nli/en |
0.8067 | 150 |
categories:nli/ja |
0.6333 | 150 |
categories:nli/tr |
0.8000 | 150 |
categories:pii/ar |
0.8467 | 150 |
categories:pii/de |
0.8400 | 150 |
categories:pii/en |
0.8933 | 150 |
categories:pii/es |
0.8533 | 150 |
categories:pii/fr |
0.8933 | 150 |
categories:pii/it |
0.8400 | 150 |
categories:pii/ja |
0.8467 | 150 |
categories:pii/nl |
0.7600 | 150 |
categories:pii/ru |
0.9067 | 150 |
categories:pii/sv |
0.8333 | 150 |
categories:pii/zh |
0.9000 | 150 |
categories:reading/en |
0.9267 | 150 |
categories:safety/en |
0.7267 | 150 |
categories:sentiment/en |
0.8000 | 150 |
categories:sentiment/zh |
0.8000 | 150 |
categories:similarity/pt |
0.8333 | 150 |
categories:stance/en |
0.8933 | 150 |
categories:topic/en |
0.6067 | 150 |
categories:urgency/en |
0.8933 | 150 |
farstail/fa |
0.7000 | 150 |
go_emotions/en |
0.4733 | 150 |
hwu64/en |
0.7867 | 150 |
indonli/id |
0.6667 | 150 |
multi_hatecheck/ar |
0.5733 | 150 |
multi_hatecheck/de |
0.6867 | 150 |
multi_hatecheck/en |
0.6133 | 150 |
multi_hatecheck/es |
0.6200 | 150 |
multi_hatecheck/fr |
0.6867 | 150 |
multi_hatecheck/hi |
0.5733 | 150 |
multi_hatecheck/it |
0.6600 | 150 |
multi_hatecheck/nl |
0.6533 | 150 |
multi_hatecheck/pl |
0.6333 | 150 |
multi_hatecheck/pt |
0.6467 | 150 |
multi_hatecheck/zh |
0.6467 | 150 |
multilingual_sentiments/ar |
0.5600 | 150 |
multilingual_sentiments/de |
0.5800 | 150 |
multilingual_sentiments/en |
0.6933 | 150 |
multilingual_sentiments/es |
0.5267 | 150 |
multilingual_sentiments/fr |
0.5867 | 150 |
multilingual_sentiments/hi |
0.5333 | 150 |
multilingual_sentiments/id |
0.7467 | 150 |
multilingual_sentiments/it |
0.5667 | 150 |
multilingual_sentiments/ja |
0.6600 | 150 |
multilingual_sentiments/ms |
0.5333 | 150 |
multilingual_sentiments/pt |
0.6333 | 150 |
multilingual_sentiments/zh |
0.6400 | 150 |
semrel/ar |
0.2467 | 150 |
semrel/en |
0.2000 | 150 |
semrel/hi |
0.2267 | 150 |
sib200/ar |
0.7867 | 150 |
sib200/de |
0.8000 | 150 |
sib200/en |
0.8333 | 150 |
sib200/hi |
0.7067 | 150 |
test/ag_news |
0.9295 | 2000 |
test/banking77 |
0.9140 | 2000 |
test/emotion |
0.5040 | 2000 |
test/typed_decisions |
0.7630 | 2000 |
Reproduce these numbers: REPRODUCE.md.
Training
Multi-task fine-tuning with train_multitask.py --clean from checkpoint laya-multilingual-big1, best
epoch 12/raw selected on validation data only. Training data: only sources whose
licence permits commercial use and imposes no ShareAlike or copyleft terms (Banking77, MASSIVE,
typed-decisions replay, a licence-audited tasksource mixture, Nemotron-Safety-Guard, IndicGuard,
MINDS-14, SNIPS), every one listed with its licence in
DATA_LICENSES.md. Evaluation test rows were removed from the training data.
Intended use and limits
- Classification-style decisions over short texts and JSON: routing, triage, moderation, intent, yes/no checks, ordinal ratings. It does not generate text.
- Accuracy varies by task and language (see the table). Reading comprehension (Belebele) and semantic similarity are weak; do not use it for them without your own evaluation.
- Use the confidence: Statim's
min_confidenceoption marks low-confidence answers withescalate: trueso a person can review them. Do not automate high-stakes decisions about people without human review.
Licence
The weights may be used under any one of: PolyForm Noncommercial 1.0.0, PolyForm Small Business 1.0.0 (free commercial use below 100 people and 1 M USD revenue), PolyForm Free Trial 1.0.0 (any company, fewer than 32 days), or a Statim commercial licence (COMMERCIAL.md). Texts in LICENSE-MODEL.md. The Statim engine is Apache-2.0.
Built on Laya (Apache-2.0) and mmBERT-base (MIT). Training data attribution: Banking77 (Casanueva et al., 2020, PolyAI), MASSIVE (FitzGerald et al., 2022, Amazon), and the CC-BY sources in DATA_LICENSES.md. Statim is independent and not affiliated with the Laya authors.