---
license: other
license_name: statim-weights
license_link: https://huggingface.co/Beko2210/statim-decide-multilingual-base/blob/main/LICENSE-MODEL.md
language:
- ar
- de
- en
- es
- fr
- hi
- it
- ja
- pl
- ru
- tr
- zh
- id
- ms
- pt
- nl
- fa
library_name: gguf
pipeline_tag: zero-shot-classification
base_model: convaiinnovations/laya-multilingual
datasets:
- PolyAI/banking77
- AmazonScience/massive
- LocalLLaMA/typed-decisions
- tasksource/tasksource-jev-typed-decisions
- nvidia/Nemotron-Safety-Guard-Dataset-v3
- l3cube-pune/IndicGuard
- PolyAI/minds14
- benayas/snips
tags:
- statim
- gguf
- decision-making
- text-classification
- zero-shot-classification
- mmBERT-base
model-index:
- name: statim-decide-multilingual-base
results:
- task:
type: text-classification
dataset:
name: typed-decisions (test split, first 2,000 decisions; its train split is
replay data)
type: LocalLLaMA/typed-decisions
metrics:
- type: accuracy
value: 0.763
- task:
type: text-classification
dataset:
name: Banking77 (test split, first 2,000 rows, all 77 intents in one question)
type: PolyAI/banking77
metrics:
- type: accuracy
value: 0.914
- task:
type: text-classification
dataset:
name: MASSIVE intents (mean over 12 languages, 150 seeded stratified test rows
each)
type: AmazonScience/massive
metrics:
- type: accuracy
value: 0.7995
- task:
type: text-classification
dataset:
name: AG News (zero-shot (never trained on), first 2,000 test rows)
type: fancyzhx/ag_news
metrics:
- type: accuracy
value: 0.9295
- task:
type: text-classification
dataset:
name: DAIR Emotion (zero-shot, first 2,000 test rows)
type: dair-ai/emotion
metrics:
- type: accuracy
value: 0.504
- task:
type: text-classification
dataset:
name: HWU64 intents (English, 150 rows; rows overlapping MASSIVE removed)
type: hwu64
metrics:
- type: accuracy
value: 0.7867
- task:
type: text-classification
dataset:
name: SIB-200 topics (zero-shot, mean over 4 languages, 150 rows each)
type: Davlan/sib200
metrics:
- type: accuracy
value: 0.7817
- task:
type: text-classification
dataset:
name: Sentiment (zero-shot, mean over 12 languages, 150 rows each)
type: tyqiangz/multilingual-sentiments
metrics:
- type: accuracy
value: 0.605
- task:
type: text-classification
dataset:
name: HateCheck (zero-shot, mean over 11 languages, 150 rows each)
type: mteb/multi-hatecheck
metrics:
- type: accuracy
value: 0.6358
- task:
type: text-classification
dataset:
name: Belebele reading (zero-shot, mean over 4 languages, 150 rows each)
type: facebook/belebele
metrics:
- type: accuracy
value: 0.2583
---
# Statim Decide Multilingual Base
A decision model for [Statim](https://github.com/BEKO2210/statim), the native C++ engine for typed decisions: ask any text a
**choice**, a **score** or a **yes/no** question and get calibrated answers from one forward pass, on
CPU or GPU, without Python at runtime. Version **0.7.0**, fine-tuned from
[`convaiinnovations/laya-multilingual`](https://huggingface.co/convaiinnovations/laya-multilingual) (mmBERT-base encoder).
One support ticket, three typed answers, one forward pass: [the 60-second film](https://beko2210.github.io/statim/#film).
## Quick start
```sh
# Statim release binary: https://github.com/BEKO2210/statim/releases
huggingface-cli download Beko2210/statim-decide-multilingual-base statim-decide-multilingual-base-q8_0.gguf --local-dir models
./statim serve -m multilingual=models/statim-decide-multilingual-base-q8_0.gguf --port 8080
curl -s localhost:8080/v1/systemone -d '{
"state": {"subject": "Duplicate charge on invoice #4411",
"body": "We were billed twice for March. Please refund the second charge."},
"questions": {
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, refunds", "technical": "bugs, outages", "sales": "pricing, contracts"}},
"urgency": {"type": "score", "instructions": "How urgent is this request?",
"criteria": ["not urgent", "soon", "critical"]},
"refund": {"type": "noul", "instructions": "Does the user explicitly request a refund?"}}}'
```
Open `http://127.0.0.1:8080/` for the playground. API reference: [docs/API.md](https://github.com/BEKO2210/statim/blob/main/docs/API.md).
## Files
| File | Size | Use |
|---|---|---|
| `statim-decide-multilingual-base-f32.gguf` | 0.91 GB | reference precision, exact on GPU |
| `statim-decide-multilingual-base-q8_0.gguf` | 0.36 GB | recommended for CPU: 4x smaller |
| `checkpoint/` | 0.68 GB | Laya-format checkpoint for fine-tuning and the Python reference |
Checksums in `SHA256SUMS`. The f32 file reproduces the reference implementation within 1e-4 on
Statim's parity tests; q8_0 is smaller and faster on CPU with slightly different logits.
## Evaluation
Measured by Statim's no-harm gate ([`tools/finetune/gate.py`](https://github.com/BEKO2210/statim/blob/main/tools/finetune/gate.py))
on held-out test data the model selection never looked at. Against the checkpoint it was trained from, on 89 held-out suites: **23 significant gains, 65 within noise, 0 regressions** (gains: more than two combined binomial standard errors; regressions: significant after Holm-Bonferroni over all suites; 1 nominal drop beyond two standard errors did not stay significant).
| Suite | Role | This model | Base checkpoint | Protocol |
|---|---|---|---|---|
| typed-decisions | trained | **0.7630** | 0.7585 | test split, first 2,000 decisions; its train split is replay data |
| Banking77 | trained | **0.9140** | 0.9035 | test split, first 2,000 rows, all 77 intents in one question |
| MASSIVE intents | trained | **0.7995** | 0.7717 | mean over 12 languages, 150 seeded stratified test rows each |
| AG News | held out | **0.9295** | 0.9315 | zero-shot (never trained on), first 2,000 test rows |
| DAIR Emotion | held out | **0.5040** | 0.5265 | zero-shot, first 2,000 test rows |
| HWU64 intents | held out | **0.7867** | 0.8200 | English, 150 rows; rows overlapping MASSIVE removed |
| SIB-200 topics | held out | **0.7817** | 0.7217 | zero-shot, mean over 4 languages, 150 rows each |
| Sentiment | held out | **0.6050** | 0.5933 | zero-shot, mean over 12 languages, 150 rows each |
| HateCheck | held out | **0.6358** | 0.6467 | zero-shot, mean over 11 languages, 150 rows each |
| Belebele reading | held out | **0.2583** | 0.3100 | zero-shot, mean over 4 languages, 150 rows each |
### Decision categories
One held-out suite per decision category, built from splits of the training sources that the mixture never loads; any text that also occurs in the training mixture is dropped. 150 items per language, macro over languages.
| Category | Languages | This model | Base checkpoint |
|---|---|---|---|
| complaint | en | **0.767** | 0.607 |
| emotion | de, en, es, fr, hi, zh | **0.586** | 0.520 |
| fact check | en | **0.313** | 0.260 |
| formality | ja, tr | **0.773** | 0.503 |
| intent | en, nl, tr | **0.753** | 0.516 |
| nli | en, ja, tr | **0.747** | 0.704 |
| pii | ar, de, en, es, fr, it, ja, nl, ru, sv, zh | **0.856** | 0.595 |
| reading | en | **0.927** | 0.560 |
| safety | en | **0.727** | 0.600 |
| sentiment | en, zh | **0.800** | 0.740 |
| similarity | pt | **0.833** | 0.627 |
| stance | en | **0.893** | 0.660 |
| topic | en | **0.607** | 0.267 |
| urgency | en | **0.893** | 0.660 |
Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768,
laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881;
Banking77 supervised MPNet 0.941; MASSIVE XLM-R base 0.857 over 12 languages (full train set). Sources:
[docs/ROADMAP.md](https://github.com/BEKO2210/statim/blob/main/docs/ROADMAP.md).
All 89 held-out suites
| Suite | Accuracy | Rows |
|---|---|---|
| `amazon_massive_intent/ar` | 0.7067 | 150 |
| `amazon_massive_intent/de` | 0.7867 | 150 |
| `amazon_massive_intent/en` | 0.8267 | 150 |
| `amazon_massive_intent/es` | 0.8133 | 150 |
| `amazon_massive_intent/fr` | 0.8200 | 150 |
| `amazon_massive_intent/hi` | 0.7867 | 150 |
| `amazon_massive_intent/it` | 0.8200 | 150 |
| `amazon_massive_intent/ja` | 0.8267 | 150 |
| `amazon_massive_intent/pl` | 0.8000 | 150 |
| `amazon_massive_intent/ru` | 0.8333 | 150 |
| `amazon_massive_intent/tr` | 0.7867 | 150 |
| `amazon_massive_intent/zh-CN` | 0.7867 | 150 |
| `belebele/ar` | 0.2933 | 150 |
| `belebele/de` | 0.2133 | 150 |
| `belebele/en` | 0.2467 | 150 |
| `belebele/hi` | 0.2800 | 150 |
| `categories:complaint/en` | 0.7667 | 150 |
| `categories:emotion/de` | 0.4533 | 150 |
| `categories:emotion/en` | 0.5733 | 150 |
| `categories:emotion/es` | 0.5733 | 150 |
| `categories:emotion/fr` | 0.6600 | 150 |
| `categories:emotion/hi` | 0.7133 | 150 |
| `categories:emotion/zh` | 0.5400 | 150 |
| `categories:fact_check/en` | 0.3133 | 150 |
| `categories:formality/ja` | 0.5467 | 150 |
| `categories:formality/tr` | 1.0000 | 150 |
| `categories:intent/en` | 0.8200 | 150 |
| `categories:intent/nl` | 0.4400 | 150 |
| `categories:intent/tr` | 1.0000 | 150 |
| `categories:nli/en` | 0.8067 | 150 |
| `categories:nli/ja` | 0.6333 | 150 |
| `categories:nli/tr` | 0.8000 | 150 |
| `categories:pii/ar` | 0.8467 | 150 |
| `categories:pii/de` | 0.8400 | 150 |
| `categories:pii/en` | 0.8933 | 150 |
| `categories:pii/es` | 0.8533 | 150 |
| `categories:pii/fr` | 0.8933 | 150 |
| `categories:pii/it` | 0.8400 | 150 |
| `categories:pii/ja` | 0.8467 | 150 |
| `categories:pii/nl` | 0.7600 | 150 |
| `categories:pii/ru` | 0.9067 | 150 |
| `categories:pii/sv` | 0.8333 | 150 |
| `categories:pii/zh` | 0.9000 | 150 |
| `categories:reading/en` | 0.9267 | 150 |
| `categories:safety/en` | 0.7267 | 150 |
| `categories:sentiment/en` | 0.8000 | 150 |
| `categories:sentiment/zh` | 0.8000 | 150 |
| `categories:similarity/pt` | 0.8333 | 150 |
| `categories:stance/en` | 0.8933 | 150 |
| `categories:topic/en` | 0.6067 | 150 |
| `categories:urgency/en` | 0.8933 | 150 |
| `farstail/fa` | 0.7000 | 150 |
| `go_emotions/en` | 0.4733 | 150 |
| `hwu64/en` | 0.7867 | 150 |
| `indonli/id` | 0.6667 | 150 |
| `multi_hatecheck/ar` | 0.5733 | 150 |
| `multi_hatecheck/de` | 0.6867 | 150 |
| `multi_hatecheck/en` | 0.6133 | 150 |
| `multi_hatecheck/es` | 0.6200 | 150 |
| `multi_hatecheck/fr` | 0.6867 | 150 |
| `multi_hatecheck/hi` | 0.5733 | 150 |
| `multi_hatecheck/it` | 0.6600 | 150 |
| `multi_hatecheck/nl` | 0.6533 | 150 |
| `multi_hatecheck/pl` | 0.6333 | 150 |
| `multi_hatecheck/pt` | 0.6467 | 150 |
| `multi_hatecheck/zh` | 0.6467 | 150 |
| `multilingual_sentiments/ar` | 0.5600 | 150 |
| `multilingual_sentiments/de` | 0.5800 | 150 |
| `multilingual_sentiments/en` | 0.6933 | 150 |
| `multilingual_sentiments/es` | 0.5267 | 150 |
| `multilingual_sentiments/fr` | 0.5867 | 150 |
| `multilingual_sentiments/hi` | 0.5333 | 150 |
| `multilingual_sentiments/id` | 0.7467 | 150 |
| `multilingual_sentiments/it` | 0.5667 | 150 |
| `multilingual_sentiments/ja` | 0.6600 | 150 |
| `multilingual_sentiments/ms` | 0.5333 | 150 |
| `multilingual_sentiments/pt` | 0.6333 | 150 |
| `multilingual_sentiments/zh` | 0.6400 | 150 |
| `semrel/ar` | 0.2467 | 150 |
| `semrel/en` | 0.2000 | 150 |
| `semrel/hi` | 0.2267 | 150 |
| `sib200/ar` | 0.7867 | 150 |
| `sib200/de` | 0.8000 | 150 |
| `sib200/en` | 0.8333 | 150 |
| `sib200/hi` | 0.7067 | 150 |
| `test/ag_news` | 0.9295 | 2000 |
| `test/banking77` | 0.9140 | 2000 |
| `test/emotion` | 0.5040 | 2000 |
| `test/typed_decisions` | 0.7630 | 2000 |
Reproduce these numbers: [REPRODUCE.md](https://github.com/BEKO2210/statim/blob/main/REPRODUCE.md).
## Training
Multi-task fine-tuning with `train_multitask.py --clean` from checkpoint `laya-multilingual-big1`, best
epoch `12/raw` selected on validation data only. Training data: only sources whose
licence permits commercial use and imposes no ShareAlike or copyleft terms (Banking77, MASSIVE,
typed-decisions replay, a licence-audited tasksource mixture, Nemotron-Safety-Guard, IndicGuard,
MINDS-14, SNIPS), every one listed with its licence in
[DATA_LICENSES.md](DATA_LICENSES.md). Evaluation test rows were removed from the training data.
## Intended use and limits
- Classification-style decisions over short texts and JSON: routing, triage, moderation, intent,
yes/no checks, ordinal ratings. It does not generate text.
- Accuracy varies by task and language (see the table). Reading comprehension (Belebele) and
semantic similarity are weak; do not use it for them without your own evaluation.
- Use the confidence: Statim's `min_confidence` option marks low-confidence answers with
`escalate: true` so a person can review them. Do not automate high-stakes decisions about people
without human review.
## Licence
The weights may be used under any one of: PolyForm Noncommercial 1.0.0, PolyForm Small Business
1.0.0 (free commercial use below 100 people and 1 M USD revenue), PolyForm Free Trial 1.0.0 (any
company, fewer than 32 days), or a Statim commercial licence
([COMMERCIAL.md](https://github.com/BEKO2210/statim/blob/main/COMMERCIAL.md)). Texts in [LICENSE-MODEL.md](LICENSE-MODEL.md).
The Statim engine is Apache-2.0.
Built on [Laya](https://huggingface.co/convaiinnovations/laya-multilingual) (Apache-2.0) and
[mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) (MIT).
Training data attribution: Banking77 (Casanueva et al., 2020, PolyAI), MASSIVE (FitzGerald et al.,
2022, Amazon), and the CC-BY sources in DATA_LICENSES.md. Statim is independent and not affiliated
with the Laya authors.