--- license: other license_name: statim-weights license_link: https://huggingface.co/Beko2210/statim-decide-multilingual-base/blob/main/LICENSE-MODEL.md language: - ar - de - en - es - fr - hi - it - ja - pl - ru - tr - zh - id - ms - pt - nl - fa library_name: gguf pipeline_tag: zero-shot-classification base_model: convaiinnovations/laya-multilingual datasets: - PolyAI/banking77 - AmazonScience/massive - LocalLLaMA/typed-decisions - tasksource/tasksource-jev-typed-decisions - nvidia/Nemotron-Safety-Guard-Dataset-v3 - l3cube-pune/IndicGuard - PolyAI/minds14 - benayas/snips tags: - statim - gguf - decision-making - text-classification - zero-shot-classification - mmBERT-base model-index: - name: statim-decide-multilingual-base results: - task: type: text-classification dataset: name: typed-decisions (test split, first 2,000 decisions; its train split is replay data) type: LocalLLaMA/typed-decisions metrics: - type: accuracy value: 0.763 - task: type: text-classification dataset: name: Banking77 (test split, first 2,000 rows, all 77 intents in one question) type: PolyAI/banking77 metrics: - type: accuracy value: 0.914 - task: type: text-classification dataset: name: MASSIVE intents (mean over 12 languages, 150 seeded stratified test rows each) type: AmazonScience/massive metrics: - type: accuracy value: 0.7995 - task: type: text-classification dataset: name: AG News (zero-shot (never trained on), first 2,000 test rows) type: fancyzhx/ag_news metrics: - type: accuracy value: 0.9295 - task: type: text-classification dataset: name: DAIR Emotion (zero-shot, first 2,000 test rows) type: dair-ai/emotion metrics: - type: accuracy value: 0.504 - task: type: text-classification dataset: name: HWU64 intents (English, 150 rows; rows overlapping MASSIVE removed) type: hwu64 metrics: - type: accuracy value: 0.7867 - task: type: text-classification dataset: name: SIB-200 topics (zero-shot, mean over 4 languages, 150 rows each) type: Davlan/sib200 metrics: - type: accuracy value: 0.7817 - task: type: text-classification dataset: name: Sentiment (zero-shot, mean over 12 languages, 150 rows each) type: tyqiangz/multilingual-sentiments metrics: - type: accuracy value: 0.605 - task: type: text-classification dataset: name: HateCheck (zero-shot, mean over 11 languages, 150 rows each) type: mteb/multi-hatecheck metrics: - type: accuracy value: 0.6358 - task: type: text-classification dataset: name: Belebele reading (zero-shot, mean over 4 languages, 150 rows each) type: facebook/belebele metrics: - type: accuracy value: 0.2583 --- # Statim Decide Multilingual Base A decision model for [Statim](https://github.com/BEKO2210/statim), the native C++ engine for typed decisions: ask any text a **choice**, a **score** or a **yes/no** question and get calibrated answers from one forward pass, on CPU or GPU, without Python at runtime. Version **0.7.0**, fine-tuned from [`convaiinnovations/laya-multilingual`](https://huggingface.co/convaiinnovations/laya-multilingual) (mmBERT-base encoder). One support ticket, three typed answers, one forward pass: [the 60-second film](https://beko2210.github.io/statim/#film). ## Quick start ```sh # Statim release binary: https://github.com/BEKO2210/statim/releases huggingface-cli download Beko2210/statim-decide-multilingual-base statim-decide-multilingual-base-q8_0.gguf --local-dir models ./statim serve -m multilingual=models/statim-decide-multilingual-base-q8_0.gguf --port 8080 curl -s localhost:8080/v1/systemone -d '{ "state": {"subject": "Duplicate charge on invoice #4411", "body": "We were billed twice for March. Please refund the second charge."}, "questions": { "department": {"type": "choice", "instructions": "Which department should handle this?", "criteria": {"billing": "invoices, refunds", "technical": "bugs, outages", "sales": "pricing, contracts"}}, "urgency": {"type": "score", "instructions": "How urgent is this request?", "criteria": ["not urgent", "soon", "critical"]}, "refund": {"type": "noul", "instructions": "Does the user explicitly request a refund?"}}}' ``` Open `http://127.0.0.1:8080/` for the playground. API reference: [docs/API.md](https://github.com/BEKO2210/statim/blob/main/docs/API.md). ## Files | File | Size | Use | |---|---|---| | `statim-decide-multilingual-base-f32.gguf` | 0.91 GB | reference precision, exact on GPU | | `statim-decide-multilingual-base-q8_0.gguf` | 0.36 GB | recommended for CPU: 4x smaller | | `checkpoint/` | 0.68 GB | Laya-format checkpoint for fine-tuning and the Python reference | Checksums in `SHA256SUMS`. The f32 file reproduces the reference implementation within 1e-4 on Statim's parity tests; q8_0 is smaller and faster on CPU with slightly different logits. ## Evaluation Measured by Statim's no-harm gate ([`tools/finetune/gate.py`](https://github.com/BEKO2210/statim/blob/main/tools/finetune/gate.py)) on held-out test data the model selection never looked at. Against the checkpoint it was trained from, on 89 held-out suites: **23 significant gains, 65 within noise, 0 regressions** (gains: more than two combined binomial standard errors; regressions: significant after Holm-Bonferroni over all suites; 1 nominal drop beyond two standard errors did not stay significant). | Suite | Role | This model | Base checkpoint | Protocol | |---|---|---|---|---| | typed-decisions | trained | **0.7630** | 0.7585 | test split, first 2,000 decisions; its train split is replay data | | Banking77 | trained | **0.9140** | 0.9035 | test split, first 2,000 rows, all 77 intents in one question | | MASSIVE intents | trained | **0.7995** | 0.7717 | mean over 12 languages, 150 seeded stratified test rows each | | AG News | held out | **0.9295** | 0.9315 | zero-shot (never trained on), first 2,000 test rows | | DAIR Emotion | held out | **0.5040** | 0.5265 | zero-shot, first 2,000 test rows | | HWU64 intents | held out | **0.7867** | 0.8200 | English, 150 rows; rows overlapping MASSIVE removed | | SIB-200 topics | held out | **0.7817** | 0.7217 | zero-shot, mean over 4 languages, 150 rows each | | Sentiment | held out | **0.6050** | 0.5933 | zero-shot, mean over 12 languages, 150 rows each | | HateCheck | held out | **0.6358** | 0.6467 | zero-shot, mean over 11 languages, 150 rows each | | Belebele reading | held out | **0.2583** | 0.3100 | zero-shot, mean over 4 languages, 150 rows each | ### Decision categories One held-out suite per decision category, built from splits of the training sources that the mixture never loads; any text that also occurs in the training mixture is dropped. 150 items per language, macro over languages. | Category | Languages | This model | Base checkpoint | |---|---|---|---| | complaint | en | **0.767** | 0.607 | | emotion | de, en, es, fr, hi, zh | **0.586** | 0.520 | | fact check | en | **0.313** | 0.260 | | formality | ja, tr | **0.773** | 0.503 | | intent | en, nl, tr | **0.753** | 0.516 | | nli | en, ja, tr | **0.747** | 0.704 | | pii | ar, de, en, es, fr, it, ja, nl, ru, sv, zh | **0.856** | 0.595 | | reading | en | **0.927** | 0.560 | | safety | en | **0.727** | 0.600 | | sentiment | en, zh | **0.800** | 0.740 | | similarity | pt | **0.833** | 0.627 | | stance | en | **0.893** | 0.660 | | topic | en | **0.607** | 0.267 | | urgency | en | **0.893** | 0.660 | Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768, laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881; Banking77 supervised MPNet 0.941; MASSIVE XLM-R base 0.857 over 12 languages (full train set). Sources: [docs/ROADMAP.md](https://github.com/BEKO2210/statim/blob/main/docs/ROADMAP.md).
All 89 held-out suites | Suite | Accuracy | Rows | |---|---|---| | `amazon_massive_intent/ar` | 0.7067 | 150 | | `amazon_massive_intent/de` | 0.7867 | 150 | | `amazon_massive_intent/en` | 0.8267 | 150 | | `amazon_massive_intent/es` | 0.8133 | 150 | | `amazon_massive_intent/fr` | 0.8200 | 150 | | `amazon_massive_intent/hi` | 0.7867 | 150 | | `amazon_massive_intent/it` | 0.8200 | 150 | | `amazon_massive_intent/ja` | 0.8267 | 150 | | `amazon_massive_intent/pl` | 0.8000 | 150 | | `amazon_massive_intent/ru` | 0.8333 | 150 | | `amazon_massive_intent/tr` | 0.7867 | 150 | | `amazon_massive_intent/zh-CN` | 0.7867 | 150 | | `belebele/ar` | 0.2933 | 150 | | `belebele/de` | 0.2133 | 150 | | `belebele/en` | 0.2467 | 150 | | `belebele/hi` | 0.2800 | 150 | | `categories:complaint/en` | 0.7667 | 150 | | `categories:emotion/de` | 0.4533 | 150 | | `categories:emotion/en` | 0.5733 | 150 | | `categories:emotion/es` | 0.5733 | 150 | | `categories:emotion/fr` | 0.6600 | 150 | | `categories:emotion/hi` | 0.7133 | 150 | | `categories:emotion/zh` | 0.5400 | 150 | | `categories:fact_check/en` | 0.3133 | 150 | | `categories:formality/ja` | 0.5467 | 150 | | `categories:formality/tr` | 1.0000 | 150 | | `categories:intent/en` | 0.8200 | 150 | | `categories:intent/nl` | 0.4400 | 150 | | `categories:intent/tr` | 1.0000 | 150 | | `categories:nli/en` | 0.8067 | 150 | | `categories:nli/ja` | 0.6333 | 150 | | `categories:nli/tr` | 0.8000 | 150 | | `categories:pii/ar` | 0.8467 | 150 | | `categories:pii/de` | 0.8400 | 150 | | `categories:pii/en` | 0.8933 | 150 | | `categories:pii/es` | 0.8533 | 150 | | `categories:pii/fr` | 0.8933 | 150 | | `categories:pii/it` | 0.8400 | 150 | | `categories:pii/ja` | 0.8467 | 150 | | `categories:pii/nl` | 0.7600 | 150 | | `categories:pii/ru` | 0.9067 | 150 | | `categories:pii/sv` | 0.8333 | 150 | | `categories:pii/zh` | 0.9000 | 150 | | `categories:reading/en` | 0.9267 | 150 | | `categories:safety/en` | 0.7267 | 150 | | `categories:sentiment/en` | 0.8000 | 150 | | `categories:sentiment/zh` | 0.8000 | 150 | | `categories:similarity/pt` | 0.8333 | 150 | | `categories:stance/en` | 0.8933 | 150 | | `categories:topic/en` | 0.6067 | 150 | | `categories:urgency/en` | 0.8933 | 150 | | `farstail/fa` | 0.7000 | 150 | | `go_emotions/en` | 0.4733 | 150 | | `hwu64/en` | 0.7867 | 150 | | `indonli/id` | 0.6667 | 150 | | `multi_hatecheck/ar` | 0.5733 | 150 | | `multi_hatecheck/de` | 0.6867 | 150 | | `multi_hatecheck/en` | 0.6133 | 150 | | `multi_hatecheck/es` | 0.6200 | 150 | | `multi_hatecheck/fr` | 0.6867 | 150 | | `multi_hatecheck/hi` | 0.5733 | 150 | | `multi_hatecheck/it` | 0.6600 | 150 | | `multi_hatecheck/nl` | 0.6533 | 150 | | `multi_hatecheck/pl` | 0.6333 | 150 | | `multi_hatecheck/pt` | 0.6467 | 150 | | `multi_hatecheck/zh` | 0.6467 | 150 | | `multilingual_sentiments/ar` | 0.5600 | 150 | | `multilingual_sentiments/de` | 0.5800 | 150 | | `multilingual_sentiments/en` | 0.6933 | 150 | | `multilingual_sentiments/es` | 0.5267 | 150 | | `multilingual_sentiments/fr` | 0.5867 | 150 | | `multilingual_sentiments/hi` | 0.5333 | 150 | | `multilingual_sentiments/id` | 0.7467 | 150 | | `multilingual_sentiments/it` | 0.5667 | 150 | | `multilingual_sentiments/ja` | 0.6600 | 150 | | `multilingual_sentiments/ms` | 0.5333 | 150 | | `multilingual_sentiments/pt` | 0.6333 | 150 | | `multilingual_sentiments/zh` | 0.6400 | 150 | | `semrel/ar` | 0.2467 | 150 | | `semrel/en` | 0.2000 | 150 | | `semrel/hi` | 0.2267 | 150 | | `sib200/ar` | 0.7867 | 150 | | `sib200/de` | 0.8000 | 150 | | `sib200/en` | 0.8333 | 150 | | `sib200/hi` | 0.7067 | 150 | | `test/ag_news` | 0.9295 | 2000 | | `test/banking77` | 0.9140 | 2000 | | `test/emotion` | 0.5040 | 2000 | | `test/typed_decisions` | 0.7630 | 2000 |
Reproduce these numbers: [REPRODUCE.md](https://github.com/BEKO2210/statim/blob/main/REPRODUCE.md). ## Training Multi-task fine-tuning with `train_multitask.py --clean` from checkpoint `laya-multilingual-big1`, best epoch `12/raw` selected on validation data only. Training data: only sources whose licence permits commercial use and imposes no ShareAlike or copyleft terms (Banking77, MASSIVE, typed-decisions replay, a licence-audited tasksource mixture, Nemotron-Safety-Guard, IndicGuard, MINDS-14, SNIPS), every one listed with its licence in [DATA_LICENSES.md](DATA_LICENSES.md). Evaluation test rows were removed from the training data. ## Intended use and limits - Classification-style decisions over short texts and JSON: routing, triage, moderation, intent, yes/no checks, ordinal ratings. It does not generate text. - Accuracy varies by task and language (see the table). Reading comprehension (Belebele) and semantic similarity are weak; do not use it for them without your own evaluation. - Use the confidence: Statim's `min_confidence` option marks low-confidence answers with `escalate: true` so a person can review them. Do not automate high-stakes decisions about people without human review. ## Licence The weights may be used under any one of: PolyForm Noncommercial 1.0.0, PolyForm Small Business 1.0.0 (free commercial use below 100 people and 1 M USD revenue), PolyForm Free Trial 1.0.0 (any company, fewer than 32 days), or a Statim commercial licence ([COMMERCIAL.md](https://github.com/BEKO2210/statim/blob/main/COMMERCIAL.md)). Texts in [LICENSE-MODEL.md](LICENSE-MODEL.md). The Statim engine is Apache-2.0. Built on [Laya](https://huggingface.co/convaiinnovations/laya-multilingual) (Apache-2.0) and [mmBERT-base](https://huggingface.co/jhu-clsp/mmBERT-base) (MIT). Training data attribution: Banking77 (Casanueva et al., 2020, PolyAI), MASSIVE (FitzGerald et al., 2022, Amazon), and the CC-BY sources in DATA_LICENSES.md. Statim is independent and not affiliated with the Laya authors.