Beko2210's picture
Statim Decide Multilingual Base 0.7.0
30fa9ad verified
|
Raw
History Blame Contribute Delete
14.1 kB
metadata
license: other
license_name: statim-weights
license_link: >-
  https://huggingface.co/Beko2210/statim-decide-multilingual-base/blob/main/LICENSE-MODEL.md
language:
  - ar
  - de
  - en
  - es
  - fr
  - hi
  - it
  - ja
  - pl
  - ru
  - tr
  - zh
  - id
  - ms
  - pt
  - nl
  - fa
library_name: gguf
pipeline_tag: zero-shot-classification
base_model: convaiinnovations/laya-multilingual
datasets:
  - PolyAI/banking77
  - AmazonScience/massive
  - LocalLLaMA/typed-decisions
  - tasksource/tasksource-jev-typed-decisions
  - nvidia/Nemotron-Safety-Guard-Dataset-v3
  - l3cube-pune/IndicGuard
  - PolyAI/minds14
  - benayas/snips
tags:
  - statim
  - gguf
  - decision-making
  - text-classification
  - zero-shot-classification
  - mmBERT-base
model-index:
  - name: statim-decide-multilingual-base
    results:
      - task:
          type: text-classification
        dataset:
          name: >-
            typed-decisions (test split, first 2,000 decisions; its train split
            is replay data)
          type: LocalLLaMA/typed-decisions
        metrics:
          - type: accuracy
            value: 0.763
      - task:
          type: text-classification
        dataset:
          name: >-
            Banking77 (test split, first 2,000 rows, all 77 intents in one
            question)
          type: PolyAI/banking77
        metrics:
          - type: accuracy
            value: 0.914
      - task:
          type: text-classification
        dataset:
          name: >-
            MASSIVE intents (mean over 12 languages, 150 seeded stratified test
            rows each)
          type: AmazonScience/massive
        metrics:
          - type: accuracy
            value: 0.7995
      - task:
          type: text-classification
        dataset:
          name: AG News (zero-shot (never trained on), first 2,000 test rows)
          type: fancyzhx/ag_news
        metrics:
          - type: accuracy
            value: 0.9295
      - task:
          type: text-classification
        dataset:
          name: DAIR Emotion (zero-shot, first 2,000 test rows)
          type: dair-ai/emotion
        metrics:
          - type: accuracy
            value: 0.504
      - task:
          type: text-classification
        dataset:
          name: HWU64 intents (English, 150 rows; rows overlapping MASSIVE removed)
          type: hwu64
        metrics:
          - type: accuracy
            value: 0.7867
      - task:
          type: text-classification
        dataset:
          name: SIB-200 topics (zero-shot, mean over 4 languages, 150 rows each)
          type: Davlan/sib200
        metrics:
          - type: accuracy
            value: 0.7817
      - task:
          type: text-classification
        dataset:
          name: Sentiment (zero-shot, mean over 12 languages, 150 rows each)
          type: tyqiangz/multilingual-sentiments
        metrics:
          - type: accuracy
            value: 0.605
      - task:
          type: text-classification
        dataset:
          name: HateCheck (zero-shot, mean over 11 languages, 150 rows each)
          type: mteb/multi-hatecheck
        metrics:
          - type: accuracy
            value: 0.6358
      - task:
          type: text-classification
        dataset:
          name: Belebele reading (zero-shot, mean over 4 languages, 150 rows each)
          type: facebook/belebele
        metrics:
          - type: accuracy
            value: 0.2583

Statim Decide Multilingual Base

A decision model for Statim, the native C++ engine for typed decisions: ask any text a choice, a score or a yes/no question and get calibrated answers from one forward pass, on CPU or GPU, without Python at runtime. Version 0.7.0, fine-tuned from convaiinnovations/laya-multilingual (mmBERT-base encoder).

One support ticket, three typed answers, one forward pass: the 60-second film.

Quick start

# Statim release binary: https://github.com/BEKO2210/statim/releases
huggingface-cli download Beko2210/statim-decide-multilingual-base statim-decide-multilingual-base-q8_0.gguf --local-dir models
./statim serve -m multilingual=models/statim-decide-multilingual-base-q8_0.gguf --port 8080
curl -s localhost:8080/v1/systemone -d '{
  "state": {"subject": "Duplicate charge on invoice #4411",
            "body": "We were billed twice for March. Please refund the second charge."},
  "questions": {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
      "criteria": {"billing": "invoices, refunds", "technical": "bugs, outages", "sales": "pricing, contracts"}},
    "urgency": {"type": "score", "instructions": "How urgent is this request?",
      "criteria": ["not urgent", "soon", "critical"]},
    "refund": {"type": "noul", "instructions": "Does the user explicitly request a refund?"}}}'

Open http://127.0.0.1:8080/ for the playground. API reference: docs/API.md.

Files

File Size Use
statim-decide-multilingual-base-f32.gguf 0.91 GB reference precision, exact on GPU
statim-decide-multilingual-base-q8_0.gguf 0.36 GB recommended for CPU: 4x smaller
checkpoint/ 0.68 GB Laya-format checkpoint for fine-tuning and the Python reference

Checksums in SHA256SUMS. The f32 file reproduces the reference implementation within 1e-4 on Statim's parity tests; q8_0 is smaller and faster on CPU with slightly different logits.

Evaluation

Measured by Statim's no-harm gate (tools/finetune/gate.py) on held-out test data the model selection never looked at. Against the checkpoint it was trained from, on 89 held-out suites: 23 significant gains, 65 within noise, 0 regressions (gains: more than two combined binomial standard errors; regressions: significant after Holm-Bonferroni over all suites; 1 nominal drop beyond two standard errors did not stay significant).

Suite Role This model Base checkpoint Protocol
typed-decisions trained 0.7630 0.7585 test split, first 2,000 decisions; its train split is replay data
Banking77 trained 0.9140 0.9035 test split, first 2,000 rows, all 77 intents in one question
MASSIVE intents trained 0.7995 0.7717 mean over 12 languages, 150 seeded stratified test rows each
AG News held out 0.9295 0.9315 zero-shot (never trained on), first 2,000 test rows
DAIR Emotion held out 0.5040 0.5265 zero-shot, first 2,000 test rows
HWU64 intents held out 0.7867 0.8200 English, 150 rows; rows overlapping MASSIVE removed
SIB-200 topics held out 0.7817 0.7217 zero-shot, mean over 4 languages, 150 rows each
Sentiment held out 0.6050 0.5933 zero-shot, mean over 12 languages, 150 rows each
HateCheck held out 0.6358 0.6467 zero-shot, mean over 11 languages, 150 rows each
Belebele reading held out 0.2583 0.3100 zero-shot, mean over 4 languages, 150 rows each

Decision categories

One held-out suite per decision category, built from splits of the training sources that the mixture never loads; any text that also occurs in the training mixture is dropped. 150 items per language, macro over languages.

Category Languages This model Base checkpoint
complaint en 0.767 0.607
emotion de, en, es, fr, hi, zh 0.586 0.520
fact check en 0.313 0.260
formality ja, tr 0.773 0.503
intent en, nl, tr 0.753 0.516
nli en, ja, tr 0.747 0.704
pii ar, de, en, es, fr, it, ja, nl, ru, sv, zh 0.856 0.595
reading en 0.927 0.560
safety en 0.727 0.600
sentiment en, zh 0.800 0.740
similarity pt 0.833 0.627
stance en 0.893 0.660
topic en 0.607 0.267
urgency en 0.893 0.660

Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768, laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881; Banking77 supervised MPNet 0.941; MASSIVE XLM-R base 0.857 over 12 languages (full train set). Sources: docs/ROADMAP.md.

All 89 held-out suites
Suite Accuracy Rows
amazon_massive_intent/ar 0.7067 150
amazon_massive_intent/de 0.7867 150
amazon_massive_intent/en 0.8267 150
amazon_massive_intent/es 0.8133 150
amazon_massive_intent/fr 0.8200 150
amazon_massive_intent/hi 0.7867 150
amazon_massive_intent/it 0.8200 150
amazon_massive_intent/ja 0.8267 150
amazon_massive_intent/pl 0.8000 150
amazon_massive_intent/ru 0.8333 150
amazon_massive_intent/tr 0.7867 150
amazon_massive_intent/zh-CN 0.7867 150
belebele/ar 0.2933 150
belebele/de 0.2133 150
belebele/en 0.2467 150
belebele/hi 0.2800 150
categories:complaint/en 0.7667 150
categories:emotion/de 0.4533 150
categories:emotion/en 0.5733 150
categories:emotion/es 0.5733 150
categories:emotion/fr 0.6600 150
categories:emotion/hi 0.7133 150
categories:emotion/zh 0.5400 150
categories:fact_check/en 0.3133 150
categories:formality/ja 0.5467 150
categories:formality/tr 1.0000 150
categories:intent/en 0.8200 150
categories:intent/nl 0.4400 150
categories:intent/tr 1.0000 150
categories:nli/en 0.8067 150
categories:nli/ja 0.6333 150
categories:nli/tr 0.8000 150
categories:pii/ar 0.8467 150
categories:pii/de 0.8400 150
categories:pii/en 0.8933 150
categories:pii/es 0.8533 150
categories:pii/fr 0.8933 150
categories:pii/it 0.8400 150
categories:pii/ja 0.8467 150
categories:pii/nl 0.7600 150
categories:pii/ru 0.9067 150
categories:pii/sv 0.8333 150
categories:pii/zh 0.9000 150
categories:reading/en 0.9267 150
categories:safety/en 0.7267 150
categories:sentiment/en 0.8000 150
categories:sentiment/zh 0.8000 150
categories:similarity/pt 0.8333 150
categories:stance/en 0.8933 150
categories:topic/en 0.6067 150
categories:urgency/en 0.8933 150
farstail/fa 0.7000 150
go_emotions/en 0.4733 150
hwu64/en 0.7867 150
indonli/id 0.6667 150
multi_hatecheck/ar 0.5733 150
multi_hatecheck/de 0.6867 150
multi_hatecheck/en 0.6133 150
multi_hatecheck/es 0.6200 150
multi_hatecheck/fr 0.6867 150
multi_hatecheck/hi 0.5733 150
multi_hatecheck/it 0.6600 150
multi_hatecheck/nl 0.6533 150
multi_hatecheck/pl 0.6333 150
multi_hatecheck/pt 0.6467 150
multi_hatecheck/zh 0.6467 150
multilingual_sentiments/ar 0.5600 150
multilingual_sentiments/de 0.5800 150
multilingual_sentiments/en 0.6933 150
multilingual_sentiments/es 0.5267 150
multilingual_sentiments/fr 0.5867 150
multilingual_sentiments/hi 0.5333 150
multilingual_sentiments/id 0.7467 150
multilingual_sentiments/it 0.5667 150
multilingual_sentiments/ja 0.6600 150
multilingual_sentiments/ms 0.5333 150
multilingual_sentiments/pt 0.6333 150
multilingual_sentiments/zh 0.6400 150
semrel/ar 0.2467 150
semrel/en 0.2000 150
semrel/hi 0.2267 150
sib200/ar 0.7867 150
sib200/de 0.8000 150
sib200/en 0.8333 150
sib200/hi 0.7067 150
test/ag_news 0.9295 2000
test/banking77 0.9140 2000
test/emotion 0.5040 2000
test/typed_decisions 0.7630 2000

Reproduce these numbers: REPRODUCE.md.

Training

Multi-task fine-tuning with train_multitask.py --clean from checkpoint laya-multilingual-big1, best epoch 12/raw selected on validation data only. Training data: only sources whose licence permits commercial use and imposes no ShareAlike or copyleft terms (Banking77, MASSIVE, typed-decisions replay, a licence-audited tasksource mixture, Nemotron-Safety-Guard, IndicGuard, MINDS-14, SNIPS), every one listed with its licence in DATA_LICENSES.md. Evaluation test rows were removed from the training data.

Intended use and limits

  • Classification-style decisions over short texts and JSON: routing, triage, moderation, intent, yes/no checks, ordinal ratings. It does not generate text.
  • Accuracy varies by task and language (see the table). Reading comprehension (Belebele) and semantic similarity are weak; do not use it for them without your own evaluation.
  • Use the confidence: Statim's min_confidence option marks low-confidence answers with escalate: true so a person can review them. Do not automate high-stakes decisions about people without human review.

Licence

The weights may be used under any one of: PolyForm Noncommercial 1.0.0, PolyForm Small Business 1.0.0 (free commercial use below 100 people and 1 M USD revenue), PolyForm Free Trial 1.0.0 (any company, fewer than 32 days), or a Statim commercial licence (COMMERCIAL.md). Texts in LICENSE-MODEL.md. The Statim engine is Apache-2.0.

Built on Laya (Apache-2.0) and mmBERT-base (MIT). Training data attribution: Banking77 (Casanueva et al., 2020, PolyAI), MASSIVE (FitzGerald et al., 2022, Amazon), and the CC-BY sources in DATA_LICENSES.md. Statim is independent and not affiliated with the Laya authors.