Beko2210 commited on
Commit
30fa9ad
·
verified ·
1 Parent(s): a6bea4f

Statim Decide Multilingual Base 0.7.0

Browse files
DATA_LICENSES.md CHANGED
@@ -200,6 +200,124 @@ Licence as recorded per row by tasksource; source names follow its build manifes
200
 
201
  Filter statistics of this build: rows 2,500,000, not commercial 1,124,896, audit excluded 489,942, not permissive 219,847, too long 6,489, test duplicate 7.
202
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
203
  ## Used only for evaluation (never trained on)
204
 
205
  | Dataset | Licence |
 
200
 
201
  Filter statistics of this build: rows 2,500,000, not commercial 1,124,896, audit excluded 489,942, not permissive 219,847, too long 6,489, test duplicate 7.
202
 
203
+ ### Mixture v6 sources (111 sources, 534,231 items, at most 6200 per source)
204
+
205
+ Licence checked at the source for every entry (registry `tools/finetune/sources/v6-keep.json`, review notes and attribution in `tools/finetune/sources/v6-research.md`). Every text that occurs in an evaluation suite was removed first (862,438 banned texts).
206
+
207
+ | Source | Category | Licence | Languages | Items |
208
+ |---|---|---|---|---:|
209
+ | [3nesdeniz/agentic-prompt-injection-5k](https://huggingface.co/datasets/3nesdeniz/agentic-prompt-injection-5k) | prompt-injection | CC-BY-4.0 | en | 6,200 |
210
+ | [3nesdeniz/turkish-conversation-prompt-injection](https://huggingface.co/datasets/3nesdeniz/turkish-conversation-prompt-injection) | prompt-injection | CC-BY-4.0 | tr | 530 |
211
+ | [Adilbai/kz-gov-complaints-data-kz-ru](https://huggingface.co/datasets/Adilbai/kz-gov-complaints-data-kz-ru) | complaint, sentiment | Apache-2.0 | ru, kk | 1,200 |
212
+ | [adiprog14/lingrow-support-tickets](https://huggingface.co/datasets/adiprog14/lingrow-support-tickets) | complaint | MIT | en | 6,200 |
213
+ | [ai4bharat/IndicSentiment](https://huggingface.co/datasets/ai4bharat/IndicSentiment) | sentiment | CC0-1.0 (AI4Bharat/IndicBERT README; HF card has none) | en, hi, bn, mr, ta, te, ur, gu +6 | 6,200 |
214
+ | [alaminxpro/university-students-complaints](https://huggingface.co/datasets/alaminxpro/university-students-complaints) | complaint | CC-BY-4.0 | en | 546 |
215
+ | [allenai/prosocial-dialog](https://huggingface.co/datasets/allenai/prosocial-dialog) | safety-moderation | CC-BY-4.0 | en | 6,200 |
216
+ | [alusci/sms-otp-spam-dataset](https://huggingface.co/datasets/alusci/sms-otp-spam-dataset) | spam-sms | MIT (templated; low value) | en | 6,200 |
217
+ | [amyrmahdy/decima-synthetic-decisions](https://huggingface.co/datasets/amyrmahdy/decima-synthetic-decisions) | typed-decisions | cc-by-4.0 (card; fully synthetic, teacher Gemma-4-26B-A4B-it, Apache-2.0 model card) | en, fa, ar, ru | 6,200 |
218
+ | [ankitkupadhyay/XNLI](https://huggingface.co/datasets/ankitkupadhyay/XNLI) | nli | apache-2.0 (card); content inherits MNLI/OANC terms | ar, bg, de, el, en, es, fr, hi +7 | 6,200 |
219
+ | [Anthropic/hh-rlhf](https://huggingface.co/datasets/Anthropic/hh-rlhf) | safety-moderation | MIT | en | 6,200 |
220
+ | [AshenFdo/synthetic_blood_request_urgency_dataset](https://huggingface.co/datasets/AshenFdo/synthetic_blood_request_urgency_dataset) | urgency | mit | en | 2,500 |
221
+ | [Avature/Job-Title-Similarity](https://huggingface.co/datasets/Avature/Job-Title-Similarity) | similarity | apache-2.0 | de, en, es, fr, it, ja, nl, pl +3 | 4,563 |
222
+ | [BEE-spoke-data/consumer-finance-complaints](https://huggingface.co/datasets/BEE-spoke-data/consumer-finance-complaints) | complaint | CC0-1.0 card; CFPB US federal data, narratives published with consumer opt-in consent | en | 6,200 |
223
+ | [boun-tabi/nli_tr](https://huggingface.co/datasets/boun-tabi/nli_tr) | nli | same terms as MultiNLI (GitHub boun-tabi/NLI-TR README) | tr | 6,200 |
224
+ | [brighter-dataset/BRIGHTER-emotion-categories](https://huggingface.co/datasets/brighter-dataset/BRIGHTER-emotion-categories) | emotion | cc-by-4.0 | hi | 3,629 |
225
+ | [brighter-dataset/BRIGHTER-emotion-categories](https://huggingface.co/datasets/brighter-dataset/BRIGHTER-emotion-categories) | emotion | cc-by-4.0 | mr | 3,768 |
226
+ | [clips/VaccinChatNL](https://huggingface.co/datasets/clips/VaccinChatNL) | intent-dialogue-act | CC-BY-4.0 | nl | 6,200 |
227
+ | [cngchis/Support-Ticket-Router-12K-Cleaned](https://huggingface.co/datasets/cngchis/Support-Ticket-Router-12K-Cleaned) | complaint | Apache-2.0 | en | 6,200 |
228
+ | [CohereForAI/aya_redteaming](https://huggingface.co/datasets/CohereForAI/aya_redteaming) | safety-moderation | Apache-2.0 | en, fr, es, ru, ar, hi, sr, tl | 494 |
229
+ | [CohereLabs/aya_dataset](https://huggingface.co/datasets/CohereLabs/aya_dataset) | language-id | apache-2.0 | 65 incl. ar, de, en, fr, hi, it, ja, nl, pl, pt, ru, es, tr, zh | 6,200 |
230
+ | [community-datasets/re_dial](https://huggingface.co/datasets/community-datasets/re_dial) | sentiment | CC-BY-4.0 | en | 6,200 |
231
+ | [community-datasets/tapaco](https://huggingface.co/datasets/community-datasets/tapaco) | similarity | cc-by-2.0 (Tatoeba CC-BY 2.0 FR) | en, de, fr, es, it, pt, nl, pl +6 | 6,200 |
232
+ | [Console-AI/IT-helpdesk-synthetic-tickets](https://huggingface.co/datasets/Console-AI/IT-helpdesk-synthetic-tickets) | complaint, urgency | MIT | en | 1,000 |
233
+ | ConvLab/crosswoz (github thu-coai/CrossWOZ) | intent-dialogue-act | Apache-2.0 | zh | 6,200 |
234
+ | [ddrg/super_eurlex](https://huggingface.co/datasets/ddrg/super_eurlex) | topic | MIT card; EUR-Lex reuse authorised incl. commercial with attribution (Decision 2011/833/EU) | bg, cs, da, de, el, en, es, et +16 | 6,200 |
235
+ | [declare-lab/CategoricalHarmfulQA](https://huggingface.co/datasets/declare-lab/CategoricalHarmfulQA) | safety-moderation | Apache-2.0 | en, zh, vi | 550 |
236
+ | [dell-research-harvard/headlines-semantic-similarity](https://huggingface.co/datasets/dell-research-harvard/headlines-semantic-similarity) | similarity | cc-by-2.0 (off-copyright US newspapers) | en | 6,200 |
237
+ | [dhruv0808/indic_sentiment_analyzer](https://huggingface.co/datasets/dhruv0808/indic_sentiment_analyzer) | sentiment | CC-BY-4.0 | en, hi, te, ta, kn, or, bn, gu +4 | 6,200 |
238
+ | [dvgodoy/CUAD_v1_Contract_Understanding_clause_classification](https://huggingface.co/datasets/dvgodoy/CUAD_v1_Contract_Understanding_clause_classification) | topic | CC-BY-4.0 | en | 6,200 |
239
+ | [E3-JSI/synthetic-multi-pii-ner-v1](https://huggingface.co/datasets/E3-JSI/synthetic-multi-pii-ner-v1) | pii | mit | en, fr, de, el, nl, it, sl | 2,971 |
240
+ | [elvanalabs/sarcasm-statements-90](https://huggingface.co/datasets/elvanalabs/sarcasm-statements-90) | sarcasm | mit | en | 90 |
241
+ | [Fumika/Wikinews-multilingual](https://huggingface.co/datasets/Fumika/Wikinews-multilingual) | topic | CC-BY-2.5 (Wikinews) | en, es, fr, de, pt, pl, it, zh +25 | 6,200 |
242
+ | [gfissore/arxiv-abstracts-2021](https://huggingface.co/datasets/gfissore/arxiv-abstracts-2021) | topic | CC0-1.0 (arXiv metadata) | en | 6,200 |
243
+ | [asappresearch/abcd](https://github.com/asappresearch/abcd) | complaint | MIT (GitHub LICENSE) | en | 6,200 |
244
+ | [bvidgen/Dynamically-Generated-Hate-Speech-Dataset (v0.2.3.csv; NOT tasksource/dynahate mirror tagged gpl)](https://github.com/bvidgen/Dynamically-Generated-Hate-Speech-Dataset) | toxicity-hate | CC-BY-4.0 (upstream README) | en | 6,200 |
245
+ | [HLTCHKUST/BiToD (mirror DeepPavlov/BiToD)](https://github.com/HLTCHKUST/BiToD) | intent-dialogue-act | Apache-2.0 | en, zh | 6,200 |
246
+ | [PolyAI-LDN/task-specific-datasets/nlupp](https://github.com/PolyAI-LDN/task-specific-datasets) | intent-dialogue-act | CC-BY-4.0 | en | 705 |
247
+ | [wwbp/empathic_reactions](https://github.com/wwbp/empathic_reactions) | emotion | cc-by-4.0 | en | 3,719 |
248
+ | [GoktugD/turkish-formality-rewrite-500k](https://huggingface.co/datasets/GoktugD/turkish-formality-rewrite-500k) | formality | cc0-1.0 | tr | 6,200 |
249
+ | [GoktugD/turkish-intent-classification-1m](https://huggingface.co/datasets/GoktugD/turkish-intent-classification-1m) | intent-dialogue-act | CC0-1.0 (template-generated) | tr | 6,200 |
250
+ | [GoktugD/turkish-nli-constructed-1.5m](https://huggingface.co/datasets/GoktugD/turkish-nli-constructed-1.5m) | nli | cc0-1.0 | tr | 6,200 |
251
+ | [google-research-datasets/poem_sentiment](https://huggingface.co/datasets/google-research-datasets/poem_sentiment) | sentiment | CC-BY-4.0 | en | 892 |
252
+ | google-research-datasets/taskmaster1/2/3 (github Taskmaster TM-1..TM-4) | intent-dialogue-act | CC-BY-4.0 | en | 6,200 |
253
+ | [gretelai/gretel-pii-masking-en-v1](https://huggingface.co/datasets/gretelai/gretel-pii-masking-en-v1) | pii | apache-2.0 | en | 6,200 |
254
+ | [gretelai/synthetic_pii_finance_multilingual](https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual) | pii | apache-2.0 | en, fr, de, nl, es, it, sv | 6,200 |
255
+ | [hblim/customer-complaints](https://huggingface.co/datasets/hblim/customer-complaints) | complaint | MIT | en | 1,260 |
256
+ | [Helsinki-NLP/tatoeba](https://huggingface.co/datasets/Helsinki-NLP/tatoeba) | language-id | cc-by-2.0 | 300+ incl. all priority | 6,200 |
257
+ | [Helsinki-NLP/tatoeba](https://huggingface.co/datasets/Helsinki-NLP/tatoeba) | similarity | cc-by-2.0 (Tatoeba CC-BY 2.0 FR) | en, de, fr, es, it, pt, nl, pl +6 | 6,200 |
258
+ | [ibm-research/AttaQ](https://huggingface.co/datasets/ibm-research/AttaQ) | safety-moderation | MIT | en | 1,402 |
259
+ | [IDinsight/urgency_detection_maternal_health_synthetic](https://huggingface.co/datasets/IDinsight/urgency_detection_maternal_health_synthetic) | urgency | mit | en | 6,200 |
260
+ | jagoldz/gahd (filter via GitHub jagol/gahd gahd_disaggregated.csv) | toxicity-hate | CC-BY-4.0 | de | 5,441 |
261
+ | [jmccardle/pulse-sofroniew-emotion-concept-texts](https://huggingface.co/datasets/jmccardle/pulse-sofroniew-emotion-concept-texts) | emotion | cc-by-4.0 | en | 6,200 |
262
+ | [joelniklaus/covid19_emergency_event](https://huggingface.co/datasets/joelniklaus/covid19_emergency_event) | topic | CC0-1.0 | en, fr, hu, it, nb, nl, pl | 1,202 |
263
+ | [joelniklaus/german_argument_mining](https://huggingface.co/datasets/joelniklaus/german_argument_mining) | argument-mining | cc-by-4.0 | de | 6,200 |
264
+ | [Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset](https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset) | emotion | mit | zh | 6,200 |
265
+ | [JusteLeo/French-emotion](https://huggingface.co/datasets/JusteLeo/French-emotion) | emotion | mit | fr | 6,200 |
266
+ | [kchawla123/casino](https://huggingface.co/datasets/kchawla123/casino) | intent-dialogue-act | CC-BY-4.0 | en | 3,643 |
267
+ | [Kenshiii/synthetic-product-reviews](https://huggingface.co/datasets/Kenshiii/synthetic-product-reviews) | sentiment | CC-BY-4.0 | en | 687 |
268
+ | [KhiredNetworks/synthetic-product-reviews](https://huggingface.co/datasets/KhiredNetworks/synthetic-product-reviews) | sentiment | MIT | en | 6,200 |
269
+ | [leonvanbokhorst/synthetic-complaints-v2](https://huggingface.co/datasets/leonvanbokhorst/synthetic-complaints-v2) | complaint, sentiment | MIT | en | 6,200 |
270
+ | [liri-uzh/cfpb-complaints-mini](https://huggingface.co/datasets/liri-uzh/cfpb-complaints-mini) | complaint | CC0-1.0 (CFPB public domain) | en | 6,200 |
271
+ | [llm-for-emotion/Cultural-Emo](https://huggingface.co/datasets/llm-for-emotion/Cultural-Emo) | emotion | mit | ar, de, en, hi, es | 3,999 |
272
+ | [lyon-nlp/clustering-hal-s2s](https://huggingface.co/datasets/lyon-nlp/clustering-hal-s2s) | topic | Apache-2.0 card; HAL metadata CC0 | fr | 6,200 |
273
+ | [masakhane/InjongoIntent](https://huggingface.co/datasets/masakhane/InjongoIntent) | intent-dialogue-act | Apache-2.0 | en, am, ee, ha, ig, rw, ln, lg +9 | 6,200 |
274
+ | [matsuxr/JaGovFaqs-22k](https://huggingface.co/datasets/matsuxr/JaGovFaqs-22k) | similarity | cc-by-4.0 (Japanese government copyright policy) | ja | 6,200 |
275
+ | [maximoss/mnli-nineeleven-fr](https://huggingface.co/datasets/maximoss/mnli-nineeleven-fr) | nli | bsd-2-clause | fr | 3,988 |
276
+ | [MoritzLaurer/synthetic_zeroshot_mixtral_v0.1](https://huggingface.co/datasets/MoritzLaurer/synthetic_zeroshot_mixtral_v0.1) | nli | apache-2.0 (Mixtral-8x7B outputs) | en | 6,200 |
277
+ | [mteb/toxic_conversations_50k](https://huggingface.co/datasets/mteb/toxic_conversations_50k) | toxicity-hate | CC-BY-4.0 (Civil Comments text CC0) | en | 6,200 |
278
+ | [NABA-AI/LUB-Saudi-Arabic-Intent](https://huggingface.co/datasets/NABA-AI/LUB-Saudi-Arabic-Intent) | intent-dialogue-act | CC-BY-4.0 (synthetic) | ar | 1,500 |
279
+ | [NagaYu/deference-keigo-corpus](https://huggingface.co/datasets/NagaYu/deference-keigo-corpus) | formality | cc-by-4.0 | ja | 3,186 |
280
+ | [napsternxg/wands](https://huggingface.co/datasets/napsternxg/wands) | similarity | MIT (wayfair/WANDS) | en | 6,200 |
281
+ | [NortheasternUniversity/big_patent](https://huggingface.co/datasets/NortheasternUniversity/big_patent) | topic | CC-BY-4.0 | en | 6,200 |
282
+ | [Novora/Tri-Class-Sentiment-Synthetic](https://huggingface.co/datasets/Novora/Tri-Class-Sentiment-Synthetic) | sentiment | CC0-1.0 | en | 6,200 |
283
+ | [nvidia/Aegis-AI-Content-Safety-Dataset-1.0](https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0) | safety-moderation | CC-BY-4.0 | en | 6,200 |
284
+ | [nvidia/Aegis-AI-Content-Safety-Dataset-2.0](https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0) | safety-moderation | CC-BY-4.0 | en | 6,200 |
285
+ | [nvidia/CantTalkAboutThis-Topic-Control-Dataset](https://huggingface.co/datasets/nvidia/CantTalkAboutThis-Topic-Control-Dataset) | safety-topic-control | CC-BY-4.0 | en | 1,073 |
286
+ | [nvidia/Nemotron-PII](https://huggingface.co/datasets/nvidia/Nemotron-PII) | pii | cc-by-4.0 | en | 6,200 |
287
+ | [nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1](https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1) | prompt-injection (indirect) | CC-BY-4.0 | en | 1,220 |
288
+ | [nyu-mll/multi_nli](https://huggingface.co/datasets/nyu-mll/multi_nli) | nli | OANC licence (permissive, commercial OK) per MNLI paper/card; fiction genre mixed incl. CC-BY-SA-3.0 | en | 6,200 |
289
+ | [OpenAssistant/oasst2](https://huggingface.co/datasets/OpenAssistant/oasst2) | toxicity-moderation | Apache-2.0 | en, es, ru, zh, de, fr, pt, it +4 | 6,200 |
290
+ | [OpenSafetyLab/Salad-Data](https://huggingface.co/datasets/OpenSafetyLab/Salad-Data) | safety-moderation | Apache-2.0 | en | 6,200 |
291
+ | [OrSabbach/food-delivery-support-tickets](https://huggingface.co/datasets/OrSabbach/food-delivery-support-tickets) | complaint | MIT | en | 6,200 |
292
+ | [pacoreyes/StanceSentences](https://huggingface.co/datasets/pacoreyes/StanceSentences) | stance | apache-2.0 | en | 972 |
293
+ | pfb30/multi_woz_v22 (github budzianowski/multiwoz) | intent-dialogue-act | MIT (upstream); Apache-2.0 (card) | en | 6,200 |
294
+ | [PolyAI/minds14](https://huggingface.co/datasets/PolyAI/minds14) | complaint | CC-BY-4.0 | cs, de, en, es, fr, it, ko, nl +4 | 6,200 |
295
+ | [Process-Venue/IntentClassification_Dataset_for_AI_Assistant_Prompt_Routing_Hindi](https://huggingface.co/datasets/Process-Venue/IntentClassification_Dataset_for_AI_Assistant_Prompt_Routing_Hindi) | intent-dialogue-act | Apache-2.0 (text provenance undocumented) | hi | 4,998 |
296
+ | [reshabhs/SPML_Chatbot_Prompt_Injection](https://huggingface.co/datasets/reshabhs/SPML_Chatbot_Prompt_Injection) | prompt-injection | MIT | en | 6,200 |
297
+ | [RichardSakaguchiMS/brazilian-customer-service-conversations](https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations) | complaint, sentiment | Apache-2.0 | pt | 1,510 |
298
+ | [s2pidape/support-ticket-dataset](https://huggingface.co/datasets/s2pidape/support-ticket-dataset) | complaint | CC-BY-4.0 | en | 6,200 |
299
+ | [shreyaspullehf/emotion-dataset-20-emotions](https://huggingface.co/datasets/shreyaspullehf/emotion-dataset-20-emotions) | emotion | mit | en | 6,200 |
300
+ | [sileod/attempto-nli](https://huggingface.co/datasets/sileod/attempto-nli) | nli | apache-2.0 | en | 6,200 |
301
+ | [SINAI/ALIA-es-discriminative-stance-detection](https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-stance-detection) | stance | cc-by-4.0 | es | 2,850 |
302
+ | [stjiris/IRIS_sts](https://huggingface.co/datasets/stjiris/IRIS_sts) | similarity | mit | pt | 3,334 |
303
+ | [sutro/synthetic-product-reviews-20k](https://huggingface.co/datasets/sutro/synthetic-product-reviews-20k) | sentiment | MIT | en | 6,200 |
304
+ | [sweatSmile/sarcastic-dataset](https://huggingface.co/datasets/sweatSmile/sarcastic-dataset) | sarcasm | mit | en | 1,440 |
305
+ | [takehika/wanli-ja-nli](https://huggingface.co/datasets/takehika/wanli-ja-nli) | nli | cc-by-4.0 | ja | 6,200 |
306
+ | [tanaos/synthetic-emotion-detection-dataset-v1](https://huggingface.co/datasets/tanaos/synthetic-emotion-detection-dataset-v1) | emotion | mit | en | 6,200 |
307
+ | [tanaos/synthetic-sentiment-analysis-dataset-v1](https://huggingface.co/datasets/tanaos/synthetic-sentiment-analysis-dataset-v1) | sentiment | MIT | en | 6,200 |
308
+ | [tasksource/esci](https://huggingface.co/datasets/tasksource/esci) | similarity | apache-2.0 (amazon-science/esci-data) | en, es, ja | 6,200 |
309
+ | [tasksource/help-desk-tickets](https://huggingface.co/datasets/tasksource/help-desk-tickets) | complaint, urgency | CC-BY-4.0 (Mendeley btm76zndnt v3) | en, mixed | 357 |
310
+ | [tasksource/it-support-tickets](https://huggingface.co/datasets/tasksource/it-support-tickets) | complaint | CC-BY-4.0 (Zenodo 7648117) | en, de, pt, es | 1,568 |
311
+ | [theatticusproject/cuad-qa](https://huggingface.co/datasets/theatticusproject/cuad-qa) | reading-comprehension | CC-BY-4.0 | en | 6,200 |
312
+ | [theatticusproject/maud](https://huggingface.co/datasets/theatticusproject/maud) | reading-comprehension | CC-BY-4.0 | en | 6,200 |
313
+ | [uoe-nlp/multi3-nlu](https://huggingface.co/datasets/uoe-nlp/multi3-nlu) | intent-dialogue-act | CC-BY-4.0 | am, mr, tr, es | 5,636 |
314
+ | [urchade/synthetic-pii-ner-mistral-v1](https://huggingface.co/datasets/urchade/synthetic-pii-ner-mistral-v1) | pii | apache-2.0 | en, fr, it, de, es | 6,200 |
315
+ | [vic35get/nhtsa_complaints_dataset](https://huggingface.co/datasets/vic35get/nhtsa_complaints_dataset) | complaint | Apache-2.0 card; NHTSA US federal data | en | 6,200 |
316
+ | [Wismut/nym-pii-multilingual-data](https://huggingface.co/datasets/Wismut/nym-pii-multilingual-data) | pii | mit | en, de, fr, es, it, pt, nl, pl +14 | 6,200 |
317
+ | [WorkInTheDark/FairytaleQA](https://huggingface.co/datasets/WorkInTheDark/FairytaleQA) | reading-comprehension | Apache-2.0 | en | 6,200 |
318
+ | [YiMeng-SYSU/chinese-logic-sentiment-dataset](https://huggingface.co/datasets/YiMeng-SYSU/chinese-logic-sentiment-dataset) | sentiment | Apache-2.0 | zh | 2,176 |
319
+ | [3609356 (ClaimBuster)](https://zenodo.org/records/3609356) | claim-detection | cc-by-4.0 | en | 1,032 |
320
+
321
  ## Used only for evaluation (never trained on)
322
 
323
  | Dataset | Licence |
README.md CHANGED
@@ -47,90 +47,90 @@ model-index:
47
  dataset:
48
  name: typed-decisions (test split, first 2,000 decisions; its train split is
49
  replay data)
50
- type: typed-decisions
51
  metrics:
52
  - type: accuracy
53
- value: 0.7585
54
  - task:
55
  type: text-classification
56
  dataset:
57
  name: Banking77 (test split, first 2,000 rows, all 77 intents in one question)
58
- type: banking77
59
  metrics:
60
  - type: accuracy
61
- value: 0.9035
62
  - task:
63
  type: text-classification
64
  dataset:
65
  name: MASSIVE intents (mean over 12 languages, 150 seeded stratified test rows
66
  each)
67
- type: massive_intents
68
  metrics:
69
  - type: accuracy
70
- value: 0.7717
71
  - task:
72
  type: text-classification
73
  dataset:
74
  name: AG News (zero-shot (never trained on), first 2,000 test rows)
75
- type: ag_news
76
  metrics:
77
  - type: accuracy
78
- value: 0.9315
79
  - task:
80
  type: text-classification
81
  dataset:
82
  name: DAIR Emotion (zero-shot, first 2,000 test rows)
83
- type: dair_emotion
84
  metrics:
85
  - type: accuracy
86
- value: 0.5265
87
  - task:
88
  type: text-classification
89
  dataset:
90
  name: HWU64 intents (English, 150 rows; rows overlapping MASSIVE removed)
91
- type: hwu64_intents
92
  metrics:
93
  - type: accuracy
94
- value: 0.82
95
  - task:
96
  type: text-classification
97
  dataset:
98
  name: SIB-200 topics (zero-shot, mean over 4 languages, 150 rows each)
99
- type: sib-200_topics
100
  metrics:
101
  - type: accuracy
102
- value: 0.7217
103
  - task:
104
  type: text-classification
105
  dataset:
106
  name: Sentiment (zero-shot, mean over 12 languages, 150 rows each)
107
- type: sentiment
108
  metrics:
109
  - type: accuracy
110
- value: 0.5933
111
  - task:
112
  type: text-classification
113
  dataset:
114
  name: HateCheck (zero-shot, mean over 11 languages, 150 rows each)
115
- type: hatecheck
116
  metrics:
117
  - type: accuracy
118
- value: 0.6467
119
  - task:
120
  type: text-classification
121
  dataset:
122
  name: Belebele reading (zero-shot, mean over 4 languages, 150 rows each)
123
- type: belebele_reading
124
  metrics:
125
  - type: accuracy
126
- value: 0.31
127
  ---
128
 
129
  # Statim Decide Multilingual Base
130
 
131
  A decision model for [Statim](https://github.com/BEKO2210/statim), the native C++ engine for typed decisions: ask any text a
132
  **choice**, a **score** or a **yes/no** question and get calibrated answers from one forward pass, on
133
- CPU or GPU, without Python at runtime. Version **0.4.0**, fine-tuned from
134
  [`convaiinnovations/laya-multilingual`](https://huggingface.co/convaiinnovations/laya-multilingual) (mmBERT-base encoder).
135
 
136
  <video controls preload="none" width="100%" poster="https://beko2210.github.io/statim/images/film-16x9.webp" src="https://beko2210.github.io/statim/video/statim-flagship-60s-16x9.mp4"></video>
@@ -170,84 +170,140 @@ Statim's parity tests; q8_0 is smaller and faster on CPU with slightly different
170
  ## Evaluation
171
 
172
  Measured by Statim's no-harm gate ([`tools/finetune/gate.py`](https://github.com/BEKO2210/statim/blob/main/tools/finetune/gate.py))
173
- on held-out test data the model selection never looked at. Against the checkpoint it was trained from, on 54 held-out suites: **18 significant gains, 36 within noise, 0 regressions** (two combined binomial standard errors).
174
 
175
  | Suite | Role | This model | Base checkpoint | Protocol |
176
  |---|---|---|---|---|
177
- | typed-decisions | trained | **0.7585** | 0.3510 | test split, first 2,000 decisions; its train split is replay data |
178
- | Banking77 | trained | **0.9035** | 0.5175 | test split, first 2,000 rows, all 77 intents in one question |
179
- | MASSIVE intents | trained | **0.7717** | 0.3400 | mean over 12 languages, 150 seeded stratified test rows each |
180
- | AG News | held out | **0.9315** | 0.9380 | zero-shot (never trained on), first 2,000 test rows |
181
- | DAIR Emotion | held out | **0.5265** | 0.5320 | zero-shot, first 2,000 test rows |
182
- | HWU64 intents | held out | **0.8200** | 0.5000 | English, 150 rows; rows overlapping MASSIVE removed |
183
- | SIB-200 topics | held out | **0.7217** | 0.7484 | zero-shot, mean over 4 languages, 150 rows each |
184
- | Sentiment | held out | **0.5933** | 0.5595 | zero-shot, mean over 12 languages, 150 rows each |
185
- | HateCheck | held out | **0.6467** | 0.6303 | zero-shot, mean over 11 languages, 150 rows each |
186
- | Belebele reading | held out | **0.3100** | 0.2467 | zero-shot, mean over 4 languages, 150 rows each |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
187
 
188
  Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768,
189
  laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881;
190
- Banking77 supervised MPNet 0.941; MASSIVE XLM-R base 0.857 (full train set). Sources:
191
  [docs/ROADMAP.md](https://github.com/BEKO2210/statim/blob/main/docs/ROADMAP.md).
192
 
193
- <details><summary>All 54 held-out suites</summary>
194
 
195
  | Suite | Accuracy | Rows |
196
  |---|---|---|
197
- | `amazon_massive_intent/ar` | 0.6600 | 150 |
198
- | `amazon_massive_intent/de` | 0.7667 | 150 |
199
- | `amazon_massive_intent/en` | 0.8133 | 150 |
200
- | `amazon_massive_intent/es` | 0.7733 | 150 |
201
- | `amazon_massive_intent/fr` | 0.7600 | 150 |
202
- | `amazon_massive_intent/hi` | 0.7600 | 150 |
203
- | `amazon_massive_intent/it` | 0.7667 | 150 |
204
- | `amazon_massive_intent/ja` | 0.7800 | 150 |
205
  | `amazon_massive_intent/pl` | 0.8000 | 150 |
206
- | `amazon_massive_intent/ru` | 0.8200 | 150 |
207
- | `amazon_massive_intent/tr` | 0.7600 | 150 |
208
- | `amazon_massive_intent/zh-CN` | 0.8000 | 150 |
209
- | `belebele/ar` | 0.2467 | 150 |
210
- | `belebele/de` | 0.3400 | 150 |
211
- | `belebele/en` | 0.3467 | 150 |
212
- | `belebele/hi` | 0.3067 | 150 |
213
- | `farstail/fa` | 0.6333 | 150 |
214
- | `go_emotions/en` | 0.4667 | 150 |
215
- | `hwu64/en` | 0.8200 | 150 |
216
- | `indonli/id` | 0.7000 | 150 |
217
- | `multi_hatecheck/ar` | 0.5800 | 150 |
218
- | `multi_hatecheck/de` | 0.6667 | 150 |
219
- | `multi_hatecheck/en` | 0.6333 | 150 |
220
- | `multi_hatecheck/es` | 0.6733 | 150 |
221
- | `multi_hatecheck/fr` | 0.6667 | 150 |
222
- | `multi_hatecheck/hi` | 0.5600 | 150 |
223
- | `multi_hatecheck/it` | 0.6867 | 150 |
224
- | `multi_hatecheck/nl` | 0.6000 | 150 |
225
- | `multi_hatecheck/pl` | 0.6400 | 150 |
226
- | `multi_hatecheck/pt` | 0.7067 | 150 |
227
- | `multi_hatecheck/zh` | 0.7000 | 150 |
228
- | `multilingual_sentiments/ar` | 0.5533 | 150 |
229
- | `multilingual_sentiments/de` | 0.5600 | 150 |
230
- | `multilingual_sentiments/en` | 0.6400 | 150 |
231
- | `multilingual_sentiments/es` | 0.5467 | 150 |
232
- | `multilingual_sentiments/fr` | 0.5533 | 150 |
233
- | `multilingual_sentiments/hi` | 0.4600 | 150 |
234
- | `multilingual_sentiments/id` | 0.7133 | 150 |
235
- | `multilingual_sentiments/it` | 0.6000 | 150 |
236
- | `multilingual_sentiments/ja` | 0.6733 | 150 |
237
- | `multilingual_sentiments/ms` | 0.5267 | 150 |
238
- | `multilingual_sentiments/pt` | 0.6600 | 150 |
239
- | `multilingual_sentiments/zh` | 0.6333 | 150 |
240
- | `semrel/ar` | 0.2400 | 150 |
241
- | `semrel/en` | 0.2200 | 150 |
242
- | `semrel/hi` | 0.2200 | 150 |
243
- | `sib200/ar` | 0.7067 | 150 |
244
- | `sib200/de` | 0.7267 | 150 |
245
- | `sib200/en` | 0.7467 | 150 |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
246
  | `sib200/hi` | 0.7067 | 150 |
247
- | `test/ag_news` | 0.9315 | 2000 |
248
- | `test/banking77` | 0.9035 | 2000 |
249
- | `test/emotion` | 0.5265 | 2000 |
250
- | `test/typed_decisions` | 0.7585 | 2000 |
251
 
252
  </details>
253
 
@@ -255,8 +311,8 @@ Reproduce these numbers: [REPRODUCE.md](https://github.com/BEKO2210/statim/blob/
255
 
256
  ## Training
257
 
258
- Multi-task fine-tuning with `train_multitask.py --clean` from checkpoint `laya-multilingual`, best
259
- epoch `19/raw` selected on validation data only. Training data: only sources whose
260
  licence permits commercial use and imposes no ShareAlike or copyleft terms (Banking77, MASSIVE,
261
  typed-decisions replay, a licence-audited tasksource mixture, Nemotron-Safety-Guard, IndicGuard,
262
  MINDS-14, SNIPS), every one listed with its licence in
 
47
  dataset:
48
  name: typed-decisions (test split, first 2,000 decisions; its train split is
49
  replay data)
50
+ type: LocalLLaMA/typed-decisions
51
  metrics:
52
  - type: accuracy
53
+ value: 0.763
54
  - task:
55
  type: text-classification
56
  dataset:
57
  name: Banking77 (test split, first 2,000 rows, all 77 intents in one question)
58
+ type: PolyAI/banking77
59
  metrics:
60
  - type: accuracy
61
+ value: 0.914
62
  - task:
63
  type: text-classification
64
  dataset:
65
  name: MASSIVE intents (mean over 12 languages, 150 seeded stratified test rows
66
  each)
67
+ type: AmazonScience/massive
68
  metrics:
69
  - type: accuracy
70
+ value: 0.7995
71
  - task:
72
  type: text-classification
73
  dataset:
74
  name: AG News (zero-shot (never trained on), first 2,000 test rows)
75
+ type: fancyzhx/ag_news
76
  metrics:
77
  - type: accuracy
78
+ value: 0.9295
79
  - task:
80
  type: text-classification
81
  dataset:
82
  name: DAIR Emotion (zero-shot, first 2,000 test rows)
83
+ type: dair-ai/emotion
84
  metrics:
85
  - type: accuracy
86
+ value: 0.504
87
  - task:
88
  type: text-classification
89
  dataset:
90
  name: HWU64 intents (English, 150 rows; rows overlapping MASSIVE removed)
91
+ type: hwu64
92
  metrics:
93
  - type: accuracy
94
+ value: 0.7867
95
  - task:
96
  type: text-classification
97
  dataset:
98
  name: SIB-200 topics (zero-shot, mean over 4 languages, 150 rows each)
99
+ type: Davlan/sib200
100
  metrics:
101
  - type: accuracy
102
+ value: 0.7817
103
  - task:
104
  type: text-classification
105
  dataset:
106
  name: Sentiment (zero-shot, mean over 12 languages, 150 rows each)
107
+ type: tyqiangz/multilingual-sentiments
108
  metrics:
109
  - type: accuracy
110
+ value: 0.605
111
  - task:
112
  type: text-classification
113
  dataset:
114
  name: HateCheck (zero-shot, mean over 11 languages, 150 rows each)
115
+ type: mteb/multi-hatecheck
116
  metrics:
117
  - type: accuracy
118
+ value: 0.6358
119
  - task:
120
  type: text-classification
121
  dataset:
122
  name: Belebele reading (zero-shot, mean over 4 languages, 150 rows each)
123
+ type: facebook/belebele
124
  metrics:
125
  - type: accuracy
126
+ value: 0.2583
127
  ---
128
 
129
  # Statim Decide Multilingual Base
130
 
131
  A decision model for [Statim](https://github.com/BEKO2210/statim), the native C++ engine for typed decisions: ask any text a
132
  **choice**, a **score** or a **yes/no** question and get calibrated answers from one forward pass, on
133
+ CPU or GPU, without Python at runtime. Version **0.7.0**, fine-tuned from
134
  [`convaiinnovations/laya-multilingual`](https://huggingface.co/convaiinnovations/laya-multilingual) (mmBERT-base encoder).
135
 
136
  <video controls preload="none" width="100%" poster="https://beko2210.github.io/statim/images/film-16x9.webp" src="https://beko2210.github.io/statim/video/statim-flagship-60s-16x9.mp4"></video>
 
170
  ## Evaluation
171
 
172
  Measured by Statim's no-harm gate ([`tools/finetune/gate.py`](https://github.com/BEKO2210/statim/blob/main/tools/finetune/gate.py))
173
+ on held-out test data the model selection never looked at. Against the checkpoint it was trained from, on 89 held-out suites: **23 significant gains, 65 within noise, 0 regressions** (gains: more than two combined binomial standard errors; regressions: significant after Holm-Bonferroni over all suites; 1 nominal drop beyond two standard errors did not stay significant).
174
 
175
  | Suite | Role | This model | Base checkpoint | Protocol |
176
  |---|---|---|---|---|
177
+ | typed-decisions | trained | **0.7630** | 0.7585 | test split, first 2,000 decisions; its train split is replay data |
178
+ | Banking77 | trained | **0.9140** | 0.9035 | test split, first 2,000 rows, all 77 intents in one question |
179
+ | MASSIVE intents | trained | **0.7995** | 0.7717 | mean over 12 languages, 150 seeded stratified test rows each |
180
+ | AG News | held out | **0.9295** | 0.9315 | zero-shot (never trained on), first 2,000 test rows |
181
+ | DAIR Emotion | held out | **0.5040** | 0.5265 | zero-shot, first 2,000 test rows |
182
+ | HWU64 intents | held out | **0.7867** | 0.8200 | English, 150 rows; rows overlapping MASSIVE removed |
183
+ | SIB-200 topics | held out | **0.7817** | 0.7217 | zero-shot, mean over 4 languages, 150 rows each |
184
+ | Sentiment | held out | **0.6050** | 0.5933 | zero-shot, mean over 12 languages, 150 rows each |
185
+ | HateCheck | held out | **0.6358** | 0.6467 | zero-shot, mean over 11 languages, 150 rows each |
186
+ | Belebele reading | held out | **0.2583** | 0.3100 | zero-shot, mean over 4 languages, 150 rows each |
187
+
188
+ ### Decision categories
189
+
190
+ One held-out suite per decision category, built from splits of the training sources that the mixture never loads; any text that also occurs in the training mixture is dropped. 150 items per language, macro over languages.
191
+
192
+ | Category | Languages | This model | Base checkpoint |
193
+ |---|---|---|---|
194
+ | complaint | en | **0.767** | 0.607 |
195
+ | emotion | de, en, es, fr, hi, zh | **0.586** | 0.520 |
196
+ | fact check | en | **0.313** | 0.260 |
197
+ | formality | ja, tr | **0.773** | 0.503 |
198
+ | intent | en, nl, tr | **0.753** | 0.516 |
199
+ | nli | en, ja, tr | **0.747** | 0.704 |
200
+ | pii | ar, de, en, es, fr, it, ja, nl, ru, sv, zh | **0.856** | 0.595 |
201
+ | reading | en | **0.927** | 0.560 |
202
+ | safety | en | **0.727** | 0.600 |
203
+ | sentiment | en, zh | **0.800** | 0.740 |
204
+ | similarity | pt | **0.833** | 0.627 |
205
+ | stance | en | **0.893** | 0.660 |
206
+ | topic | en | **0.607** | 0.267 |
207
+ | urgency | en | **0.893** | 0.660 |
208
 
209
  Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768,
210
  laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881;
211
+ Banking77 supervised MPNet 0.941; MASSIVE XLM-R base 0.857 over 12 languages (full train set). Sources:
212
  [docs/ROADMAP.md](https://github.com/BEKO2210/statim/blob/main/docs/ROADMAP.md).
213
 
214
+ <details><summary>All 89 held-out suites</summary>
215
 
216
  | Suite | Accuracy | Rows |
217
  |---|---|---|
218
+ | `amazon_massive_intent/ar` | 0.7067 | 150 |
219
+ | `amazon_massive_intent/de` | 0.7867 | 150 |
220
+ | `amazon_massive_intent/en` | 0.8267 | 150 |
221
+ | `amazon_massive_intent/es` | 0.8133 | 150 |
222
+ | `amazon_massive_intent/fr` | 0.8200 | 150 |
223
+ | `amazon_massive_intent/hi` | 0.7867 | 150 |
224
+ | `amazon_massive_intent/it` | 0.8200 | 150 |
225
+ | `amazon_massive_intent/ja` | 0.8267 | 150 |
226
  | `amazon_massive_intent/pl` | 0.8000 | 150 |
227
+ | `amazon_massive_intent/ru` | 0.8333 | 150 |
228
+ | `amazon_massive_intent/tr` | 0.7867 | 150 |
229
+ | `amazon_massive_intent/zh-CN` | 0.7867 | 150 |
230
+ | `belebele/ar` | 0.2933 | 150 |
231
+ | `belebele/de` | 0.2133 | 150 |
232
+ | `belebele/en` | 0.2467 | 150 |
233
+ | `belebele/hi` | 0.2800 | 150 |
234
+ | `categories:complaint/en` | 0.7667 | 150 |
235
+ | `categories:emotion/de` | 0.4533 | 150 |
236
+ | `categories:emotion/en` | 0.5733 | 150 |
237
+ | `categories:emotion/es` | 0.5733 | 150 |
238
+ | `categories:emotion/fr` | 0.6600 | 150 |
239
+ | `categories:emotion/hi` | 0.7133 | 150 |
240
+ | `categories:emotion/zh` | 0.5400 | 150 |
241
+ | `categories:fact_check/en` | 0.3133 | 150 |
242
+ | `categories:formality/ja` | 0.5467 | 150 |
243
+ | `categories:formality/tr` | 1.0000 | 150 |
244
+ | `categories:intent/en` | 0.8200 | 150 |
245
+ | `categories:intent/nl` | 0.4400 | 150 |
246
+ | `categories:intent/tr` | 1.0000 | 150 |
247
+ | `categories:nli/en` | 0.8067 | 150 |
248
+ | `categories:nli/ja` | 0.6333 | 150 |
249
+ | `categories:nli/tr` | 0.8000 | 150 |
250
+ | `categories:pii/ar` | 0.8467 | 150 |
251
+ | `categories:pii/de` | 0.8400 | 150 |
252
+ | `categories:pii/en` | 0.8933 | 150 |
253
+ | `categories:pii/es` | 0.8533 | 150 |
254
+ | `categories:pii/fr` | 0.8933 | 150 |
255
+ | `categories:pii/it` | 0.8400 | 150 |
256
+ | `categories:pii/ja` | 0.8467 | 150 |
257
+ | `categories:pii/nl` | 0.7600 | 150 |
258
+ | `categories:pii/ru` | 0.9067 | 150 |
259
+ | `categories:pii/sv` | 0.8333 | 150 |
260
+ | `categories:pii/zh` | 0.9000 | 150 |
261
+ | `categories:reading/en` | 0.9267 | 150 |
262
+ | `categories:safety/en` | 0.7267 | 150 |
263
+ | `categories:sentiment/en` | 0.8000 | 150 |
264
+ | `categories:sentiment/zh` | 0.8000 | 150 |
265
+ | `categories:similarity/pt` | 0.8333 | 150 |
266
+ | `categories:stance/en` | 0.8933 | 150 |
267
+ | `categories:topic/en` | 0.6067 | 150 |
268
+ | `categories:urgency/en` | 0.8933 | 150 |
269
+ | `farstail/fa` | 0.7000 | 150 |
270
+ | `go_emotions/en` | 0.4733 | 150 |
271
+ | `hwu64/en` | 0.7867 | 150 |
272
+ | `indonli/id` | 0.6667 | 150 |
273
+ | `multi_hatecheck/ar` | 0.5733 | 150 |
274
+ | `multi_hatecheck/de` | 0.6867 | 150 |
275
+ | `multi_hatecheck/en` | 0.6133 | 150 |
276
+ | `multi_hatecheck/es` | 0.6200 | 150 |
277
+ | `multi_hatecheck/fr` | 0.6867 | 150 |
278
+ | `multi_hatecheck/hi` | 0.5733 | 150 |
279
+ | `multi_hatecheck/it` | 0.6600 | 150 |
280
+ | `multi_hatecheck/nl` | 0.6533 | 150 |
281
+ | `multi_hatecheck/pl` | 0.6333 | 150 |
282
+ | `multi_hatecheck/pt` | 0.6467 | 150 |
283
+ | `multi_hatecheck/zh` | 0.6467 | 150 |
284
+ | `multilingual_sentiments/ar` | 0.5600 | 150 |
285
+ | `multilingual_sentiments/de` | 0.5800 | 150 |
286
+ | `multilingual_sentiments/en` | 0.6933 | 150 |
287
+ | `multilingual_sentiments/es` | 0.5267 | 150 |
288
+ | `multilingual_sentiments/fr` | 0.5867 | 150 |
289
+ | `multilingual_sentiments/hi` | 0.5333 | 150 |
290
+ | `multilingual_sentiments/id` | 0.7467 | 150 |
291
+ | `multilingual_sentiments/it` | 0.5667 | 150 |
292
+ | `multilingual_sentiments/ja` | 0.6600 | 150 |
293
+ | `multilingual_sentiments/ms` | 0.5333 | 150 |
294
+ | `multilingual_sentiments/pt` | 0.6333 | 150 |
295
+ | `multilingual_sentiments/zh` | 0.6400 | 150 |
296
+ | `semrel/ar` | 0.2467 | 150 |
297
+ | `semrel/en` | 0.2000 | 150 |
298
+ | `semrel/hi` | 0.2267 | 150 |
299
+ | `sib200/ar` | 0.7867 | 150 |
300
+ | `sib200/de` | 0.8000 | 150 |
301
+ | `sib200/en` | 0.8333 | 150 |
302
  | `sib200/hi` | 0.7067 | 150 |
303
+ | `test/ag_news` | 0.9295 | 2000 |
304
+ | `test/banking77` | 0.9140 | 2000 |
305
+ | `test/emotion` | 0.5040 | 2000 |
306
+ | `test/typed_decisions` | 0.7630 | 2000 |
307
 
308
  </details>
309
 
 
311
 
312
  ## Training
313
 
314
+ Multi-task fine-tuning with `train_multitask.py --clean` from checkpoint `laya-multilingual-big1`, best
315
+ epoch `12/raw` selected on validation data only. Training data: only sources whose
316
  licence permits commercial use and imposes no ShareAlike or copyleft terms (Banking77, MASSIVE,
317
  typed-decisions replay, a licence-audited tasksource mixture, Nemotron-Safety-Guard, IndicGuard,
318
  MINDS-14, SNIPS), every one listed with its licence in
SHA256SUMS CHANGED
@@ -1,12 +1,12 @@
1
- 5c0aa16064949ae42cb1db6f563c81e14e6838483608ad1d679d418e7d864fc3 DATA_LICENSES.md
2
  1106ca74eec5e589cec780e4f19f5218ea84ceb6b0d23594d7d9b5b30a2b2322 LICENSE-MODEL.md
3
  2416135f4479640b6237a07ab323585cd7eaee91041dd4208846d22d17aceb61 NOTICE
4
- 4a28b309bfa966a9ff2fe53ed4b9e04cc9da6ef9da1e26dbca8d759c9730c41f README.md
5
  83f6916d13ef0f556ac461f28308dc2bffa7ebeadee8ec9e2db5812020ea5bb4 checkpoint/encoder/config.json
6
- e34dc66c238dc1d3606ce12c24e7deb853e9eb2319c76c99f725af274bf88f3f checkpoint/model.safetensors
7
- ae3b1cc7d74023d2d7641e7518855b17c869800bcf7552514eef5bd60f4898eb checkpoint/rl_agent_config.json
8
  609d8f4c067cd3950f88594c5a802616cea245823836ef5848ee4fc40aab5b6f checkpoint/tokenizer/tokenizer.json
9
  6c6b2d8e3c84ce0e671c129cd6b374b235d6f9863042a5836358d00a89bbb5a1 checkpoint/tokenizer/tokenizer_config.json
10
- ba88393c5f94e51c7e69d5e73ed18b583afd6d4e420247d6dfc86be3b3a817b5 evaluation/eval.json
11
- 965d77854b2a2ce8384842120042b5768cd094166f965ad79a09f4b03d81b4a6 statim-decide-multilingual-base-f32.gguf
12
- 39fa79fdc381a21cc324909b810d0cfb414f474656fa7994d561118b3fbc73fd statim-decide-multilingual-base-q8_0.gguf
 
1
+ a8e5e0bb5dfdd8d831f2d1bdfafdb1ddfecb1fc3e662ef21929ef8511389a075 DATA_LICENSES.md
2
  1106ca74eec5e589cec780e4f19f5218ea84ceb6b0d23594d7d9b5b30a2b2322 LICENSE-MODEL.md
3
  2416135f4479640b6237a07ab323585cd7eaee91041dd4208846d22d17aceb61 NOTICE
4
+ a12f4a971a7fdbfd73657f8e8a575072d3355bbd992f31aa06fedbeba9546bcd README.md
5
  83f6916d13ef0f556ac461f28308dc2bffa7ebeadee8ec9e2db5812020ea5bb4 checkpoint/encoder/config.json
6
+ c44425f14ac9d55508f73a6f371e4e2802ed59e646287c3abbec10b653a19840 checkpoint/model.safetensors
7
+ 9b889caa1ae60f0d6e343a04799f88985d705799e5dfa7e05accbc639c131841 checkpoint/rl_agent_config.json
8
  609d8f4c067cd3950f88594c5a802616cea245823836ef5848ee4fc40aab5b6f checkpoint/tokenizer/tokenizer.json
9
  6c6b2d8e3c84ce0e671c129cd6b374b235d6f9863042a5836358d00a89bbb5a1 checkpoint/tokenizer/tokenizer_config.json
10
+ d9cd02f468711f0fe0b3d7e528e095fe954bbbd6ea18b68764d96b3da84f6a58 evaluation/eval.json
11
+ 03349c619a05c78d94217c5c8279317a8ea9abc4a696d90480a370ea3a402eeb statim-decide-multilingual-base-f32.gguf
12
+ 96fb3971656ee6b596cb0c108aff4bbe1306d87707ffe9ae5f364a52b919ff1b statim-decide-multilingual-base-q8_0.gguf
checkpoint/model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e34dc66c238dc1d3606ce12c24e7deb853e9eb2319c76c99f725af274bf88f3f
3
  size 643835524
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c44425f14ac9d55508f73a6f371e4e2802ed59e646287c3abbec10b653a19840
3
  size 643835524
checkpoint/rl_agent_config.json CHANGED
@@ -9,20 +9,19 @@
9
  },
10
  "cost_wrong_act": 3.0,
11
  "amp_dtype": "bf16",
12
- "model_name": "laya-multilingual-big1",
13
  "temperature": [
14
- 1.758754849433899,
15
- 1.2504303455352783,
16
- 1.3224705457687378
17
  ],
18
  "temperature_by_options": {
19
- "choice:11+": 1.813578724861145,
20
- "noul:2": 1.3224705457687378,
21
- "choice:3-5": 1.3400647640228271,
22
- "score:3-5": 1.3033926486968994,
23
- "choice:2": 1.6684647798538208,
24
- "choice:6-10": 1.2673249244689941,
25
- "score:2": 1.3971556425094604
26
  },
27
  "training": {
28
  "updates": 15987,
@@ -33,442 +32,304 @@
33
  },
34
  "fine_tuned": true,
35
  "training_multitask": {
36
- "base": "laya-multilingual",
37
- "best_epoch": "19/raw",
38
  "dev_before": {
39
- "banking77": 0.464,
40
- "massive": 0.3017,
41
- "mixture": 0.475
42
  },
43
  "log": [
44
  {
45
  "epoch": 1,
46
  "weights": "raw",
47
- "train_ce": 1.6029,
48
- "banking77": 0.752,
49
- "massive": 0.6267,
50
- "mixture": 0.52,
51
- "mean": 0.6329,
52
- "seconds": 1090
53
  },
54
  {
55
  "epoch": 1,
56
  "weights": "ema",
57
- "train_ce": 1.6029,
58
- "banking77": 0.614,
59
- "massive": 0.4817,
60
- "mixture": 0.488,
61
- "mean": 0.5279,
62
- "seconds": 1092
63
  },
64
  {
65
  "epoch": 2,
66
  "weights": "raw",
67
- "train_ce": 1.1416,
68
- "banking77": 0.772,
69
- "massive": 0.675,
70
- "mixture": 0.529,
71
- "mean": 0.6587,
72
- "seconds": 2165
73
  },
74
  {
75
  "epoch": 2,
76
  "weights": "ema",
77
- "train_ce": 1.1416,
78
- "banking77": 0.774,
79
- "massive": 0.64,
80
- "mixture": 0.533,
81
- "mean": 0.649,
82
- "seconds": 2166
83
  },
84
  {
85
  "epoch": 3,
86
  "weights": "raw",
87
- "train_ce": 1.031,
88
- "banking77": 0.806,
89
- "massive": 0.715,
90
- "mixture": 0.545,
91
- "mean": 0.6887,
92
- "seconds": 3241
93
  },
94
  {
95
  "epoch": 3,
96
  "weights": "ema",
97
- "train_ce": 1.031,
98
- "banking77": 0.818,
99
- "massive": 0.6917,
100
- "mixture": 0.537,
101
- "mean": 0.6822,
102
- "seconds": 3242
103
  },
104
  {
105
  "epoch": 4,
106
  "weights": "raw",
107
- "train_ce": 0.952,
108
- "banking77": 0.816,
109
- "massive": 0.7367,
110
- "mixture": 0.536,
111
- "mean": 0.6962,
112
- "seconds": 4324
113
  },
114
  {
115
  "epoch": 4,
116
  "weights": "ema",
117
- "train_ce": 0.952,
118
- "banking77": 0.854,
119
- "massive": 0.74,
120
- "mixture": 0.551,
121
- "mean": 0.715,
122
- "seconds": 4326
123
  },
124
  {
125
  "epoch": 5,
126
  "weights": "raw",
127
- "train_ce": 0.9832,
128
- "banking77": 0.836,
129
- "massive": 0.74,
130
- "mixture": 0.573,
131
- "mean": 0.7163,
132
- "seconds": 5404
133
  },
134
  {
135
  "epoch": 5,
136
  "weights": "ema",
137
- "train_ce": 0.9832,
138
- "banking77": 0.848,
139
- "massive": 0.7517,
140
- "mixture": 0.545,
141
- "mean": 0.7149,
142
- "seconds": 5406
143
  },
144
  {
145
  "epoch": 6,
146
  "weights": "raw",
147
- "train_ce": 0.8739,
148
- "banking77": 0.856,
149
- "massive": 0.7517,
150
- "mixture": 0.561,
151
- "mean": 0.7229,
152
- "seconds": 6476
153
  },
154
  {
155
  "epoch": 6,
156
  "weights": "ema",
157
- "train_ce": 0.8739,
158
- "banking77": 0.874,
159
- "massive": 0.7633,
160
- "mixture": 0.571,
161
- "mean": 0.7361,
162
- "seconds": 6477
163
  },
164
  {
165
  "epoch": 7,
166
  "weights": "raw",
167
- "train_ce": 0.9355,
168
- "banking77": 0.834,
169
- "massive": 0.7633,
170
- "mixture": 0.564,
171
- "mean": 0.7204,
172
- "seconds": 7550
173
  },
174
  {
175
  "epoch": 7,
176
  "weights": "ema",
177
- "train_ce": 0.9355,
178
- "banking77": 0.872,
179
- "massive": 0.7667,
180
- "mixture": 0.568,
181
- "mean": 0.7356,
182
- "seconds": 7550
183
  },
184
  {
185
  "epoch": 8,
186
  "weights": "raw",
187
- "train_ce": 0.8557,
188
- "banking77": 0.866,
189
- "massive": 0.765,
190
- "mixture": 0.573,
191
- "mean": 0.7347,
192
- "seconds": 8627
193
  },
194
  {
195
  "epoch": 8,
196
  "weights": "ema",
197
- "train_ce": 0.8557,
198
- "banking77": 0.886,
199
- "massive": 0.77,
200
- "mixture": 0.575,
201
- "mean": 0.7437,
202
- "seconds": 8627
203
  },
204
  {
205
  "epoch": 9,
206
  "weights": "raw",
207
- "train_ce": 0.8353,
208
- "banking77": 0.882,
209
- "massive": 0.7633,
210
- "mixture": 0.564,
211
- "mean": 0.7364,
212
- "seconds": 9700
213
  },
214
  {
215
  "epoch": 9,
216
  "weights": "ema",
217
- "train_ce": 0.8353,
218
- "banking77": 0.9,
219
- "massive": 0.7883,
220
- "mixture": 0.576,
221
- "mean": 0.7548,
222
- "seconds": 9700
223
  },
224
  {
225
  "epoch": 10,
226
  "weights": "raw",
227
- "train_ce": 0.7674,
228
- "banking77": 0.868,
229
- "massive": 0.7783,
230
- "mixture": 0.569,
231
- "mean": 0.7384,
232
- "seconds": 10771
233
  },
234
  {
235
  "epoch": 10,
236
  "weights": "ema",
237
- "train_ce": 0.7674,
238
- "banking77": 0.902,
239
- "massive": 0.7867,
240
- "mixture": 0.576,
241
- "mean": 0.7549,
242
- "seconds": 10771
243
- },
244
- {
245
- "epoch": 11,
246
- "weights": "raw",
247
- "train_ce": 0.7038,
248
  "banking77": 0.892,
249
- "massive": 0.7817,
250
- "mixture": 0.578,
251
- "mean": 0.7506,
252
- "seconds": 11841
253
  },
254
  {
255
  "epoch": 11,
256
- "weights": "ema",
257
- "train_ce": 0.7038,
258
- "banking77": 0.896,
259
- "massive": 0.7917,
260
- "mixture": 0.566,
261
- "mean": 0.7512,
262
- "seconds": 11841
263
- },
264
- {
265
- "epoch": 12,
266
  "weights": "raw",
267
- "train_ce": 0.6528,
268
- "banking77": 0.878,
269
- "massive": 0.7883,
270
- "mixture": 0.588,
271
- "mean": 0.7514,
272
- "seconds": 12915
273
- },
274
- {
275
- "epoch": 12,
276
- "weights": "ema",
277
- "train_ce": 0.6528,
278
  "banking77": 0.894,
279
- "massive": 0.795,
280
- "mixture": 0.572,
281
- "mean": 0.7537,
282
- "seconds": 12915
283
  },
284
  {
285
- "epoch": 13,
286
- "weights": "raw",
287
- "train_ce": 0.6207,
288
- "banking77": 0.882,
289
- "massive": 0.8083,
290
- "mixture": 0.582,
291
- "mean": 0.7574,
292
- "seconds": 13985
293
- },
294
- {
295
- "epoch": 13,
296
  "weights": "ema",
297
- "train_ce": 0.6207,
298
- "banking77": 0.888,
299
- "massive": 0.795,
300
- "mixture": 0.574,
301
- "mean": 0.7523,
302
- "seconds": 13987
303
- },
304
- {
305
- "epoch": 14,
306
- "weights": "raw",
307
- "train_ce": 0.5874,
308
  "banking77": 0.894,
309
- "massive": 0.7917,
310
- "mixture": 0.585,
311
- "mean": 0.7569,
312
- "seconds": 15061
313
- },
314
- {
315
- "epoch": 14,
316
- "weights": "ema",
317
- "train_ce": 0.5874,
318
- "banking77": 0.892,
319
- "massive": 0.7983,
320
- "mixture": 0.572,
321
- "mean": 0.7541,
322
- "seconds": 15061
323
  },
324
  {
325
- "epoch": 15,
326
- "weights": "raw",
327
- "train_ce": 0.5589,
328
- "banking77": 0.9,
329
- "massive": 0.8017,
330
- "mixture": 0.58,
331
- "mean": 0.7606,
332
- "seconds": 16133
333
- },
334
- {
335
- "epoch": 15,
336
- "weights": "ema",
337
- "train_ce": 0.5589,
338
- "banking77": 0.898,
339
- "massive": 0.81,
340
- "mixture": 0.582,
341
- "mean": 0.7633,
342
- "seconds": 16135
343
- },
344
- {
345
- "epoch": 16,
346
- "weights": "raw",
347
- "train_ce": 0.536,
348
- "banking77": 0.898,
349
- "massive": 0.795,
350
- "mixture": 0.589,
351
- "mean": 0.7607,
352
- "seconds": 17201
353
- },
354
- {
355
- "epoch": 16,
356
- "weights": "ema",
357
- "train_ce": 0.536,
358
- "banking77": 0.902,
359
- "massive": 0.81,
360
- "mixture": 0.577,
361
- "mean": 0.763,
362
- "seconds": 17201
363
- },
364
- {
365
- "epoch": 17,
366
- "weights": "raw",
367
- "train_ce": 0.5204,
368
- "banking77": 0.9,
369
- "massive": 0.8083,
370
- "mixture": 0.589,
371
- "mean": 0.7658,
372
- "seconds": 18271
373
- },
374
- {
375
- "epoch": 17,
376
- "weights": "ema",
377
- "train_ce": 0.5204,
378
- "banking77": 0.902,
379
- "massive": 0.8083,
380
- "mixture": 0.59,
381
- "mean": 0.7668,
382
- "seconds": 18273
383
- },
384
- {
385
- "epoch": 18,
386
- "weights": "raw",
387
- "train_ce": 0.514,
388
- "banking77": 0.902,
389
- "massive": 0.805,
390
- "mixture": 0.592,
391
- "mean": 0.7663,
392
- "seconds": 19354
393
- },
394
- {
395
- "epoch": 18,
396
- "weights": "ema",
397
- "train_ce": 0.514,
398
- "banking77": 0.9,
399
- "massive": 0.805,
400
- "mixture": 0.592,
401
- "mean": 0.7657,
402
- "seconds": 19354
403
- },
404
- {
405
- "epoch": 19,
406
- "weights": "raw",
407
- "train_ce": 0.4997,
408
- "banking77": 0.898,
409
- "massive": 0.8133,
410
- "mixture": 0.599,
411
- "mean": 0.7701,
412
- "seconds": 20426
413
- },
414
- {
415
- "epoch": 19,
416
- "weights": "ema",
417
- "train_ce": 0.4997,
418
- "banking77": 0.902,
419
- "massive": 0.81,
420
- "mixture": 0.593,
421
- "mean": 0.7683,
422
- "seconds": 20428
423
- },
424
- {
425
- "epoch": 20,
426
  "weights": "raw",
427
- "train_ce": 0.4975,
428
  "banking77": 0.896,
429
- "massive": 0.81,
430
- "mixture": 0.594,
431
- "mean": 0.7667,
432
- "seconds": 21505
433
  },
434
  {
435
- "epoch": 20,
436
  "weights": "ema",
437
- "train_ce": 0.4975,
438
- "banking77": 0.898,
439
- "massive": 0.81,
440
- "mixture": 0.596,
441
- "mean": 0.768,
442
- "seconds": 21505
443
  }
444
  ],
445
  "counts": {
446
  "banking77": 9493,
447
- "massive": 102000,
448
  "typed": 5400,
449
- "mixture": 187561,
450
- "distill": 6000
 
 
 
 
 
 
 
 
 
 
451
  },
452
- "distill": 6000,
453
- "massive_per_lang": 2000,
 
454
  "sentiment_per_lang": 1200,
455
  "lr": [
456
  2.5e-05,
457
  0.0001
458
  ],
459
- "epochs": 20,
460
  "seed": 20260926,
461
  "budget": {
462
  "banking77": 12000,
463
  "massive": 16000,
464
- "mixture": 20000,
465
  "typed": 4000,
466
- "distill": 3000
 
 
 
 
 
 
 
 
 
 
467
  },
468
  "warmup": 0.06,
469
  "ema": 0.999,
470
  "patience": 3,
471
- "mixture": "mixture-v5.jsonl.gz",
 
472
  "clean": true
473
  }
474
  }
 
9
  },
10
  "cost_wrong_act": 3.0,
11
  "amp_dtype": "bf16",
12
+ "model_name": "laya-multilingual-multitask",
13
  "temperature": [
14
+ 1.3157309293746948,
15
+ 1.0678483247756958,
16
+ 1.254639983177185
17
  ],
18
  "temperature_by_options": {
19
+ "choice:11+": 1.4175829887390137,
20
+ "choice:3-5": 1.1190439462661743,
21
+ "score:3-5": 1.0678483247756958,
22
+ "noul:2": 1.254639983177185,
23
+ "choice:6-10": 1.068841576576233,
24
+ "choice:2": 1.3097771406173706
 
25
  },
26
  "training": {
27
  "updates": 15987,
 
32
  },
33
  "fine_tuned": true,
34
  "training_multitask": {
35
+ "base": "laya-multilingual-big1",
36
+ "best_epoch": "12/raw",
37
  "dev_before": {
38
+ "banking77": 0.898,
39
+ "massive": 0.8467,
40
+ "mixture": 0.4875
41
  },
42
  "log": [
43
  {
44
  "epoch": 1,
45
  "weights": "raw",
46
+ "train_ce": 0.9343,
47
+ "banking77": 0.858,
48
+ "massive": 0.825,
49
+ "mixture": 0.5793,
50
+ "mean": 0.7541,
51
+ "seconds": 1676
52
  },
53
  {
54
  "epoch": 1,
55
  "weights": "ema",
56
+ "train_ce": 0.9343,
57
+ "banking77": 0.904,
58
+ "massive": 0.8417,
59
+ "mixture": 0.5545,
60
+ "mean": 0.7667,
61
+ "seconds": 1687
62
  },
63
  {
64
  "epoch": 2,
65
  "weights": "raw",
66
+ "train_ce": 0.9273,
67
+ "banking77": 0.882,
68
+ "massive": 0.8217,
69
+ "mixture": 0.6067,
70
+ "mean": 0.7701,
71
+ "seconds": 3372
72
  },
73
  {
74
  "epoch": 2,
75
  "weights": "ema",
76
+ "train_ce": 0.9273,
77
+ "banking77": 0.904,
78
+ "massive": 0.84,
79
+ "mixture": 0.6015,
80
+ "mean": 0.7818,
81
+ "seconds": 3374
82
  },
83
  {
84
  "epoch": 3,
85
  "weights": "raw",
86
+ "train_ce": 0.8975,
87
+ "banking77": 0.882,
88
+ "massive": 0.825,
89
+ "mixture": 0.6352,
90
+ "mean": 0.7807,
91
+ "seconds": 5049
92
  },
93
  {
94
  "epoch": 3,
95
  "weights": "ema",
96
+ "train_ce": 0.8975,
97
+ "banking77": 0.902,
98
+ "massive": 0.8567,
99
+ "mixture": 0.6312,
100
+ "mean": 0.7966,
101
+ "seconds": 5049
102
  },
103
  {
104
  "epoch": 4,
105
  "weights": "raw",
106
+ "train_ce": 0.8468,
107
+ "banking77": 0.894,
108
+ "massive": 0.8333,
109
+ "mixture": 0.6432,
110
+ "mean": 0.7902,
111
+ "seconds": 6725
112
  },
113
  {
114
  "epoch": 4,
115
  "weights": "ema",
116
+ "train_ce": 0.8468,
117
+ "banking77": 0.902,
118
+ "massive": 0.8433,
119
+ "mixture": 0.6495,
120
+ "mean": 0.7983,
121
+ "seconds": 6725
122
  },
123
  {
124
  "epoch": 5,
125
  "weights": "raw",
126
+ "train_ce": 0.8105,
127
+ "banking77": 0.89,
128
+ "massive": 0.8367,
129
+ "mixture": 0.6505,
130
+ "mean": 0.7924,
131
+ "seconds": 8391
132
  },
133
  {
134
  "epoch": 5,
135
  "weights": "ema",
136
+ "train_ce": 0.8105,
137
+ "banking77": 0.898,
138
+ "massive": 0.85,
139
+ "mixture": 0.6588,
140
+ "mean": 0.8023,
141
+ "seconds": 8391
142
  },
143
  {
144
  "epoch": 6,
145
  "weights": "raw",
146
+ "train_ce": 0.7723,
147
+ "banking77": 0.894,
148
+ "massive": 0.845,
149
+ "mixture": 0.6663,
150
+ "mean": 0.8018,
151
+ "seconds": 10062
152
  },
153
  {
154
  "epoch": 6,
155
  "weights": "ema",
156
+ "train_ce": 0.7723,
157
+ "banking77": 0.894,
158
+ "massive": 0.845,
159
+ "mixture": 0.6657,
160
+ "mean": 0.8016,
161
+ "seconds": 10062
162
  },
163
  {
164
  "epoch": 7,
165
  "weights": "raw",
166
+ "train_ce": 0.7471,
167
+ "banking77": 0.904,
168
+ "massive": 0.8333,
169
+ "mixture": 0.6703,
170
+ "mean": 0.8025,
171
+ "seconds": 11743
172
  },
173
  {
174
  "epoch": 7,
175
  "weights": "ema",
176
+ "train_ce": 0.7471,
177
+ "banking77": 0.894,
178
+ "massive": 0.8433,
179
+ "mixture": 0.6698,
180
+ "mean": 0.8024,
181
+ "seconds": 11745
182
  },
183
  {
184
  "epoch": 8,
185
  "weights": "raw",
186
+ "train_ce": 0.7179,
187
+ "banking77": 0.89,
188
+ "massive": 0.84,
189
+ "mixture": 0.6815,
190
+ "mean": 0.8038,
191
+ "seconds": 13428
192
  },
193
  {
194
  "epoch": 8,
195
  "weights": "ema",
196
+ "train_ce": 0.7179,
197
+ "banking77": 0.898,
198
+ "massive": 0.84,
199
+ "mixture": 0.6773,
200
+ "mean": 0.8051,
201
+ "seconds": 13431
202
  },
203
  {
204
  "epoch": 9,
205
  "weights": "raw",
206
+ "train_ce": 0.6946,
207
+ "banking77": 0.894,
208
+ "massive": 0.835,
209
+ "mixture": 0.6877,
210
+ "mean": 0.8056,
211
+ "seconds": 15113
212
  },
213
  {
214
  "epoch": 9,
215
  "weights": "ema",
216
+ "train_ce": 0.6946,
217
+ "banking77": 0.89,
218
+ "massive": 0.84,
219
+ "mixture": 0.6862,
220
+ "mean": 0.8054,
221
+ "seconds": 15114
222
  },
223
  {
224
  "epoch": 10,
225
  "weights": "raw",
226
+ "train_ce": 0.6722,
227
+ "banking77": 0.894,
228
+ "massive": 0.8383,
229
+ "mixture": 0.6893,
230
+ "mean": 0.8072,
231
+ "seconds": 16795
232
  },
233
  {
234
  "epoch": 10,
235
  "weights": "ema",
236
+ "train_ce": 0.6722,
 
 
 
 
 
 
 
 
 
 
237
  "banking77": 0.892,
238
+ "massive": 0.84,
239
+ "mixture": 0.6905,
240
+ "mean": 0.8075,
241
+ "seconds": 16796
242
  },
243
  {
244
  "epoch": 11,
 
 
 
 
 
 
 
 
 
 
245
  "weights": "raw",
246
+ "train_ce": 0.6562,
 
 
 
 
 
 
 
 
 
 
247
  "banking77": 0.894,
248
+ "massive": 0.8467,
249
+ "mixture": 0.6938,
250
+ "mean": 0.8115,
251
+ "seconds": 18472
252
  },
253
  {
254
+ "epoch": 11,
 
 
 
 
 
 
 
 
 
 
255
  "weights": "ema",
256
+ "train_ce": 0.6562,
 
 
 
 
 
 
 
 
 
 
257
  "banking77": 0.894,
258
+ "massive": 0.8433,
259
+ "mixture": 0.6932,
260
+ "mean": 0.8102,
261
+ "seconds": 18473
 
 
 
 
 
 
 
 
 
 
262
  },
263
  {
264
+ "epoch": 12,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
265
  "weights": "raw",
266
+ "train_ce": 0.6488,
267
  "banking77": 0.896,
268
+ "massive": 0.8483,
269
+ "mixture": 0.6968,
270
+ "mean": 0.8137,
271
+ "seconds": 20158
272
  },
273
  {
274
+ "epoch": 12,
275
  "weights": "ema",
276
+ "train_ce": 0.6488,
277
+ "banking77": 0.896,
278
+ "massive": 0.8467,
279
+ "mixture": 0.696,
280
+ "mean": 0.8129,
281
+ "seconds": 20159
282
  }
283
  ],
284
  "counts": {
285
  "banking77": 9493,
286
+ "massive": 560371,
287
  "typed": 5400,
288
+ "mixture:safety": 77867,
289
+ "mixture:intent": 53461,
290
+ "mixture": 191609,
291
+ "mixture:complaint": 55918,
292
+ "mixture:emotion": 41770,
293
+ "mixture:similarity": 50686,
294
+ "mixture:sentiment": 46531,
295
+ "mixture:other": 97480,
296
+ "mixture:topic": 38020,
297
+ "mixture:reading": 18443,
298
+ "mixture:nli": 46891,
299
+ "distill": 24000
300
  },
301
+ "distill": 24000,
302
+ "massive_per_lang": 11000,
303
+ "massive_langs": null,
304
  "sentiment_per_lang": 1200,
305
  "lr": [
306
  2.5e-05,
307
  0.0001
308
  ],
309
+ "epochs": 12,
310
  "seed": 20260926,
311
  "budget": {
312
  "banking77": 12000,
313
  "massive": 16000,
314
+ "mixture": 24000,
315
  "typed": 4000,
316
+ "distill": 9000,
317
+ "mixture:safety": 2975,
318
+ "mixture:intent": 2465,
319
+ "mixture:complaint": 2521,
320
+ "mixture:emotion": 2179,
321
+ "mixture:similarity": 2400,
322
+ "mixture:sentiment": 2299,
323
+ "mixture:other": 3328,
324
+ "mixture:topic": 2078,
325
+ "mixture:reading": 1448,
326
+ "mixture:nli": 2308
327
  },
328
  "warmup": 0.06,
329
  "ema": 0.999,
330
  "patience": 3,
331
+ "optim": "adamw",
332
+ "mixture": "mixture-v8.jsonl.gz",
333
  "clean": true
334
  }
335
  }
evaluation/eval.json CHANGED
@@ -1,282 +1,509 @@
1
  {
2
- "model": "models/laya-multilingual-big1",
3
  "validation": {
4
  "banking77": 0.898,
5
- "massive": 0.81,
6
- "sentiment": 0.58,
7
- "emotion": 0.49,
8
- "ag_news": 0.924
9
  },
10
  "heldout": {
11
  "test/ag_news": {
12
- "acc": 0.9315,
13
  "n": 2000,
14
- "ece": 0.0088
15
  },
16
  "test/emotion": {
17
- "acc": 0.5265,
18
  "n": 2000,
19
- "ece": 0.2127
20
  },
21
  "test/banking77": {
22
- "acc": 0.9035,
23
  "n": 2000,
24
- "ece": 0.085
25
  },
26
  "test/typed_decisions": {
27
- "acc": 0.7585,
28
  "n": 2000,
29
  "ece": null
30
  },
31
  "amazon_massive_intent/de": {
32
- "acc": 0.7667,
33
  "n": 150,
34
- "ece": 0.1134
35
  },
36
  "amazon_massive_intent/en": {
37
- "acc": 0.8133,
38
  "n": 150,
39
- "ece": 0.0751
40
  },
41
  "amazon_massive_intent/fr": {
42
- "acc": 0.76,
43
  "n": 150,
44
- "ece": 0.0894
45
  },
46
  "amazon_massive_intent/es": {
47
- "acc": 0.7733,
48
  "n": 150,
49
- "ece": 0.1043
50
  },
51
  "amazon_massive_intent/it": {
52
- "acc": 0.7667,
53
  "n": 150,
54
- "ece": 0.1365
55
  },
56
  "amazon_massive_intent/tr": {
57
- "acc": 0.76,
58
  "n": 150,
59
- "ece": 0.1207
60
  },
61
  "amazon_massive_intent/pl": {
62
  "acc": 0.8,
63
  "n": 150,
64
- "ece": 0.106
65
  },
66
  "amazon_massive_intent/ru": {
67
- "acc": 0.82,
68
  "n": 150,
69
- "ece": 0.119
70
  },
71
  "amazon_massive_intent/ja": {
72
- "acc": 0.78,
73
  "n": 150,
74
- "ece": 0.0907
75
  },
76
  "amazon_massive_intent/zh-CN": {
77
- "acc": 0.8,
78
  "n": 150,
79
- "ece": 0.1124
80
  },
81
  "amazon_massive_intent/ar": {
82
- "acc": 0.66,
83
  "n": 150,
84
- "ece": 0.1182
85
  },
86
  "amazon_massive_intent/hi": {
87
- "acc": 0.76,
88
  "n": 150,
89
- "ece": 0.1192
90
  },
91
  "multilingual_sentiments/ar": {
92
- "acc": 0.5533,
93
  "n": 150,
94
- "ece": 0.1422
95
  },
96
  "multilingual_sentiments/zh": {
97
- "acc": 0.6333,
98
  "n": 150,
99
- "ece": 0.1794
100
  },
101
  "multilingual_sentiments/en": {
102
- "acc": 0.64,
103
  "n": 150,
104
- "ece": 0.1111
105
  },
106
  "multilingual_sentiments/fr": {
107
- "acc": 0.5533,
108
  "n": 150,
109
- "ece": 0.1315
110
  },
111
  "multilingual_sentiments/de": {
112
- "acc": 0.56,
113
  "n": 150,
114
- "ece": 0.1269
115
  },
116
  "multilingual_sentiments/hi": {
117
- "acc": 0.46,
118
  "n": 150,
119
- "ece": 0.2051
120
  },
121
  "multilingual_sentiments/id": {
122
- "acc": 0.7133,
123
  "n": 150,
124
- "ece": 0.1191
125
  },
126
  "multilingual_sentiments/it": {
127
- "acc": 0.6,
128
  "n": 150,
129
- "ece": 0.1432
130
  },
131
  "multilingual_sentiments/ja": {
132
- "acc": 0.6733,
133
  "n": 150,
134
- "ece": 0.1392
135
  },
136
  "multilingual_sentiments/ms": {
137
- "acc": 0.5267,
138
  "n": 150,
139
- "ece": 0.189
140
  },
141
  "multilingual_sentiments/pt": {
142
- "acc": 0.66,
143
  "n": 150,
144
- "ece": 0.0952
145
  },
146
  "multilingual_sentiments/es": {
147
- "acc": 0.5467,
148
  "n": 150,
149
  "ece": 0.2224
150
  },
151
  "go_emotions/en": {
152
- "acc": 0.4667,
153
  "n": 150,
154
- "ece": 0.1269
155
  },
156
  "multi_hatecheck/en": {
157
- "acc": 0.6333,
158
  "n": 150,
159
- "ece": 0.0572
160
  },
161
  "multi_hatecheck/de": {
162
- "acc": 0.6667,
163
  "n": 150,
164
- "ece": 0.1041
165
  },
166
  "multi_hatecheck/fr": {
167
- "acc": 0.6667,
168
  "n": 150,
169
- "ece": 0.0257
170
  },
171
  "multi_hatecheck/es": {
172
- "acc": 0.6733,
173
  "n": 150,
174
- "ece": 0.038
175
  },
176
  "multi_hatecheck/it": {
177
- "acc": 0.6867,
178
  "n": 150,
179
- "ece": 0.0648
180
  },
181
  "multi_hatecheck/nl": {
182
- "acc": 0.6,
183
  "n": 150,
184
- "ece": 0.1074
185
  },
186
  "multi_hatecheck/pl": {
187
- "acc": 0.64,
188
  "n": 150,
189
- "ece": 0.0847
190
  },
191
  "multi_hatecheck/pt": {
192
- "acc": 0.7067,
193
  "n": 150,
194
- "ece": 0.0865
195
  },
196
  "multi_hatecheck/zh": {
197
- "acc": 0.7,
198
  "n": 150,
199
- "ece": 0.0713
200
  },
201
  "multi_hatecheck/ar": {
202
- "acc": 0.58,
203
  "n": 150,
204
- "ece": 0.085
205
  },
206
  "multi_hatecheck/hi": {
207
- "acc": 0.56,
208
  "n": 150,
209
- "ece": 0.0877
210
  },
211
  "sib200/en": {
212
- "acc": 0.7467,
213
  "n": 150,
214
- "ece": 0.0713
215
  },
216
  "sib200/de": {
217
- "acc": 0.7267,
218
  "n": 150,
219
- "ece": 0.0726
220
  },
221
  "sib200/ar": {
222
- "acc": 0.7067,
223
  "n": 150,
224
- "ece": 0.0872
225
  },
226
  "sib200/hi": {
227
  "acc": 0.7067,
228
  "n": 150,
229
- "ece": 0.0964
230
  },
231
  "hwu64/en": {
232
- "acc": 0.82,
233
  "n": 150,
234
- "ece": 0.0925
235
  },
236
  "indonli/id": {
237
- "acc": 0.7,
238
  "n": 150,
239
- "ece": 0.1418
240
  },
241
  "farstail/fa": {
242
- "acc": 0.6333,
243
  "n": 150,
244
- "ece": 0.1371
245
  },
246
  "belebele/en": {
247
- "acc": 0.3467,
248
  "n": 150,
249
- "ece": 0.1656
250
  },
251
  "belebele/de": {
252
- "acc": 0.34,
253
  "n": 150,
254
- "ece": 0.1268
255
  },
256
  "belebele/ar": {
257
- "acc": 0.2467,
258
  "n": 150,
259
- "ece": 0.1813
260
  },
261
  "belebele/hi": {
262
- "acc": 0.3067,
263
  "n": 150,
264
- "ece": 0.1124
265
  },
266
  "semrel/en": {
267
- "acc": 0.22,
268
  "n": 150,
269
- "ece": 0.1619
270
  },
271
  "semrel/ar": {
272
- "acc": 0.24,
273
  "n": 150,
274
- "ece": 0.1085
275
  },
276
  "semrel/hi": {
277
- "acc": 0.22,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
278
  "n": 150,
279
- "ece": 0.1612
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
280
  }
281
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
282
  }
 
1
  {
2
+ "model": "models/laya-multilingual-v9",
3
  "validation": {
4
  "banking77": 0.898,
5
+ "massive": 0.8133,
6
+ "sentiment": 0.5667,
7
+ "emotion": 0.484,
8
+ "ag_news": 0.908
9
  },
10
  "heldout": {
11
  "test/ag_news": {
12
+ "acc": 0.9295,
13
  "n": 2000,
14
+ "ece": 0.0282
15
  },
16
  "test/emotion": {
17
+ "acc": 0.504,
18
  "n": 2000,
19
+ "ece": 0.2148
20
  },
21
  "test/banking77": {
22
+ "acc": 0.914,
23
  "n": 2000,
24
+ "ece": 0.0796
25
  },
26
  "test/typed_decisions": {
27
+ "acc": 0.763,
28
  "n": 2000,
29
  "ece": null
30
  },
31
  "amazon_massive_intent/de": {
32
+ "acc": 0.7867,
33
  "n": 150,
34
+ "ece": 0.0726
35
  },
36
  "amazon_massive_intent/en": {
37
+ "acc": 0.8267,
38
  "n": 150,
39
+ "ece": 0.0544
40
  },
41
  "amazon_massive_intent/fr": {
42
+ "acc": 0.82,
43
  "n": 150,
44
+ "ece": 0.0611
45
  },
46
  "amazon_massive_intent/es": {
47
+ "acc": 0.8133,
48
  "n": 150,
49
+ "ece": 0.0746
50
  },
51
  "amazon_massive_intent/it": {
52
+ "acc": 0.82,
53
  "n": 150,
54
+ "ece": 0.0877
55
  },
56
  "amazon_massive_intent/tr": {
57
+ "acc": 0.7867,
58
  "n": 150,
59
+ "ece": 0.0614
60
  },
61
  "amazon_massive_intent/pl": {
62
  "acc": 0.8,
63
  "n": 150,
64
+ "ece": 0.0585
65
  },
66
  "amazon_massive_intent/ru": {
67
+ "acc": 0.8333,
68
  "n": 150,
69
+ "ece": 0.0718
70
  },
71
  "amazon_massive_intent/ja": {
72
+ "acc": 0.8267,
73
  "n": 150,
74
+ "ece": 0.0875
75
  },
76
  "amazon_massive_intent/zh-CN": {
77
+ "acc": 0.7867,
78
  "n": 150,
79
+ "ece": 0.0666
80
  },
81
  "amazon_massive_intent/ar": {
82
+ "acc": 0.7067,
83
  "n": 150,
84
+ "ece": 0.0606
85
  },
86
  "amazon_massive_intent/hi": {
87
+ "acc": 0.7867,
88
  "n": 150,
89
+ "ece": 0.0609
90
  },
91
  "multilingual_sentiments/ar": {
92
+ "acc": 0.56,
93
  "n": 150,
94
+ "ece": 0.1499
95
  },
96
  "multilingual_sentiments/zh": {
97
+ "acc": 0.64,
98
  "n": 150,
99
+ "ece": 0.1621
100
  },
101
  "multilingual_sentiments/en": {
102
+ "acc": 0.6933,
103
  "n": 150,
104
+ "ece": 0.0874
105
  },
106
  "multilingual_sentiments/fr": {
107
+ "acc": 0.5867,
108
  "n": 150,
109
+ "ece": 0.1192
110
  },
111
  "multilingual_sentiments/de": {
112
+ "acc": 0.58,
113
  "n": 150,
114
+ "ece": 0.1361
115
  },
116
  "multilingual_sentiments/hi": {
117
+ "acc": 0.5333,
118
  "n": 150,
119
+ "ece": 0.1885
120
  },
121
  "multilingual_sentiments/id": {
122
+ "acc": 0.7467,
123
  "n": 150,
124
+ "ece": 0.0827
125
  },
126
  "multilingual_sentiments/it": {
127
+ "acc": 0.5667,
128
  "n": 150,
129
+ "ece": 0.1605
130
  },
131
  "multilingual_sentiments/ja": {
132
+ "acc": 0.66,
133
  "n": 150,
134
+ "ece": 0.1772
135
  },
136
  "multilingual_sentiments/ms": {
137
+ "acc": 0.5333,
138
  "n": 150,
139
+ "ece": 0.197
140
  },
141
  "multilingual_sentiments/pt": {
142
+ "acc": 0.6333,
143
  "n": 150,
144
+ "ece": 0.1496
145
  },
146
  "multilingual_sentiments/es": {
147
+ "acc": 0.5267,
148
  "n": 150,
149
  "ece": 0.2224
150
  },
151
  "go_emotions/en": {
152
+ "acc": 0.4733,
153
  "n": 150,
154
+ "ece": 0.1692
155
  },
156
  "multi_hatecheck/en": {
157
+ "acc": 0.6133,
158
  "n": 150,
159
+ "ece": 0.108
160
  },
161
  "multi_hatecheck/de": {
162
+ "acc": 0.6867,
163
  "n": 150,
164
+ "ece": 0.0632
165
  },
166
  "multi_hatecheck/fr": {
167
+ "acc": 0.6867,
168
  "n": 150,
169
+ "ece": 0.1177
170
  },
171
  "multi_hatecheck/es": {
172
+ "acc": 0.62,
173
  "n": 150,
174
+ "ece": 0.0724
175
  },
176
  "multi_hatecheck/it": {
177
+ "acc": 0.66,
178
  "n": 150,
179
+ "ece": 0.0974
180
  },
181
  "multi_hatecheck/nl": {
182
+ "acc": 0.6533,
183
  "n": 150,
184
+ "ece": 0.0851
185
  },
186
  "multi_hatecheck/pl": {
187
+ "acc": 0.6333,
188
  "n": 150,
189
+ "ece": 0.0867
190
  },
191
  "multi_hatecheck/pt": {
192
+ "acc": 0.6467,
193
  "n": 150,
194
+ "ece": 0.0782
195
  },
196
  "multi_hatecheck/zh": {
197
+ "acc": 0.6467,
198
  "n": 150,
199
+ "ece": 0.0846
200
  },
201
  "multi_hatecheck/ar": {
202
+ "acc": 0.5733,
203
  "n": 150,
204
+ "ece": 0.0855
205
  },
206
  "multi_hatecheck/hi": {
207
+ "acc": 0.5733,
208
  "n": 150,
209
+ "ece": 0.1385
210
  },
211
  "sib200/en": {
212
+ "acc": 0.8333,
213
  "n": 150,
214
+ "ece": 0.0795
215
  },
216
  "sib200/de": {
217
+ "acc": 0.8,
218
  "n": 150,
219
+ "ece": 0.0728
220
  },
221
  "sib200/ar": {
222
+ "acc": 0.7867,
223
  "n": 150,
224
+ "ece": 0.1031
225
  },
226
  "sib200/hi": {
227
  "acc": 0.7067,
228
  "n": 150,
229
+ "ece": 0.1364
230
  },
231
  "hwu64/en": {
232
+ "acc": 0.7867,
233
  "n": 150,
234
+ "ece": 0.0727
235
  },
236
  "indonli/id": {
237
+ "acc": 0.6667,
238
  "n": 150,
239
+ "ece": 0.1434
240
  },
241
  "farstail/fa": {
242
+ "acc": 0.7,
243
  "n": 150,
244
+ "ece": 0.1228
245
  },
246
  "belebele/en": {
247
+ "acc": 0.2467,
248
  "n": 150,
249
+ "ece": 0.3155
250
  },
251
  "belebele/de": {
252
+ "acc": 0.2133,
253
  "n": 150,
254
+ "ece": 0.3476
255
  },
256
  "belebele/ar": {
257
+ "acc": 0.2933,
258
  "n": 150,
259
+ "ece": 0.247
260
  },
261
  "belebele/hi": {
262
+ "acc": 0.28,
263
  "n": 150,
264
+ "ece": 0.3109
265
  },
266
  "semrel/en": {
267
+ "acc": 0.2,
268
  "n": 150,
269
+ "ece": 0.2866
270
  },
271
  "semrel/ar": {
272
+ "acc": 0.2467,
273
  "n": 150,
274
+ "ece": 0.2724
275
  },
276
  "semrel/hi": {
277
+ "acc": 0.2267,
278
+ "n": 150,
279
+ "ece": 0.2519
280
+ },
281
+ "categories:sentiment/en": {
282
+ "acc": 0.8,
283
+ "n": 150,
284
+ "ece": 0.0704,
285
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
286
+ },
287
+ "categories:sentiment/zh": {
288
+ "acc": 0.8,
289
+ "n": 150,
290
+ "ece": 0.0818,
291
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
292
+ },
293
+ "categories:emotion/de": {
294
+ "acc": 0.4533,
295
+ "n": 150,
296
+ "ece": 0.0991,
297
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
298
+ },
299
+ "categories:emotion/en": {
300
+ "acc": 0.5733,
301
+ "n": 150,
302
+ "ece": 0.1161,
303
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
304
+ },
305
+ "categories:emotion/es": {
306
+ "acc": 0.5733,
307
+ "n": 150,
308
+ "ece": 0.0795,
309
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
310
+ },
311
+ "categories:emotion/fr": {
312
+ "acc": 0.66,
313
  "n": 150,
314
+ "ece": 0.0923,
315
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
316
+ },
317
+ "categories:emotion/hi": {
318
+ "acc": 0.7133,
319
+ "n": 150,
320
+ "ece": 0.1035,
321
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
322
+ },
323
+ "categories:emotion/zh": {
324
+ "acc": 0.54,
325
+ "n": 150,
326
+ "ece": 0.0978,
327
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
328
+ },
329
+ "categories:complaint/en": {
330
+ "acc": 0.7667,
331
+ "n": 150,
332
+ "ece": 0.0751,
333
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
334
+ },
335
+ "categories:nli/en": {
336
+ "acc": 0.8067,
337
+ "n": 150,
338
+ "ece": 0.0613,
339
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
340
+ },
341
+ "categories:nli/ja": {
342
+ "acc": 0.6333,
343
+ "n": 150,
344
+ "ece": 0.12,
345
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
346
+ },
347
+ "categories:nli/tr": {
348
+ "acc": 0.8,
349
+ "n": 150,
350
+ "ece": 0.0736,
351
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
352
+ },
353
+ "categories:safety/en": {
354
+ "acc": 0.7267,
355
+ "n": 150,
356
+ "ece": 0.0738,
357
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
358
+ },
359
+ "categories:reading/en": {
360
+ "acc": 0.9267,
361
+ "n": 150,
362
+ "ece": 0.0317,
363
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
364
+ },
365
+ "categories:similarity/pt": {
366
+ "acc": 0.8333,
367
+ "n": 150,
368
+ "ece": 0.0741,
369
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
370
+ },
371
+ "categories:topic/en": {
372
+ "acc": 0.6067,
373
+ "n": 150,
374
+ "ece": 0.1069,
375
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
376
+ },
377
+ "categories:intent/en": {
378
+ "acc": 0.82,
379
+ "n": 150,
380
+ "ece": 0.1027,
381
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
382
+ },
383
+ "categories:intent/nl": {
384
+ "acc": 0.44,
385
+ "n": 150,
386
+ "ece": 0.1407,
387
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
388
+ },
389
+ "categories:intent/tr": {
390
+ "acc": 1.0,
391
+ "n": 150,
392
+ "ece": 0.0014,
393
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
394
+ },
395
+ "categories:stance/en": {
396
+ "acc": 0.8933,
397
+ "n": 150,
398
+ "ece": 0.0921,
399
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
400
+ },
401
+ "categories:formality/ja": {
402
+ "acc": 0.5467,
403
+ "n": 150,
404
+ "ece": 0.0616,
405
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
406
+ },
407
+ "categories:formality/tr": {
408
+ "acc": 1.0,
409
+ "n": 150,
410
+ "ece": 0.0667,
411
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
412
+ },
413
+ "categories:urgency/en": {
414
+ "acc": 0.8933,
415
+ "n": 150,
416
+ "ece": 0.1511,
417
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
418
+ },
419
+ "categories:fact_check/en": {
420
+ "acc": 0.3133,
421
+ "n": 150,
422
+ "ece": 0.3145,
423
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
424
+ },
425
+ "categories:pii/ar": {
426
+ "acc": 0.8467,
427
+ "n": 150,
428
+ "ece": 0.0645,
429
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
430
+ },
431
+ "categories:pii/de": {
432
+ "acc": 0.84,
433
+ "n": 150,
434
+ "ece": 0.0387,
435
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
436
+ },
437
+ "categories:pii/en": {
438
+ "acc": 0.8933,
439
+ "n": 150,
440
+ "ece": 0.0967,
441
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
442
+ },
443
+ "categories:pii/es": {
444
+ "acc": 0.8533,
445
+ "n": 150,
446
+ "ece": 0.06,
447
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
448
+ },
449
+ "categories:pii/fr": {
450
+ "acc": 0.8933,
451
+ "n": 150,
452
+ "ece": 0.0387,
453
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
454
+ },
455
+ "categories:pii/it": {
456
+ "acc": 0.84,
457
+ "n": 150,
458
+ "ece": 0.0638,
459
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
460
+ },
461
+ "categories:pii/ja": {
462
+ "acc": 0.8467,
463
+ "n": 150,
464
+ "ece": 0.0413,
465
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
466
+ },
467
+ "categories:pii/nl": {
468
+ "acc": 0.76,
469
+ "n": 150,
470
+ "ece": 0.1127,
471
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
472
+ },
473
+ "categories:pii/ru": {
474
+ "acc": 0.9067,
475
+ "n": 150,
476
+ "ece": 0.0725,
477
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
478
+ },
479
+ "categories:pii/sv": {
480
+ "acc": 0.8333,
481
+ "n": 150,
482
+ "ece": 0.0326,
483
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
484
+ },
485
+ "categories:pii/zh": {
486
+ "acc": 0.9,
487
+ "n": 150,
488
+ "ece": 0.0608,
489
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6"
490
  }
491
+ },
492
+ "reported": {
493
+ "categories:emotion/pt": {
494
+ "acc": 0.5067,
495
+ "n": 150,
496
+ "ece": 0.1138,
497
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6",
498
+ "note": "no v6 training source has emotion labels in Portuguese: zero-shot"
499
+ },
500
+ "categories:emotion/ru": {
501
+ "acc": 0.6933,
502
+ "n": 150,
503
+ "ece": 0.0784,
504
+ "pool": "e0c48b24784ef0c9+b8a87e85f3bf72b6",
505
+ "note": "no v6 training source has emotion labels in Russian: zero-shot"
506
+ }
507
+ },
508
+ "mixture": "/home/belkis/Schreibtisch/My WorkSpace/statim/data/mixture-v8.jsonl.gz"
509
  }
statim-decide-multilingual-base-f32.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:965d77854b2a2ce8384842120042b5768cd094166f965ad79a09f4b03d81b4a6
3
- size 908307072
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:03349c619a05c78d94217c5c8279317a8ea9abc4a696d90480a370ea3a402eeb
3
+ size 908307040
statim-decide-multilingual-base-q8_0.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:39fa79fdc381a21cc324909b810d0cfb414f474656fa7994d561118b3fbc73fd
3
- size 356674176
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:96fb3971656ee6b596cb0c108aff4bbe1306d87707ffe9ae5f364a52b919ff1b
3
+ size 356674144