elnachto commited on
Commit
6c2dad2
·
verified ·
1 Parent(s): 194786c

Upload laya-triage model

Browse files
Files changed (3) hide show
  1. README.md +45 -30
  2. model.safetensors +1 -1
  3. rl_agent_config.json +8 -8
README.md CHANGED
@@ -32,49 +32,64 @@ metrics:
32
 
33
  Classifies GitHub issues written in any language as **bug**, **feature**, **question** or **docs**. A fine-tune of [Laya multilingual](https://huggingface.co/convaiinnovations/laya-multilingual) (mmBERT-base) used by the [laya-triage](https://github.com/elnachto/laya-triage) GitHub Action for non-English issues, next to the English model [laya-triage-en](https://huggingface.co/elnachto/laya-triage-en).
34
 
 
 
35
  ## Results
36
 
37
- The same 500 NLBSE'23 validation issues, machine-translated with NLLB-200 into 13 languages. Accuracy (±3 points per language):
38
-
39
- | Language | Laya base | Jev (TypeSafe, hosted) | **laya-triage** |
40
- |---|---|---|---|
41
- | English | 89.2% | 88.2% | **90.6%** |
42
- | German | 74.6% | 85.8% | **87.4%** |
43
- | Vietnamese | 71.2% | 84.4% | **86.6%** |
44
- | Chinese | 68.2% | 83.8% | **86.2%** |
45
- | Japanese | 67.8% | 84.4% | **86.0%** |
46
- | Turkish | 68.2% | 85.0% | **85.8%** |
47
- | Indonesian | 73.8% | 86.0% | 85.6% |
48
- | Spanish | 73.8% | 85.6% | 85.2% |
49
- | Hindi | 66.2% | 85.4% | 85.0% |
50
- | Portuguese | 74.0% | 86.6% | 84.8% |
51
- | Russian | 70.8% | 85.8% | 84.6% |
52
- | Korean | 64.2% | 83.8% | **84.4%** |
53
- | French | 71.6% | 85.2% | 84.2% |
54
- | Arabic | 67.4% | 84.8% | 84.0% |
55
-
56
- laya-triage and Jev are within noise of each other across languages; both are far ahead of the untuned base.
57
-
58
- Translations can flatter a model trained on translations, so we also checked real issues: on 367 non-English issues opened in 2026 (never seen, written by people, not translated) accuracy went from **47.1% to 65.7%** without class priors, and to 61.0% with the priors the action applies. That set is balanced on purpose, which penalizes priors; on typical repositories the priors help.
 
 
 
 
 
 
 
 
 
59
 
60
  ## How to use
61
 
62
  Use it through the Laya router together with the English model, with the exact training question. See [laya-triage-en](https://huggingface.co/elnachto/laya-triage-en#how-to-use) for the code; this checkpoint is the `"multilingual"` entry.
63
 
 
 
 
 
64
  ## Training
65
 
66
  - Base: `convaiinnovations/laya-multilingual` (mmBERT-base, 322M parameters).
67
- - Data: 2,000 NLBSE'23 training issues (500 per class) translated into 13 languages with NLLB-200 distilled 600M, plus the English originals and 10,000 extra English issues: 35,200 examples. Code blocks, stack traces and error messages were left untranslated, as in real issues.
68
- - Validation: 200 issues held out in all 14 languages (no issue appears in both splits in any language).
69
- - One epoch, same recipe as the English model; the second epoch overfit and was discarded.
70
- - Calibration: temperature 1.203 (the base model was overconfident), ECE 0.077 → 0.053.
71
 
72
  ## Limitations
73
 
74
- - Machine-translated training data: real issues mix languages, code and English error messages more than translations do.
75
- - 13 languages measured; others are supported by the base model but not evaluated.
76
- - Trained on balanced classes, so it expects class priors: multiply the probabilities by the priors stored in `rl_agent_config.json` (`laya_triage.priores`: bug 0.526, feature 0.370, question 0.060, docs 0.044) and renormalize, as the action does.
77
 
78
  ## Credits
79
 
80
- Base model by [Convai Innovations](https://huggingface.co/convaiinnovations/laya-multilingual) (Apache-2.0). Translation with [NLLB-200](https://huggingface.co/facebook/nllb-200-distilled-600M). Data from the [NLBSE'23 tool competition](https://github.com/nlbse2023/issue-report-classification). Built by [elnachto](https://github.com/elnachto).
 
32
 
33
  Classifies GitHub issues written in any language as **bug**, **feature**, **question** or **docs**. A fine-tune of [Laya multilingual](https://huggingface.co/convaiinnovations/laya-multilingual) (mmBERT-base) used by the [laya-triage](https://github.com/elnachto/laya-triage) GitHub Action for non-English issues, next to the English model [laya-triage-en](https://huggingface.co/elnachto/laya-triage-en).
34
 
35
+ This is the v1.1 model: the v1.0 multilingual model, fine-tuned on real recent issues from active repositories.
36
+
37
  ## Results
38
 
39
+ ### Real non-English issues
40
+
41
+ Issues written by people, not translated, that the router sent to this model. No repository in these sets was used for training.
42
+
43
+ | Set | v1.0 | **v1.1** |
44
+ |---|---|---|
45
+ | Recent issues from 288 active repositories (1,031 routed here), adapted to each repository | 74.0% | **77.6%** |
46
+ | Issues opened in 2026 (367 routed here), priors set to the set's balanced mix | 65.7% | **71.1%** |
47
+
48
+ ### 14 languages
49
+
50
+ The same 500 NLBSE'23 validation issues, machine-translated with NLLB-200. Accuracy of the whole system (router + both models) with the natural class mix, ±3 points per language:
51
+
52
+ | Language | Laya base | Jev (TypeSafe, hosted) | laya-triage v1.0 | **laya-triage v1.1** |
53
+ |---|---|---|---|---|
54
+ | English | 89.2% | 88.2% | 90.6% | **92.4%** |
55
+ | Vietnamese | 71.2% | 84.4% | 86.6% | **87.6%** |
56
+ | German | 74.6% | 85.8% | 87.4% | **86.4%** |
57
+ | Turkish | 68.2% | 85.0% | 85.8% | **86.4%** |
58
+ | Portuguese | 74.0% | 86.6% | 84.8% | 86.4% |
59
+ | Indonesian | 73.8% | 86.0% | 85.6% | **86.2%** |
60
+ | Spanish | 73.8% | 85.6% | 85.2% | **86.2%** |
61
+ | Russian | 70.8% | 85.8% | 84.6% | 85.6% |
62
+ | French | 71.6% | 85.2% | 84.2% | **85.4%** |
63
+ | Hindi | 66.2% | 85.4% | 85.0% | 85.2% |
64
+ | Arabic | 67.4% | 84.8% | 84.0% | **85.0%** |
65
+ | Chinese | 68.2% | 83.8% | 86.2% | **84.8%** |
66
+ | Japanese | 67.8% | 84.4% | 86.0% | **84.8%** |
67
+ | Korean | 64.2% | 83.8% | 84.4% | **84.8%** |
68
+
69
+ v1.1 is ahead of Jev in 11 of 14 languages (in bold); most gaps between the two are within noise. Average over the 13 non-English languages: 85.8% for v1.1, 85.4% for v1.0 and 85.1% for Jev.
70
 
71
  ## How to use
72
 
73
  Use it through the Laya router together with the English model, with the exact training question. See [laya-triage-en](https://huggingface.co/elnachto/laya-triage-en#how-to-use) for the code; this checkpoint is the `"multilingual"` entry.
74
 
75
+ ### Class balance
76
+
77
+ v1.1 was trained on bug 53.2%, feature 30.8%, question 11.0% and docs 5.0%. The mix is stored in `rl_agent_config.json` under `laya_triage.mezcla_base`. To match a repository's own mix, multiply the probabilities by your class frequencies divided by these and renormalize. The action does this automatically with `class-priors`.
78
+
79
  ## Training
80
 
81
  - Base: `convaiinnovations/laya-multilingual` (mmBERT-base, 322M parameters).
82
+ - v1.0: 2,000 NLBSE'23 training issues (500 per class) translated into 13 languages with NLLB-200 distilled 600M, plus the English originals and 10,000 extra English issues: 35,200 examples. Code blocks, stack traces and error messages were left untranslated, as in real issues.
83
+ - v1.1: v1.0 fine-tuned on 97,593 issues, 47,593 recent issues from active repositories labeled by a maintainer plus 50,000 NLBSE'23 issues so it keeps what it already knew. Repositories used for validation and testing were excluded. Validation accuracy on recent issues from unseen repositories went from 65.1% to 72.2%.
84
+ - Same recipe as the English model: AdamW, lr 2.5e-5 encoder / 1e-4 head, label smoothing 0.1, bf16.
85
+ - Calibration: temperature 1.013, ECE 0.037 on recent validation issues.
86
 
87
  ## Limitations
88
 
89
+ - Most non-English issues in real repositories are in Chinese; the other languages are measured mainly on translations, which are cleaner than real issues that mix languages, code and English error messages.
90
+ - 14 languages measured; others are supported by the base model but not evaluated.
91
+ - **question** and **docs** remain the hardest classes, as in the English model.
92
 
93
  ## Credits
94
 
95
+ Base model by [Convai Innovations](https://huggingface.co/convaiinnovations/laya-multilingual) (Apache-2.0). Translation with [NLLB-200](https://huggingface.co/facebook/nllb-200-distilled-600M). Data from the [NLBSE'23 tool competition](https://github.com/nlbse2023/issue-report-classification) and public GitHub issues. Built by [elnachto](https://github.com/elnachto).
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:67429ad745be450ce8b539e135affe5b8057692f84278c1cbd081070f0f78bd9
3
  size 643835692
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:be222fb925b5d9d1c18bad3268004f76f4b97c75f24cae0b9d12382878275111
3
  size 643835692
rl_agent_config.json CHANGED
@@ -11,7 +11,7 @@
11
  "amp_dtype": "bf16",
12
  "model_name": "rl-agent",
13
  "temperature": [
14
- 1.2034,
15
  1.0,
16
  1.0
17
  ],
@@ -26,14 +26,14 @@
26
  "variante": "sin_other",
27
  "epoca": 1,
28
  "validacion": {
29
- "exactitud": 0.6886,
30
- "f1_macro": 0.6783
31
  },
32
- "priores": {
33
- "bug": 0.526,
34
- "feature": 0.37,
35
- "question": 0.06,
36
- "docs": 0.044
37
  }
38
  }
39
  }
 
11
  "amp_dtype": "bf16",
12
  "model_name": "rl-agent",
13
  "temperature": [
14
+ 1.0133,
15
  1.0,
16
  1.0
17
  ],
 
26
  "variante": "sin_other",
27
  "epoca": 1,
28
  "validacion": {
29
+ "exactitud": 0.7223,
30
+ "f1_macro": 0.6368
31
  },
32
+ "mezcla_base": {
33
+ "bug": 0.532,
34
+ "feature": 0.308,
35
+ "question": 0.11,
36
+ "docs": 0.05
37
  }
38
  }
39
  }