Prasanna85 commited on
Commit
bbfde5c
·
verified ·
1 Parent(s): 76ece1f

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +224 -0
README.md ADDED
@@ -0,0 +1,224 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: convaiinnovations/laya
4
+ language:
5
+ - en
6
+ tags:
7
+ - laya
8
+ - text-classification
9
+ - github-issues
10
+ - issue-triage
11
+ - nlbse
12
+ ---
13
+
14
+ # laya-issue-triage
15
+
16
+ <!-- Generated by results/phase3/make_model_card.py from files under results/ and config/; do not edit by hand. -->
17
+
18
+ ## Summary
19
+
20
+ A fine-tuned [Laya](https://huggingface.co/convaiinnovations/laya) decision model that classifies a newly
21
+ opened GitHub issue as one of `bug`, `feature`, `question` from its title and body. It is the model behind the laya-triage
22
+ GitHub Action, which auto-applies a label only when `answer_confidence` reaches a per-label threshold and
23
+ otherwise escalates the issue to a human.
24
+
25
+ - Revision `76ece1fb0eb8b32bd5d8c509293c1692a2534805`: weights of training run R1 (`a704b3eadd185f1fa028576cfa50605b6eada4a4`) with the choice
26
+ temperature refit on validation data (4.0188 → 2.6968); nothing else differs.
27
+ - On the NLBSE'24 test split (1,500 issues, 5 repositories) it reaches cross-repo macro-F1
28
+ **0.8020**, accuracy 0.8020, and calibrated ECE 0.0454.
29
+ - With the shipped thresholds, it auto-labels 0.2607 of test issues (only `bug`) at precision
30
+ **0.8747**. This **misses the 0.90 precision target** (see Gating).
31
+
32
+ ## Intended use
33
+
34
+ - Suggesting or applying a type label (`bug`, `feature`, `question`) to new issues in English-language
35
+ software repositories, through the laya-triage Action with the confidence gate and a human escalation path.
36
+ The Action's default config sets `mode: dry-run`.
37
+ - Research on calibrated, gated issue classification.
38
+
39
+ ## Not intended use
40
+
41
+ - Fully automatic labelling without a confidence gate or human review.
42
+ - Any decision about people (for example rating contributors or prioritising by author).
43
+ - Issue types other than these three (documentation, security reports, and so on), or non-English issues.
44
+ - Repositories very different from the 5 in the training data without first checking precision on a
45
+ sample of that repository's own issues.
46
+
47
+ ## How to load
48
+
49
+ CPU, fp32, pinned revision. Download once, then run offline. The question below, including its wording and
50
+ the option order `bug`, `feature`, `question`, is part of the model's contract: it was frozen before training and must be
51
+ passed exactly as written. The state is the issue cleaned by laya-triage's `preprocess()` (see Training data).
52
+
53
+ ```python
54
+ import os
55
+ os.environ["HF_HUB_OFFLINE"] = "1" # after one online download of this revision
56
+ os.environ.pop("LAYA_CPU_AMP", None) # keep CPU inference in fp32
57
+
58
+ import laya
59
+
60
+ REPO = "Prasanna85/laya-issue-triage"
61
+ REVISION = "76ece1fb0eb8b32bd5d8c509293c1692a2534805"
62
+ ISSUE_TYPE_QUESTION = {
63
+ "issue_type": {
64
+ "type": "choice",
65
+ "instructions": "You are triaging a newly opened GitHub issue. Using its `title` and `body`, decide which type of issue it is.",
66
+ "criteria": {
67
+ "bug": "reports a defect: a crash, error message, failing build or test, regression, or behaviour that contradicts the documentation",
68
+ "feature": "proposes something new: a new capability, option or API, or an improvement to how existing behaviour works",
69
+ "question": "asks for help: how to use or configure something, why it behaves a certain way, or troubleshooting the author's own setup"
70
+ }
71
+ }
72
+ }
73
+
74
+ agent = laya.load(REPO, revision=REVISION, device="cpu") # max_len 1024, head_max_len 256
75
+ state = {"title": "...", "body": "..."} # from laya_triage.preprocess.preprocess(title, body)
76
+ answer = agent.predict(state, ISSUE_TYPE_QUESTION)["answers"]["issue_type"]
77
+ answer["choice"], answer["answer_confidence"], answer["probabilities"]
78
+ ```
79
+
80
+ Gate on `answer_confidence` (the calibrated probability of the chosen label), not on `confidence`.
81
+
82
+ ## Training data
83
+
84
+ - [NLBSE'24 issue report classification](https://github.com/nlbse2024/issue-report-classification): issues
85
+ from bitcoin, react, vscode, opencv and tensorflow, labelled `bug`, `feature` or `question`.
86
+ - Of the 1,500 official-train rows, 4 are dropped (their content also appears in
87
+ the official test set). 300 rows are held out as validation (stratified by repo × label, seed 42),
88
+ which leaves **1,196 training rows** (bug 398, feature 398, question 400).
89
+ The official test set (1,500 rows) is used only for the single final evaluation.
90
+ - Preprocessing (shared by training, evaluation and the Action): strip HTML comments; collapse fenced code
91
+ blocks longer than 20 lines to their first 10 and last 5 lines; replace images and bare
92
+ URLs with placeholders; normalise whitespace; cap the body at 6,000 characters. The state
93
+ is `{"title": ..., "body": ...}`, and Laya truncates it to 1024 tokens (11.1% of
94
+ test states are truncated).
95
+ - **License and redistribution.** The dataset's upstream `LICENSE` file is empty, so no license is granted.
96
+ The issue text belongs to the source projects and their authors. No data is redistributed with this model
97
+ or its repository; only synthetic IDs and content hashes are published.
98
+
99
+ ## Training procedure
100
+
101
+ | Setting | Value |
102
+ |---|---|
103
+ | Base | `convaiinnovations/laya` @ `55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851` |
104
+ | Recipe | Laya Kaggle notebook (RLCD policy-gradient term + soft cross-entropy on one-hot gold), label smoothing ε 0.0 |
105
+ | Epochs / optimizer steps | 4 / 68 |
106
+ | Effective batch | 64 (8 per micro-batch × 2 GPUs × 4 accumulation steps) |
107
+ | Sequence budget | max_len 1024, head_max_len 256 |
108
+ | Seed | 42 (GPU kernels are not deterministic, so runs are seeded but not bit-for-bit reproducible) |
109
+ | Hardware / time | Kaggle, 2× T4; train time 712 s |
110
+ | Versions (training) | laya 0.3.23, transformers 5.18.0, Python 3.12, torch 2.10.0+cu128 |
111
+ | Versions (evaluation) | laya 0.3.23, torch 2.14.1, transformers 5.18.0, Python 3.11.17; CPU fp32 |
112
+ | Calibration | The notebook's choice temperature 4.0188 (fitted on 119 items held out from the training rows) was replaced by **2.6968**, refit by NLL on the 300 validation issues with `laya.calibrate`. Weights unchanged. |
113
+
114
+ ## Evaluation
115
+
116
+ Single run on the NLBSE'24 official test split (1,500 issues: 100 per class in each of 5 repositories), CPU fp32,
117
+ after the thresholds were frozen. The headline metric is **cross-repo macro-F1**: the mean of the 5
118
+ per-repository macro-F1 scores, as in the NLBSE'24 competition.
119
+
120
+ | System | Cross-repo macro-F1 | Pooled macro-F1 | Accuracy | F1 bug | F1 feature | F1 question | ECE |
121
+ |---|---:|---:|---:|---:|---:|---:|---:|
122
+ | **M1 (this model)** | 0.8020 | 0.8023 | 0.8020 | 0.8033 | 0.8358 | 0.7677 | 0.0454 |
123
+ | B1 TF-IDF + logistic regression | 0.7655 | 0.7666 | 0.7673 | 0.7843 | 0.7925 | 0.7230 | 0.0493 |
124
+ | B2 Laya base, zero-shot, 512/192 (native) | 0.6108 | 0.6160 | 0.6453 | 0.6693 | 0.7757 | 0.4029 | 0.0684 |
125
+ | B2 Laya base, zero-shot, 1024/256 | 0.6257 | 0.6310 | 0.6567 | 0.6773 | 0.7827 | 0.4330 | 0.0647 |
126
+ | B3 SetFit (NLBSE'24, published) | 0.8270 | 0.8263 [c] | 0.8267 [c] | 0.8425 [c] | 0.8555 [c] | 0.7809 [c] | — |
127
+ | B3 RoBERTa (NLBSE'24, published) | 0.7923 | 0.7926 [c] | 0.7927 [c] | 0.8052 [c] | 0.8064 [c] | 0.7663 [c] | — |
128
+ | B3 fastText (NLBSE'24, published) | 0.7184 | 0.7193 [c] | 0.7193 [c] | 0.7362 [c] | 0.7390 [c] | 0.6827 [c] | — |
129
+
130
+ - [c] Derived, not published. NLBSE'24 publishes per-repository, per-class P/R/F1 only; pooled values were
131
+ recovered exactly from them, because every test repository has 100 issues per class.
132
+ - **Protocol differences.** B3 trains one classifier per repository on that repository's official-train rows
133
+ (all of them, including the 4 that overlap test), uses raw text, and is copied rather than re-run.
134
+ M1 and B1 are one model across all 5 repositories, trained on 1,196 rows with model selection,
135
+ temperature and thresholds fitted on the 300-issue validation split. B2 is the base checkpoint
136
+ without fine-tuning. B3 publishes no probabilities, so it has no ECE.
137
+ - Per class (M1), precision / recall: bug 0.8326 / 0.7760; feature 0.8317 / 0.8400; question 0.7467 / 0.7900.
138
+
139
+ Per-repository macro-F1 (test):
140
+
141
+ | System | bitcoin | react | vscode | opencv | tensorflow | Cross-repo |
142
+ |---|---:|---:|---:|---:|---:|---:|
143
+ | **M1** | 0.7645 | 0.8445 | 0.7581 | 0.8002 | 0.8425 | 0.8020 |
144
+ | B1 TF-IDF + logistic regression | 0.7489 | 0.7903 | 0.7490 | 0.7383 | 0.8012 | 0.7655 |
145
+ | B2 Laya base, zero-shot, 512/192 (native) | 0.6213 | 0.7081 | 0.5053 | 0.5358 | 0.6833 | 0.6108 |
146
+ | B2 Laya base, zero-shot, 1024/256 | 0.6298 | 0.7358 | 0.5106 | 0.5622 | 0.6901 | 0.6257 |
147
+
148
+ Calibration (test, 15 bins):
149
+
150
+ | Probabilities | ECE | Brier |
151
+ |---|---:|---:|
152
+ | Calibrated (T = 2.6968, shipped) | 0.0454 | 0.2890 |
153
+ | Uncalibrated (T = 1) | 0.1465 | 0.3301 |
154
+
155
+ Paired comparisons (test): difference in cross-repo macro-F1, with a 95% CI from a paired bootstrap
156
+ (2,000 resamples of issues, seed 42), and the exact McNemar p-value.
157
+
158
+ | Comparison | Δ cross-repo macro-F1 | 95% CI | McNemar p |
159
+ |---|---:|---|---:|
160
+ | M1 vs B1 TF-IDF + logistic regression | +0.0364 | [+0.0151, +0.0588] | 0.0026 |
161
+ | M1 vs B2 Laya base, zero-shot, 512/192 (native) | +0.1912 | [+0.1661, +0.2167] | 4e-30 |
162
+ | M1 vs B2 Laya base, zero-shot, 1024/256 | +0.1763 | [+0.1505, +0.2016] | 1.3e-26 |
163
+
164
+ M1 beats B1 and both B2 settings on the headline metric, and all three CIs exclude 0. Relative to the published SetFit
165
+ baseline (B3) it is -2.5 points (not a paired comparison; see the protocol differences).
166
+
167
+ ## Gating
168
+
169
+ Per-label thresholds on `answer_confidence` were fitted on validation before the test run: the smallest
170
+ threshold whose bootstrap lower bound of precision (5th percentile of 2,000 resamples, at
171
+ least 20 issues) is ≥ 0.90; 1.01 means the label is never auto-applied.
172
+
173
+ | Label | Threshold | Val lower bound | Applied on test | Test coverage | Test precision |
174
+ |---|---|---:|---:|---:|---:|
175
+ | bug | 0.6033 | 0.9091 | 391 | 0.2607 | 0.8747 |
176
+ | feature | never (1.01) | none ≥ 0.90 | 0 | 0.0000 | — |
177
+ | question | never (1.01) | none ≥ 0.90 | 0 | 0.0000 | — |
178
+ | **all** | | | 391 of 1,500 | **0.2607** | **0.8747** |
179
+
180
+ **The 0.90 precision target was missed on test.** The `bug` threshold had validation precision
181
+ 0.9481 (lower bound 0.9091, 77 issues, in-sample) but gave
182
+ 0.8747 on test. `feature` and `question` never reached the lower bound on validation, so they are
183
+ always escalated (project success criterion 3: met for 0 of 3 labels). For comparison, the point-estimate
184
+ thresholds from validation would auto-label 0.8747 of test issues at precision
185
+ 0.8491. Anyone using this model should verify precision on their own repository before
186
+ switching the Action from dry-run to apply.
187
+
188
+ ## Limitations
189
+
190
+ - **Balanced data.** Test has exactly 100 issues per class per repository; train is near-balanced
191
+ (bug 398, feature 398, question 400). Real repositories have very different class mixes, so precision at a given threshold will differ.
192
+ - **Random, not temporal, split.** The numbers do not measure drift to future issues.
193
+ - **5 repositories.** Large, well-known projects only; performance elsewhere is unknown.
194
+ - **Label noise**, especially for `question`. Labels come from upstream repository labels. Some issues are
195
+ plausibly mislabelled, and the author's own template choice often disagrees with a `question` label.
196
+ `question` is the weakest class (test F1 0.7677).
197
+ - **Validation overstated test performance.** Cross-repo macro-F1 was 0.8691 on the
198
+ 300-issue validation split and 0.8020 on test.
199
+ - **Thresholds fitted on validation do not transfer exactly.** The `bug` threshold lost
200
+ 7.3 points of precision from validation to test.
201
+ - **English only.** Trained and evaluated on English issues.
202
+ - **Latency was measured on a local CPU, not on a GitHub runner** (local Apple M4 Pro CPU, fp32, model preloaded):
203
+ p50 178 ms and p95 436 ms per issue at 8 threads
204
+ (test run); p50 288 ms and p95 690 ms at 2 threads (100 validation issues).
205
+ Cold load: `laya.load` 3.14 s, peak RSS 2,863 MiB.
206
+ - **11.1% of test issues are truncated** at 1024 tokens; text beyond that
207
+ is not seen.
208
+
209
+ ## Citation
210
+
211
+ Laya:
212
+
213
+ > Convai Innovations / NandhaKishorM. Laya: non-autoregressive decision models, v0.3.23 (Apache-2.0). https://github.com/NandhaKishorM/laya . Checkpoints: https://huggingface.co/convaiinnovations/laya
214
+
215
+ NLBSE'24:
216
+
217
+ ```bibtex
218
+ @inproceedings{nlbse2024,
219
+ author={Kallis, Rafael and Colavito, Giuseppe and Al-Kaswan, Ali and Pascarella, Luca and Chaparro, Oscar and Rani, Pooja},
220
+ title={The NLBSE'24 Tool Competition},
221
+ booktitle={Proceedings of The 3rd International Workshop on Natural Language-based Software Engineering (NLBSE'24)},
222
+ year={2024}
223
+ }
224
+ ```