Instructions to use Prasanna85/laya-issue-triage with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use Prasanna85/laya-issue-triage with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Create README.md
Browse files
README.md
ADDED
|
@@ -0,0 +1,224 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: convaiinnovations/laya
|
| 4 |
+
language:
|
| 5 |
+
- en
|
| 6 |
+
tags:
|
| 7 |
+
- laya
|
| 8 |
+
- text-classification
|
| 9 |
+
- github-issues
|
| 10 |
+
- issue-triage
|
| 11 |
+
- nlbse
|
| 12 |
+
---
|
| 13 |
+
|
| 14 |
+
# laya-issue-triage
|
| 15 |
+
|
| 16 |
+
<!-- Generated by results/phase3/make_model_card.py from files under results/ and config/; do not edit by hand. -->
|
| 17 |
+
|
| 18 |
+
## Summary
|
| 19 |
+
|
| 20 |
+
A fine-tuned [Laya](https://huggingface.co/convaiinnovations/laya) decision model that classifies a newly
|
| 21 |
+
opened GitHub issue as one of `bug`, `feature`, `question` from its title and body. It is the model behind the laya-triage
|
| 22 |
+
GitHub Action, which auto-applies a label only when `answer_confidence` reaches a per-label threshold and
|
| 23 |
+
otherwise escalates the issue to a human.
|
| 24 |
+
|
| 25 |
+
- Revision `76ece1fb0eb8b32bd5d8c509293c1692a2534805`: weights of training run R1 (`a704b3eadd185f1fa028576cfa50605b6eada4a4`) with the choice
|
| 26 |
+
temperature refit on validation data (4.0188 → 2.6968); nothing else differs.
|
| 27 |
+
- On the NLBSE'24 test split (1,500 issues, 5 repositories) it reaches cross-repo macro-F1
|
| 28 |
+
**0.8020**, accuracy 0.8020, and calibrated ECE 0.0454.
|
| 29 |
+
- With the shipped thresholds, it auto-labels 0.2607 of test issues (only `bug`) at precision
|
| 30 |
+
**0.8747**. This **misses the 0.90 precision target** (see Gating).
|
| 31 |
+
|
| 32 |
+
## Intended use
|
| 33 |
+
|
| 34 |
+
- Suggesting or applying a type label (`bug`, `feature`, `question`) to new issues in English-language
|
| 35 |
+
software repositories, through the laya-triage Action with the confidence gate and a human escalation path.
|
| 36 |
+
The Action's default config sets `mode: dry-run`.
|
| 37 |
+
- Research on calibrated, gated issue classification.
|
| 38 |
+
|
| 39 |
+
## Not intended use
|
| 40 |
+
|
| 41 |
+
- Fully automatic labelling without a confidence gate or human review.
|
| 42 |
+
- Any decision about people (for example rating contributors or prioritising by author).
|
| 43 |
+
- Issue types other than these three (documentation, security reports, and so on), or non-English issues.
|
| 44 |
+
- Repositories very different from the 5 in the training data without first checking precision on a
|
| 45 |
+
sample of that repository's own issues.
|
| 46 |
+
|
| 47 |
+
## How to load
|
| 48 |
+
|
| 49 |
+
CPU, fp32, pinned revision. Download once, then run offline. The question below, including its wording and
|
| 50 |
+
the option order `bug`, `feature`, `question`, is part of the model's contract: it was frozen before training and must be
|
| 51 |
+
passed exactly as written. The state is the issue cleaned by laya-triage's `preprocess()` (see Training data).
|
| 52 |
+
|
| 53 |
+
```python
|
| 54 |
+
import os
|
| 55 |
+
os.environ["HF_HUB_OFFLINE"] = "1" # after one online download of this revision
|
| 56 |
+
os.environ.pop("LAYA_CPU_AMP", None) # keep CPU inference in fp32
|
| 57 |
+
|
| 58 |
+
import laya
|
| 59 |
+
|
| 60 |
+
REPO = "Prasanna85/laya-issue-triage"
|
| 61 |
+
REVISION = "76ece1fb0eb8b32bd5d8c509293c1692a2534805"
|
| 62 |
+
ISSUE_TYPE_QUESTION = {
|
| 63 |
+
"issue_type": {
|
| 64 |
+
"type": "choice",
|
| 65 |
+
"instructions": "You are triaging a newly opened GitHub issue. Using its `title` and `body`, decide which type of issue it is.",
|
| 66 |
+
"criteria": {
|
| 67 |
+
"bug": "reports a defect: a crash, error message, failing build or test, regression, or behaviour that contradicts the documentation",
|
| 68 |
+
"feature": "proposes something new: a new capability, option or API, or an improvement to how existing behaviour works",
|
| 69 |
+
"question": "asks for help: how to use or configure something, why it behaves a certain way, or troubleshooting the author's own setup"
|
| 70 |
+
}
|
| 71 |
+
}
|
| 72 |
+
}
|
| 73 |
+
|
| 74 |
+
agent = laya.load(REPO, revision=REVISION, device="cpu") # max_len 1024, head_max_len 256
|
| 75 |
+
state = {"title": "...", "body": "..."} # from laya_triage.preprocess.preprocess(title, body)
|
| 76 |
+
answer = agent.predict(state, ISSUE_TYPE_QUESTION)["answers"]["issue_type"]
|
| 77 |
+
answer["choice"], answer["answer_confidence"], answer["probabilities"]
|
| 78 |
+
```
|
| 79 |
+
|
| 80 |
+
Gate on `answer_confidence` (the calibrated probability of the chosen label), not on `confidence`.
|
| 81 |
+
|
| 82 |
+
## Training data
|
| 83 |
+
|
| 84 |
+
- [NLBSE'24 issue report classification](https://github.com/nlbse2024/issue-report-classification): issues
|
| 85 |
+
from bitcoin, react, vscode, opencv and tensorflow, labelled `bug`, `feature` or `question`.
|
| 86 |
+
- Of the 1,500 official-train rows, 4 are dropped (their content also appears in
|
| 87 |
+
the official test set). 300 rows are held out as validation (stratified by repo × label, seed 42),
|
| 88 |
+
which leaves **1,196 training rows** (bug 398, feature 398, question 400).
|
| 89 |
+
The official test set (1,500 rows) is used only for the single final evaluation.
|
| 90 |
+
- Preprocessing (shared by training, evaluation and the Action): strip HTML comments; collapse fenced code
|
| 91 |
+
blocks longer than 20 lines to their first 10 and last 5 lines; replace images and bare
|
| 92 |
+
URLs with placeholders; normalise whitespace; cap the body at 6,000 characters. The state
|
| 93 |
+
is `{"title": ..., "body": ...}`, and Laya truncates it to 1024 tokens (11.1% of
|
| 94 |
+
test states are truncated).
|
| 95 |
+
- **License and redistribution.** The dataset's upstream `LICENSE` file is empty, so no license is granted.
|
| 96 |
+
The issue text belongs to the source projects and their authors. No data is redistributed with this model
|
| 97 |
+
or its repository; only synthetic IDs and content hashes are published.
|
| 98 |
+
|
| 99 |
+
## Training procedure
|
| 100 |
+
|
| 101 |
+
| Setting | Value |
|
| 102 |
+
|---|---|
|
| 103 |
+
| Base | `convaiinnovations/laya` @ `55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851` |
|
| 104 |
+
| Recipe | Laya Kaggle notebook (RLCD policy-gradient term + soft cross-entropy on one-hot gold), label smoothing ε 0.0 |
|
| 105 |
+
| Epochs / optimizer steps | 4 / 68 |
|
| 106 |
+
| Effective batch | 64 (8 per micro-batch × 2 GPUs × 4 accumulation steps) |
|
| 107 |
+
| Sequence budget | max_len 1024, head_max_len 256 |
|
| 108 |
+
| Seed | 42 (GPU kernels are not deterministic, so runs are seeded but not bit-for-bit reproducible) |
|
| 109 |
+
| Hardware / time | Kaggle, 2× T4; train time 712 s |
|
| 110 |
+
| Versions (training) | laya 0.3.23, transformers 5.18.0, Python 3.12, torch 2.10.0+cu128 |
|
| 111 |
+
| Versions (evaluation) | laya 0.3.23, torch 2.14.1, transformers 5.18.0, Python 3.11.17; CPU fp32 |
|
| 112 |
+
| Calibration | The notebook's choice temperature 4.0188 (fitted on 119 items held out from the training rows) was replaced by **2.6968**, refit by NLL on the 300 validation issues with `laya.calibrate`. Weights unchanged. |
|
| 113 |
+
|
| 114 |
+
## Evaluation
|
| 115 |
+
|
| 116 |
+
Single run on the NLBSE'24 official test split (1,500 issues: 100 per class in each of 5 repositories), CPU fp32,
|
| 117 |
+
after the thresholds were frozen. The headline metric is **cross-repo macro-F1**: the mean of the 5
|
| 118 |
+
per-repository macro-F1 scores, as in the NLBSE'24 competition.
|
| 119 |
+
|
| 120 |
+
| System | Cross-repo macro-F1 | Pooled macro-F1 | Accuracy | F1 bug | F1 feature | F1 question | ECE |
|
| 121 |
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 122 |
+
| **M1 (this model)** | 0.8020 | 0.8023 | 0.8020 | 0.8033 | 0.8358 | 0.7677 | 0.0454 |
|
| 123 |
+
| B1 TF-IDF + logistic regression | 0.7655 | 0.7666 | 0.7673 | 0.7843 | 0.7925 | 0.7230 | 0.0493 |
|
| 124 |
+
| B2 Laya base, zero-shot, 512/192 (native) | 0.6108 | 0.6160 | 0.6453 | 0.6693 | 0.7757 | 0.4029 | 0.0684 |
|
| 125 |
+
| B2 Laya base, zero-shot, 1024/256 | 0.6257 | 0.6310 | 0.6567 | 0.6773 | 0.7827 | 0.4330 | 0.0647 |
|
| 126 |
+
| B3 SetFit (NLBSE'24, published) | 0.8270 | 0.8263 [c] | 0.8267 [c] | 0.8425 [c] | 0.8555 [c] | 0.7809 [c] | — |
|
| 127 |
+
| B3 RoBERTa (NLBSE'24, published) | 0.7923 | 0.7926 [c] | 0.7927 [c] | 0.8052 [c] | 0.8064 [c] | 0.7663 [c] | — |
|
| 128 |
+
| B3 fastText (NLBSE'24, published) | 0.7184 | 0.7193 [c] | 0.7193 [c] | 0.7362 [c] | 0.7390 [c] | 0.6827 [c] | — |
|
| 129 |
+
|
| 130 |
+
- [c] Derived, not published. NLBSE'24 publishes per-repository, per-class P/R/F1 only; pooled values were
|
| 131 |
+
recovered exactly from them, because every test repository has 100 issues per class.
|
| 132 |
+
- **Protocol differences.** B3 trains one classifier per repository on that repository's official-train rows
|
| 133 |
+
(all of them, including the 4 that overlap test), uses raw text, and is copied rather than re-run.
|
| 134 |
+
M1 and B1 are one model across all 5 repositories, trained on 1,196 rows with model selection,
|
| 135 |
+
temperature and thresholds fitted on the 300-issue validation split. B2 is the base checkpoint
|
| 136 |
+
without fine-tuning. B3 publishes no probabilities, so it has no ECE.
|
| 137 |
+
- Per class (M1), precision / recall: bug 0.8326 / 0.7760; feature 0.8317 / 0.8400; question 0.7467 / 0.7900.
|
| 138 |
+
|
| 139 |
+
Per-repository macro-F1 (test):
|
| 140 |
+
|
| 141 |
+
| System | bitcoin | react | vscode | opencv | tensorflow | Cross-repo |
|
| 142 |
+
|---|---:|---:|---:|---:|---:|---:|
|
| 143 |
+
| **M1** | 0.7645 | 0.8445 | 0.7581 | 0.8002 | 0.8425 | 0.8020 |
|
| 144 |
+
| B1 TF-IDF + logistic regression | 0.7489 | 0.7903 | 0.7490 | 0.7383 | 0.8012 | 0.7655 |
|
| 145 |
+
| B2 Laya base, zero-shot, 512/192 (native) | 0.6213 | 0.7081 | 0.5053 | 0.5358 | 0.6833 | 0.6108 |
|
| 146 |
+
| B2 Laya base, zero-shot, 1024/256 | 0.6298 | 0.7358 | 0.5106 | 0.5622 | 0.6901 | 0.6257 |
|
| 147 |
+
|
| 148 |
+
Calibration (test, 15 bins):
|
| 149 |
+
|
| 150 |
+
| Probabilities | ECE | Brier |
|
| 151 |
+
|---|---:|---:|
|
| 152 |
+
| Calibrated (T = 2.6968, shipped) | 0.0454 | 0.2890 |
|
| 153 |
+
| Uncalibrated (T = 1) | 0.1465 | 0.3301 |
|
| 154 |
+
|
| 155 |
+
Paired comparisons (test): difference in cross-repo macro-F1, with a 95% CI from a paired bootstrap
|
| 156 |
+
(2,000 resamples of issues, seed 42), and the exact McNemar p-value.
|
| 157 |
+
|
| 158 |
+
| Comparison | Δ cross-repo macro-F1 | 95% CI | McNemar p |
|
| 159 |
+
|---|---:|---|---:|
|
| 160 |
+
| M1 vs B1 TF-IDF + logistic regression | +0.0364 | [+0.0151, +0.0588] | 0.0026 |
|
| 161 |
+
| M1 vs B2 Laya base, zero-shot, 512/192 (native) | +0.1912 | [+0.1661, +0.2167] | 4e-30 |
|
| 162 |
+
| M1 vs B2 Laya base, zero-shot, 1024/256 | +0.1763 | [+0.1505, +0.2016] | 1.3e-26 |
|
| 163 |
+
|
| 164 |
+
M1 beats B1 and both B2 settings on the headline metric, and all three CIs exclude 0. Relative to the published SetFit
|
| 165 |
+
baseline (B3) it is -2.5 points (not a paired comparison; see the protocol differences).
|
| 166 |
+
|
| 167 |
+
## Gating
|
| 168 |
+
|
| 169 |
+
Per-label thresholds on `answer_confidence` were fitted on validation before the test run: the smallest
|
| 170 |
+
threshold whose bootstrap lower bound of precision (5th percentile of 2,000 resamples, at
|
| 171 |
+
least 20 issues) is ≥ 0.90; 1.01 means the label is never auto-applied.
|
| 172 |
+
|
| 173 |
+
| Label | Threshold | Val lower bound | Applied on test | Test coverage | Test precision |
|
| 174 |
+
|---|---|---:|---:|---:|---:|
|
| 175 |
+
| bug | 0.6033 | 0.9091 | 391 | 0.2607 | 0.8747 |
|
| 176 |
+
| feature | never (1.01) | none ≥ 0.90 | 0 | 0.0000 | — |
|
| 177 |
+
| question | never (1.01) | none ≥ 0.90 | 0 | 0.0000 | — |
|
| 178 |
+
| **all** | | | 391 of 1,500 | **0.2607** | **0.8747** |
|
| 179 |
+
|
| 180 |
+
**The 0.90 precision target was missed on test.** The `bug` threshold had validation precision
|
| 181 |
+
0.9481 (lower bound 0.9091, 77 issues, in-sample) but gave
|
| 182 |
+
0.8747 on test. `feature` and `question` never reached the lower bound on validation, so they are
|
| 183 |
+
always escalated (project success criterion 3: met for 0 of 3 labels). For comparison, the point-estimate
|
| 184 |
+
thresholds from validation would auto-label 0.8747 of test issues at precision
|
| 185 |
+
0.8491. Anyone using this model should verify precision on their own repository before
|
| 186 |
+
switching the Action from dry-run to apply.
|
| 187 |
+
|
| 188 |
+
## Limitations
|
| 189 |
+
|
| 190 |
+
- **Balanced data.** Test has exactly 100 issues per class per repository; train is near-balanced
|
| 191 |
+
(bug 398, feature 398, question 400). Real repositories have very different class mixes, so precision at a given threshold will differ.
|
| 192 |
+
- **Random, not temporal, split.** The numbers do not measure drift to future issues.
|
| 193 |
+
- **5 repositories.** Large, well-known projects only; performance elsewhere is unknown.
|
| 194 |
+
- **Label noise**, especially for `question`. Labels come from upstream repository labels. Some issues are
|
| 195 |
+
plausibly mislabelled, and the author's own template choice often disagrees with a `question` label.
|
| 196 |
+
`question` is the weakest class (test F1 0.7677).
|
| 197 |
+
- **Validation overstated test performance.** Cross-repo macro-F1 was 0.8691 on the
|
| 198 |
+
300-issue validation split and 0.8020 on test.
|
| 199 |
+
- **Thresholds fitted on validation do not transfer exactly.** The `bug` threshold lost
|
| 200 |
+
7.3 points of precision from validation to test.
|
| 201 |
+
- **English only.** Trained and evaluated on English issues.
|
| 202 |
+
- **Latency was measured on a local CPU, not on a GitHub runner** (local Apple M4 Pro CPU, fp32, model preloaded):
|
| 203 |
+
p50 178 ms and p95 436 ms per issue at 8 threads
|
| 204 |
+
(test run); p50 288 ms and p95 690 ms at 2 threads (100 validation issues).
|
| 205 |
+
Cold load: `laya.load` 3.14 s, peak RSS 2,863 MiB.
|
| 206 |
+
- **11.1% of test issues are truncated** at 1024 tokens; text beyond that
|
| 207 |
+
is not seen.
|
| 208 |
+
|
| 209 |
+
## Citation
|
| 210 |
+
|
| 211 |
+
Laya:
|
| 212 |
+
|
| 213 |
+
> Convai Innovations / NandhaKishorM. Laya: non-autoregressive decision models, v0.3.23 (Apache-2.0). https://github.com/NandhaKishorM/laya . Checkpoints: https://huggingface.co/convaiinnovations/laya
|
| 214 |
+
|
| 215 |
+
NLBSE'24:
|
| 216 |
+
|
| 217 |
+
```bibtex
|
| 218 |
+
@inproceedings{nlbse2024,
|
| 219 |
+
author={Kallis, Rafael and Colavito, Giuseppe and Al-Kaswan, Ali and Pascarella, Luca and Chaparro, Oscar and Rani, Pooja},
|
| 220 |
+
title={The NLBSE'24 Tool Competition},
|
| 221 |
+
booktitle={Proceedings of The 3rd International Workshop on Natural Language-based Software Engineering (NLBSE'24)},
|
| 222 |
+
year={2024}
|
| 223 |
+
}
|
| 224 |
+
```
|