--- license: apache-2.0 base_model: convaiinnovations/laya language: - en tags: - laya - text-classification - github-issues - issue-triage - nlbse --- # laya-issue-triage ## Summary A fine-tuned [Laya](https://huggingface.co/convaiinnovations/laya) decision model that classifies a newly opened GitHub issue as one of `bug`, `feature`, `question` from its title and body. It is the model behind the [laya-triage](https://github.com/Prasanna-KS-85/laya-triage) GitHub Action, which auto-applies a label only when `answer_confidence` reaches a per-label threshold and otherwise escalates the issue to a human. Source code, install instructions and the full evaluation write-up: https://github.com/Prasanna-KS-85/laya-triage. - Revision `76ece1fb0eb8b32bd5d8c509293c1692a2534805`: weights of training run R1 (`a704b3eadd185f1fa028576cfa50605b6eada4a4`) with the choice temperature refit on validation data (4.0188 → 2.6968); nothing else differs. - On the NLBSE'24 test split (1,500 issues, 5 repositories) it reaches cross-repo macro-F1 **0.8020**, accuracy 0.8020, and calibrated ECE 0.0454. - With the shipped thresholds, it auto-labels 0.2607 of test issues (only `bug`) at precision **0.8747**. This **misses the 0.90 precision target** (see Gating). ## Intended use - Suggesting or applying a type label (`bug`, `feature`, `question`) to new issues in English-language software repositories, through the laya-triage Action with the confidence gate and a human escalation path. The Action's default config sets `mode: dry-run`. - Research on calibrated, gated issue classification. ## Not intended use - Fully automatic labelling without a confidence gate or human review. - Any decision about people (for example rating contributors or prioritising by author). - Issue types other than these three (documentation, security reports, and so on), or non-English issues. - Repositories very different from the 5 in the training data without first checking precision on a sample of that repository's own issues. ## How to load CPU, fp32, pinned revision. Download once, then run offline. The question below, including its wording and the option order `bug`, `feature`, `question`, is part of the model's contract: it was frozen before training and must be passed exactly as written. The state is the issue cleaned by laya-triage's `preprocess()` (see Training data). ```python import os os.environ["HF_HUB_OFFLINE"] = "1" # after one online download of this revision os.environ.pop("LAYA_CPU_AMP", None) # keep CPU inference in fp32 import laya REPO = "Prasanna85/laya-issue-triage" REVISION = "76ece1fb0eb8b32bd5d8c509293c1692a2534805" ISSUE_TYPE_QUESTION = { "issue_type": { "type": "choice", "instructions": "You are triaging a newly opened GitHub issue. Using its `title` and `body`, decide which type of issue it is.", "criteria": { "bug": "reports a defect: a crash, error message, failing build or test, regression, or behaviour that contradicts the documentation", "feature": "proposes something new: a new capability, option or API, or an improvement to how existing behaviour works", "question": "asks for help: how to use or configure something, why it behaves a certain way, or troubleshooting the author's own setup" } } } agent = laya.load(REPO, revision=REVISION, device="cpu") # max_len 1024, head_max_len 256 state = {"title": "...", "body": "..."} # from laya_triage.preprocess.preprocess(title, body) answer = agent.predict(state, ISSUE_TYPE_QUESTION)["answers"]["issue_type"] answer["choice"], answer["answer_confidence"], answer["probabilities"] ``` Gate on `answer_confidence` (the calibrated probability of the chosen label), not on `confidence`. ### With the laya-triage package The [laya-triage](https://github.com/Prasanna-KS-85/laya-triage) package wraps the same steps: it brings the frozen question and `preprocess()`, so nothing is retyped. Install it with `pip install "laya_triage @ git+https://github.com/Prasanna-KS-85/laya-triage@v1.0.0"` (on Linux, install the CPU-only torch wheel first; see the repository's docs/USING_THE_MODEL.md), then: ```python """Classify one issue with the Laya Triage model on CPU (fp32). Install: see docs/USING_THE_MODEL.md.""" # The package brings the frozen question and the shared preprocess(): nothing is retyped. from laya_triage.classifier import Classifier from laya_triage.preprocess import preprocess REPO = "Prasanna85/laya-issue-triage" REVISION = "76ece1fb0eb8b32bd5d8c509293c1692a2534805" # pinned: the revision the Action uses MAX_LEN = 1024 # must equal the training budget BUG_THRESHOLD = 0.6033 # the Action's default gate; feature and question are never auto-applied model = Classifier(REPO, REVISION, MAX_LEN) # downloads once, then loads from the local cache state = preprocess("App closes when I open the export dialog", "Steps: open a notebook, choose File > Export. The app exits with KeyError: 'last_format'.") result = model.classify(state, issue_number=0) print(result.probabilities) # {'bug': ..., 'feature': ..., 'question': ...} confident_bug = result.label == "bug" and result.answer_confidence >= BUG_THRESHOLD print("apply bug" if confident_bug else "escalate to a human") ``` ## Training data - [NLBSE'24 issue report classification](https://github.com/nlbse2024/issue-report-classification): issues from bitcoin, react, vscode, opencv and tensorflow, labelled `bug`, `feature` or `question`. - Of the 1,500 official-train rows, 4 are dropped (their content also appears in the official test set). 300 rows are held out as validation (stratified by repo × label, seed 42), which leaves **1,196 training rows** (bug 398, feature 398, question 400). The official test set (1,500 rows) is used only for the single final evaluation. - Preprocessing (shared by training, evaluation and the Action): strip HTML comments; collapse fenced code blocks longer than 20 lines to their first 10 and last 5 lines; replace images and bare URLs with placeholders; normalise whitespace; cap the body at 6,000 characters. The state is `{"title": ..., "body": ...}`, and Laya truncates it to 1024 tokens (11.1% of test states are truncated). - **License and redistribution.** The dataset's upstream `LICENSE` file is empty, so no license is granted. The issue text belongs to the source projects and their authors. No data is redistributed with this model or its repository; only synthetic IDs and content hashes are published. ## Training procedure | Setting | Value | |---|---| | Base | `convaiinnovations/laya` @ `55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851` | | Recipe | Laya Kaggle notebook (RLCD policy-gradient term + soft cross-entropy on one-hot gold), label smoothing ε 0.0 | | Epochs / optimizer steps | 4 / 68 | | Effective batch | 64 (8 per micro-batch × 2 GPUs × 4 accumulation steps) | | Sequence budget | max_len 1024, head_max_len 256 | | Seed | 42 (GPU kernels are not deterministic, so runs are seeded but not bit-for-bit reproducible) | | Hardware / time | Kaggle, 2× T4; train time 712 s | | Versions (training) | laya 0.3.23, transformers 5.18.0, Python 3.12, torch 2.10.0+cu128 | | Versions (evaluation) | laya 0.3.23, torch 2.14.1, transformers 5.18.0, Python 3.11.17; CPU fp32 | | Calibration | The notebook's choice temperature 4.0188 (fitted on 119 items held out from the training rows) was replaced by **2.6968**, refit by NLL on the 300 validation issues with `laya.calibrate`. Weights unchanged. | ## Evaluation Single run on the NLBSE'24 official test split (1,500 issues: 100 per class in each of 5 repositories), CPU fp32, after the thresholds were frozen. The headline metric is **cross-repo macro-F1**: the mean of the 5 per-repository macro-F1 scores, as in the NLBSE'24 competition. | System | Cross-repo macro-F1 | Pooled macro-F1 | Accuracy | F1 bug | F1 feature | F1 question | ECE | |---|---:|---:|---:|---:|---:|---:|---:| | **M1 (this model)** | 0.8020 | 0.8023 | 0.8020 | 0.8033 | 0.8358 | 0.7677 | 0.0454 | | B1 TF-IDF + logistic regression | 0.7655 | 0.7666 | 0.7673 | 0.7843 | 0.7925 | 0.7230 | 0.0493 | | B2 Laya base, zero-shot, 512/192 (native) | 0.6108 | 0.6160 | 0.6453 | 0.6693 | 0.7757 | 0.4029 | 0.0684 | | B2 Laya base, zero-shot, 1024/256 | 0.6257 | 0.6310 | 0.6567 | 0.6773 | 0.7827 | 0.4330 | 0.0647 | | B3 SetFit (NLBSE'24, published) | 0.8270 | 0.8263 [c] | 0.8267 [c] | 0.8425 [c] | 0.8555 [c] | 0.7809 [c] | — | | B3 RoBERTa (NLBSE'24, published) | 0.7923 | 0.7926 [c] | 0.7927 [c] | 0.8052 [c] | 0.8064 [c] | 0.7663 [c] | — | | B3 fastText (NLBSE'24, published) | 0.7184 | 0.7193 [c] | 0.7193 [c] | 0.7362 [c] | 0.7390 [c] | 0.6827 [c] | — | - [c] Derived, not published. NLBSE'24 publishes per-repository, per-class P/R/F1 only; pooled values were recovered exactly from them, because every test repository has 100 issues per class. - **Protocol differences.** B3 trains one classifier per repository, each on only that repository's 300 official-train rows (together the 5 classifiers use all 1,500 rows, including the 4 whose content also appears in test). The SetFit input is raw text; the RoBERTa and fastText setups were not inspected. B3 is copied rather than re-run. M1 and B1 are one model across all 5 repositories, trained on 1,196 rows with model selection, temperature and thresholds fitted on the 300-issue validation split. B2 is the base checkpoint without fine-tuning. B3 publishes no probabilities, so it has no ECE. - Per class (M1), precision / recall: bug 0.8326 / 0.7760; feature 0.8317 / 0.8400; question 0.7467 / 0.7900. Per-repository macro-F1 (test): | System | bitcoin | react | vscode | opencv | tensorflow | Cross-repo | |---|---:|---:|---:|---:|---:|---:| | **M1** | 0.7645 | 0.8445 | 0.7581 | 0.8002 | 0.8425 | 0.8020 | | B1 TF-IDF + logistic regression | 0.7489 | 0.7903 | 0.7490 | 0.7383 | 0.8012 | 0.7655 | | B2 Laya base, zero-shot, 512/192 (native) | 0.6213 | 0.7081 | 0.5053 | 0.5358 | 0.6833 | 0.6108 | | B2 Laya base, zero-shot, 1024/256 | 0.6298 | 0.7358 | 0.5106 | 0.5622 | 0.6901 | 0.6257 | Calibration (test, 15 bins): | Probabilities | ECE | Brier | |---|---:|---:| | Calibrated (T = 2.6968, shipped) | 0.0454 | 0.2890 | | Uncalibrated (T = 1) | 0.1465 | 0.3301 | Paired comparisons (test): difference in cross-repo macro-F1, with a 95% CI from a paired bootstrap (2,000 resamples of issues, seed 42), and the exact McNemar p-value. | Comparison | Δ cross-repo macro-F1 | 95% CI | McNemar p | |---|---:|---|---:| | M1 vs B1 TF-IDF + logistic regression | +0.0364 | [+0.0151, +0.0588] | 0.0026 | | M1 vs B2 Laya base, zero-shot, 512/192 (native) | +0.1912 | [+0.1661, +0.2167] | 4e-30 | | M1 vs B2 Laya base, zero-shot, 1024/256 | +0.1763 | [+0.1505, +0.2016] | 1.3e-26 | M1 beats B1 and both B2 settings on the headline metric, and all three CIs exclude 0. Relative to the published SetFit baseline (B3) it is -2.5 points (not a paired comparison; see the protocol differences). ## Gating Per-label thresholds on `answer_confidence` were fitted on validation before the test run: the smallest threshold whose bootstrap lower bound of precision (5th percentile of 2,000 resamples, at least 20 issues) is ≥ 0.90; 1.01 means the label is never auto-applied. | Label | Threshold | Val lower bound | Applied on test | Test coverage | Test precision | |---|---|---:|---:|---:|---:| | bug | 0.6033 | 0.9091 | 391 | 0.2607 | 0.8747 | | feature | never (1.01) | none ≥ 0.90 | 0 | 0.0000 | — | | question | never (1.01) | none ≥ 0.90 | 0 | 0.0000 | — | | **all** | | | 391 of 1,500 | **0.2607** | **0.8747** | **The 0.90 precision target was missed on test.** The `bug` threshold had validation precision 0.9481 (lower bound 0.9091, 77 issues, in-sample) but gave 0.8747 on test. `feature` and `question` never reached the lower bound on validation, so they are always escalated (project success criterion 3: met for 0 of 3 labels). For comparison, the point-estimate thresholds from validation would auto-label 0.8747 of test issues at precision 0.8491. Anyone using this model should verify precision on their own repository before switching the Action from dry-run to apply. ## Limitations - **Balanced data.** Test has exactly 100 issues per class per repository; train is near-balanced (bug 398, feature 398, question 400). Real repositories have very different class mixes, so precision at a given threshold will differ. - **Random, not temporal, split.** The numbers do not measure drift to future issues. - **5 repositories.** Large, well-known projects only; performance elsewhere is unknown. - **Label noise**, especially for `question`. Labels come from upstream repository labels. Some issues are plausibly mislabelled, and the author's own template choice often disagrees with a `question` label. `question` is the weakest class (test F1 0.7677). - **Validation overstated test performance.** Cross-repo macro-F1 was 0.8691 on the 300-issue validation split and 0.8020 on test. - **Thresholds fitted on validation do not transfer exactly.** The `bug` threshold lost 7.3 points of precision from validation to test. - **English only.** Trained and evaluated on English issues. - Exactly three classes (bug, feature, question); adding a class requires new labelled data, retraining and recalibrating. - **Latency was measured on a local CPU, not on a GitHub runner** (local Apple M4 Pro CPU, fp32, model preloaded): p50 178 ms and p95 436 ms per issue at 8 threads (test run); p50 288 ms and p95 690 ms at 2 threads (100 validation issues). Cold load: `laya.load` 3.14 s, peak RSS 2,863 MiB. - **11.1% of test issues are truncated** at 1024 tokens; text beyond that is not seen. ## Citation Laya: > Convai Innovations / NandhaKishorM. Laya: non-autoregressive decision models, v0.3.23 (Apache-2.0). https://github.com/NandhaKishorM/laya . Checkpoints: https://huggingface.co/convaiinnovations/laya NLBSE'24: ```bibtex @inproceedings{nlbse2024, author={Kallis, Rafael and Colavito, Giuseppe and Al-Kaswan, Ali and Pascarella, Luca and Chaparro, Oscar and Rani, Pooja}, title={The NLBSE'24 Tool Competition}, booktitle={Proceedings of The 3rd International Workshop on Natural Language-based Software Engineering (NLBSE'24)}, year={2024} } ```