laya-issue-triage

Summary

A fine-tuned Laya decision model that classifies a newly opened GitHub issue as one of bug, feature, question from its title and body. It is the model behind the laya-triage GitHub Action, which auto-applies a label only when answer_confidence reaches a per-label threshold and otherwise escalates the issue to a human. Source code, install instructions and the full evaluation write-up: https://github.com/Prasanna-KS-85/laya-triage.

  • Revision 76ece1fb0eb8b32bd5d8c509293c1692a2534805: weights of training run R1 (a704b3eadd185f1fa028576cfa50605b6eada4a4) with the choice temperature refit on validation data (4.0188 → 2.6968); nothing else differs.
  • On the NLBSE'24 test split (1,500 issues, 5 repositories) it reaches cross-repo macro-F1 0.8020, accuracy 0.8020, and calibrated ECE 0.0454.
  • With the shipped thresholds, it auto-labels 0.2607 of test issues (only bug) at precision 0.8747. This misses the 0.90 precision target (see Gating).

Intended use

  • Suggesting or applying a type label (bug, feature, question) to new issues in English-language software repositories, through the laya-triage Action with the confidence gate and a human escalation path. The Action's default config sets mode: dry-run.
  • Research on calibrated, gated issue classification.

Not intended use

  • Fully automatic labelling without a confidence gate or human review.
  • Any decision about people (for example rating contributors or prioritising by author).
  • Issue types other than these three (documentation, security reports, and so on), or non-English issues.
  • Repositories very different from the 5 in the training data without first checking precision on a sample of that repository's own issues.

How to load

CPU, fp32, pinned revision. Download once, then run offline. The question below, including its wording and the option order bug, feature, question, is part of the model's contract: it was frozen before training and must be passed exactly as written. The state is the issue cleaned by laya-triage's preprocess() (see Training data).

import os
os.environ["HF_HUB_OFFLINE"] = "1"  # after one online download of this revision
os.environ.pop("LAYA_CPU_AMP", None)  # keep CPU inference in fp32

import laya

REPO = "Prasanna85/laya-issue-triage"
REVISION = "76ece1fb0eb8b32bd5d8c509293c1692a2534805"
ISSUE_TYPE_QUESTION = {
    "issue_type": {
        "type": "choice",
        "instructions": "You are triaging a newly opened GitHub issue. Using its `title` and `body`, decide which type of issue it is.",
        "criteria": {
            "bug": "reports a defect: a crash, error message, failing build or test, regression, or behaviour that contradicts the documentation",
            "feature": "proposes something new: a new capability, option or API, or an improvement to how existing behaviour works",
            "question": "asks for help: how to use or configure something, why it behaves a certain way, or troubleshooting the author's own setup"
        }
    }
}

agent = laya.load(REPO, revision=REVISION, device="cpu")  # max_len 1024, head_max_len 256
state = {"title": "...", "body": "..."}  # from laya_triage.preprocess.preprocess(title, body)
answer = agent.predict(state, ISSUE_TYPE_QUESTION)["answers"]["issue_type"]
answer["choice"], answer["answer_confidence"], answer["probabilities"]

Gate on answer_confidence (the calibrated probability of the chosen label), not on confidence.

With the laya-triage package

The laya-triage package wraps the same steps: it brings the frozen question and preprocess(), so nothing is retyped. Install it with pip install "laya_triage @ git+https://github.com/Prasanna-KS-85/laya-triage@v1.0.0" (on Linux, install the CPU-only torch wheel first; see the repository's docs/USING_THE_MODEL.md), then:

"""Classify one issue with the Laya Triage model on CPU (fp32). Install: see docs/USING_THE_MODEL.md."""
# The package brings the frozen question and the shared preprocess(): nothing is retyped.
from laya_triage.classifier import Classifier
from laya_triage.preprocess import preprocess

REPO = "Prasanna85/laya-issue-triage"
REVISION = "76ece1fb0eb8b32bd5d8c509293c1692a2534805"  # pinned: the revision the Action uses
MAX_LEN = 1024  # must equal the training budget
BUG_THRESHOLD = 0.6033  # the Action's default gate; feature and question are never auto-applied

model = Classifier(REPO, REVISION, MAX_LEN)  # downloads once, then loads from the local cache
state = preprocess("App closes when I open the export dialog",
                   "Steps: open a notebook, choose File > Export. The app exits with KeyError: 'last_format'.")
result = model.classify(state, issue_number=0)
print(result.probabilities)  # {'bug': ..., 'feature': ..., 'question': ...}
confident_bug = result.label == "bug" and result.answer_confidence >= BUG_THRESHOLD
print("apply bug" if confident_bug else "escalate to a human")

Training data

  • NLBSE'24 issue report classification: issues from bitcoin, react, vscode, opencv and tensorflow, labelled bug, feature or question.
  • Of the 1,500 official-train rows, 4 are dropped (their content also appears in the official test set). 300 rows are held out as validation (stratified by repo × label, seed 42), which leaves 1,196 training rows (bug 398, feature 398, question 400). The official test set (1,500 rows) is used only for the single final evaluation.
  • Preprocessing (shared by training, evaluation and the Action): strip HTML comments; collapse fenced code blocks longer than 20 lines to their first 10 and last 5 lines; replace images and bare URLs with placeholders; normalise whitespace; cap the body at 6,000 characters. The state is {"title": ..., "body": ...}, and Laya truncates it to 1024 tokens (11.1% of test states are truncated).
  • License and redistribution. The dataset's upstream LICENSE file is empty, so no license is granted. The issue text belongs to the source projects and their authors. No data is redistributed with this model or its repository; only synthetic IDs and content hashes are published.

Training procedure

Setting Value
Base convaiinnovations/laya @ 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851
Recipe Laya Kaggle notebook (RLCD policy-gradient term + soft cross-entropy on one-hot gold), label smoothing ε 0.0
Epochs / optimizer steps 4 / 68
Effective batch 64 (8 per micro-batch × 2 GPUs × 4 accumulation steps)
Sequence budget max_len 1024, head_max_len 256
Seed 42 (GPU kernels are not deterministic, so runs are seeded but not bit-for-bit reproducible)
Hardware / time Kaggle, 2× T4; train time 712 s
Versions (training) laya 0.3.23, transformers 5.18.0, Python 3.12, torch 2.10.0+cu128
Versions (evaluation) laya 0.3.23, torch 2.14.1, transformers 5.18.0, Python 3.11.17; CPU fp32
Calibration The notebook's choice temperature 4.0188 (fitted on 119 items held out from the training rows) was replaced by 2.6968, refit by NLL on the 300 validation issues with laya.calibrate. Weights unchanged.

Evaluation

Single run on the NLBSE'24 official test split (1,500 issues: 100 per class in each of 5 repositories), CPU fp32, after the thresholds were frozen. The headline metric is cross-repo macro-F1: the mean of the 5 per-repository macro-F1 scores, as in the NLBSE'24 competition.

System Cross-repo macro-F1 Pooled macro-F1 Accuracy F1 bug F1 feature F1 question ECE
M1 (this model) 0.8020 0.8023 0.8020 0.8033 0.8358 0.7677 0.0454
B1 TF-IDF + logistic regression 0.7655 0.7666 0.7673 0.7843 0.7925 0.7230 0.0493
B2 Laya base, zero-shot, 512/192 (native) 0.6108 0.6160 0.6453 0.6693 0.7757 0.4029 0.0684
B2 Laya base, zero-shot, 1024/256 0.6257 0.6310 0.6567 0.6773 0.7827 0.4330 0.0647
B3 SetFit (NLBSE'24, published) 0.8270 0.8263 [c] 0.8267 [c] 0.8425 [c] 0.8555 [c] 0.7809 [c] —
B3 RoBERTa (NLBSE'24, published) 0.7923 0.7926 [c] 0.7927 [c] 0.8052 [c] 0.8064 [c] 0.7663 [c] —
B3 fastText (NLBSE'24, published) 0.7184 0.7193 [c] 0.7193 [c] 0.7362 [c] 0.7390 [c] 0.6827 [c] —
  • [c] Derived, not published. NLBSE'24 publishes per-repository, per-class P/R/F1 only; pooled values were recovered exactly from them, because every test repository has 100 issues per class.
  • Protocol differences. B3 trains one classifier per repository, each on only that repository's 300 official-train rows (together the 5 classifiers use all 1,500 rows, including the 4 whose content also appears in test). The SetFit input is raw text; the RoBERTa and fastText setups were not inspected. B3 is copied rather than re-run. M1 and B1 are one model across all 5 repositories, trained on 1,196 rows with model selection, temperature and thresholds fitted on the 300-issue validation split. B2 is the base checkpoint without fine-tuning. B3 publishes no probabilities, so it has no ECE.
  • Per class (M1), precision / recall: bug 0.8326 / 0.7760; feature 0.8317 / 0.8400; question 0.7467 / 0.7900.

Per-repository macro-F1 (test):

System bitcoin react vscode opencv tensorflow Cross-repo
M1 0.7645 0.8445 0.7581 0.8002 0.8425 0.8020
B1 TF-IDF + logistic regression 0.7489 0.7903 0.7490 0.7383 0.8012 0.7655
B2 Laya base, zero-shot, 512/192 (native) 0.6213 0.7081 0.5053 0.5358 0.6833 0.6108
B2 Laya base, zero-shot, 1024/256 0.6298 0.7358 0.5106 0.5622 0.6901 0.6257

Calibration (test, 15 bins):

Probabilities ECE Brier
Calibrated (T = 2.6968, shipped) 0.0454 0.2890
Uncalibrated (T = 1) 0.1465 0.3301

Paired comparisons (test): difference in cross-repo macro-F1, with a 95% CI from a paired bootstrap (2,000 resamples of issues, seed 42), and the exact McNemar p-value.

Comparison Δ cross-repo macro-F1 95% CI McNemar p
M1 vs B1 TF-IDF + logistic regression +0.0364 [+0.0151, +0.0588] 0.0026
M1 vs B2 Laya base, zero-shot, 512/192 (native) +0.1912 [+0.1661, +0.2167] 4e-30
M1 vs B2 Laya base, zero-shot, 1024/256 +0.1763 [+0.1505, +0.2016] 1.3e-26

M1 beats B1 and both B2 settings on the headline metric, and all three CIs exclude 0. Relative to the published SetFit baseline (B3) it is -2.5 points (not a paired comparison; see the protocol differences).

Gating

Per-label thresholds on answer_confidence were fitted on validation before the test run: the smallest threshold whose bootstrap lower bound of precision (5th percentile of 2,000 resamples, at least 20 issues) is ≥ 0.90; 1.01 means the label is never auto-applied.

Label Threshold Val lower bound Applied on test Test coverage Test precision
bug 0.6033 0.9091 391 0.2607 0.8747
feature never (1.01) none ≥ 0.90 0 0.0000 —
question never (1.01) none ≥ 0.90 0 0.0000 —
all 391 of 1,500 0.2607 0.8747

The 0.90 precision target was missed on test. The bug threshold had validation precision 0.9481 (lower bound 0.9091, 77 issues, in-sample) but gave 0.8747 on test. feature and question never reached the lower bound on validation, so they are always escalated (project success criterion 3: met for 0 of 3 labels). For comparison, the point-estimate thresholds from validation would auto-label 0.8747 of test issues at precision 0.8491. Anyone using this model should verify precision on their own repository before switching the Action from dry-run to apply.

Limitations

  • Balanced data. Test has exactly 100 issues per class per repository; train is near-balanced (bug 398, feature 398, question 400). Real repositories have very different class mixes, so precision at a given threshold will differ.
  • Random, not temporal, split. The numbers do not measure drift to future issues.
  • 5 repositories. Large, well-known projects only; performance elsewhere is unknown.
  • Label noise, especially for question. Labels come from upstream repository labels. Some issues are plausibly mislabelled, and the author's own template choice often disagrees with a question label. question is the weakest class (test F1 0.7677).
  • Validation overstated test performance. Cross-repo macro-F1 was 0.8691 on the 300-issue validation split and 0.8020 on test.
  • Thresholds fitted on validation do not transfer exactly. The bug threshold lost 7.3 points of precision from validation to test.
  • English only. Trained and evaluated on English issues.
  • Exactly three classes (bug, feature, question); adding a class requires new labelled data, retraining and recalibrating.
  • Latency was measured on a local CPU, not on a GitHub runner (local Apple M4 Pro CPU, fp32, model preloaded): p50 178 ms and p95 436 ms per issue at 8 threads (test run); p50 288 ms and p95 690 ms at 2 threads (100 validation issues). Cold load: laya.load 3.14 s, peak RSS 2,863 MiB.
  • 11.1% of test issues are truncated at 1024 tokens; text beyond that is not seen.

Citation

Laya:

Convai Innovations / NandhaKishorM. Laya: non-autoregressive decision models, v0.3.23 (Apache-2.0). https://github.com/NandhaKishorM/laya . Checkpoints: https://huggingface.co/convaiinnovations/laya

NLBSE'24:

@inproceedings{nlbse2024,
  author={Kallis, Rafael and Colavito, Giuseppe and Al-Kaswan, Ali and Pascarella, Luca and Chaparro, Oscar and Rani, Pooja},
  title={The NLBSE'24 Tool Competition},
  booktitle={Proceedings of The 3rd International Workshop on Natural Language-based Software Engineering (NLBSE'24)},
  year={2024}
}
Downloads last month
-
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Prasanna85/laya-issue-triage

Finetuned
(144)
this model