Prasanna85 commited on
Commit
da82718
·
verified ·
1 Parent(s): bbfde5c

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +35 -6
README.md CHANGED
@@ -18,9 +18,10 @@ tags:
18
  ## Summary
19
 
20
  A fine-tuned [Laya](https://huggingface.co/convaiinnovations/laya) decision model that classifies a newly
21
- opened GitHub issue as one of `bug`, `feature`, `question` from its title and body. It is the model behind the laya-triage
22
- GitHub Action, which auto-applies a label only when `answer_confidence` reaches a per-label threshold and
23
- otherwise escalates the issue to a human.
 
24
 
25
  - Revision `76ece1fb0eb8b32bd5d8c509293c1692a2534805`: weights of training run R1 (`a704b3eadd185f1fa028576cfa50605b6eada4a4`) with the choice
26
  temperature refit on validation data (4.0188 → 2.6968); nothing else differs.
@@ -79,6 +80,32 @@ answer["choice"], answer["answer_confidence"], answer["probabilities"]
79
 
80
  Gate on `answer_confidence` (the calibrated probability of the chosen label), not on `confidence`.
81
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
82
  ## Training data
83
 
84
  - [NLBSE'24 issue report classification](https://github.com/nlbse2024/issue-report-classification): issues
@@ -129,9 +156,10 @@ per-repository macro-F1 scores, as in the NLBSE'24 competition.
129
 
130
  - [c] Derived, not published. NLBSE'24 publishes per-repository, per-class P/R/F1 only; pooled values were
131
  recovered exactly from them, because every test repository has 100 issues per class.
132
- - **Protocol differences.** B3 trains one classifier per repository on that repository's official-train rows
133
- (all of them, including the 4 that overlap test), uses raw text, and is copied rather than re-run.
134
- M1 and B1 are one model across all 5 repositories, trained on 1,196 rows with model selection,
 
135
  temperature and thresholds fitted on the 300-issue validation split. B2 is the base checkpoint
136
  without fine-tuning. B3 publishes no probabilities, so it has no ECE.
137
  - Per class (M1), precision / recall: bug 0.8326 / 0.7760; feature 0.8317 / 0.8400; question 0.7467 / 0.7900.
@@ -199,6 +227,7 @@ switching the Action from dry-run to apply.
199
  - **Thresholds fitted on validation do not transfer exactly.** The `bug` threshold lost
200
  7.3 points of precision from validation to test.
201
  - **English only.** Trained and evaluated on English issues.
 
202
  - **Latency was measured on a local CPU, not on a GitHub runner** (local Apple M4 Pro CPU, fp32, model preloaded):
203
  p50 178 ms and p95 436 ms per issue at 8 threads
204
  (test run); p50 288 ms and p95 690 ms at 2 threads (100 validation issues).
 
18
  ## Summary
19
 
20
  A fine-tuned [Laya](https://huggingface.co/convaiinnovations/laya) decision model that classifies a newly
21
+ opened GitHub issue as one of `bug`, `feature`, `question` from its title and body. It is the model behind the
22
+ [laya-triage](https://github.com/Prasanna-KS-85/laya-triage) GitHub Action, which auto-applies a label only when `answer_confidence` reaches a
23
+ per-label threshold and otherwise escalates the issue to a human. Source code, install instructions and the
24
+ full evaluation write-up: https://github.com/Prasanna-KS-85/laya-triage.
25
 
26
  - Revision `76ece1fb0eb8b32bd5d8c509293c1692a2534805`: weights of training run R1 (`a704b3eadd185f1fa028576cfa50605b6eada4a4`) with the choice
27
  temperature refit on validation data (4.0188 → 2.6968); nothing else differs.
 
80
 
81
  Gate on `answer_confidence` (the calibrated probability of the chosen label), not on `confidence`.
82
 
83
+ ### With the laya-triage package
84
+
85
+ The [laya-triage](https://github.com/Prasanna-KS-85/laya-triage) package wraps the same steps: it brings the frozen question and `preprocess()`, so
86
+ nothing is retyped. Install it with `pip install "laya_triage @ git+https://github.com/Prasanna-KS-85/laya-triage@v1.0.0"` (on Linux, install
87
+ the CPU-only torch wheel first; see the repository's docs/USING_THE_MODEL.md), then:
88
+
89
+ ```python
90
+ """Classify one issue with the Laya Triage model on CPU (fp32). Install: see docs/USING_THE_MODEL.md."""
91
+ # The package brings the frozen question and the shared preprocess(): nothing is retyped.
92
+ from laya_triage.classifier import Classifier
93
+ from laya_triage.preprocess import preprocess
94
+
95
+ REPO = "Prasanna85/laya-issue-triage"
96
+ REVISION = "76ece1fb0eb8b32bd5d8c509293c1692a2534805" # pinned: the revision the Action uses
97
+ MAX_LEN = 1024 # must equal the training budget
98
+ BUG_THRESHOLD = 0.6033 # the Action's default gate; feature and question are never auto-applied
99
+
100
+ model = Classifier(REPO, REVISION, MAX_LEN) # downloads once, then loads from the local cache
101
+ state = preprocess("App closes when I open the export dialog",
102
+ "Steps: open a notebook, choose File > Export. The app exits with KeyError: 'last_format'.")
103
+ result = model.classify(state, issue_number=0)
104
+ print(result.probabilities) # {'bug': ..., 'feature': ..., 'question': ...}
105
+ confident_bug = result.label == "bug" and result.answer_confidence >= BUG_THRESHOLD
106
+ print("apply bug" if confident_bug else "escalate to a human")
107
+ ```
108
+
109
  ## Training data
110
 
111
  - [NLBSE'24 issue report classification](https://github.com/nlbse2024/issue-report-classification): issues
 
156
 
157
  - [c] Derived, not published. NLBSE'24 publishes per-repository, per-class P/R/F1 only; pooled values were
158
  recovered exactly from them, because every test repository has 100 issues per class.
159
+ - **Protocol differences.** B3 trains one classifier per repository, each on only that repository's
160
+ 300 official-train rows (together the 5 classifiers use all 1,500 rows, including the
161
+ 4 whose content also appears in test). The SetFit input is raw text; the RoBERTa and fastText
162
+ setups were not inspected. B3 is copied rather than re-run. M1 and B1 are one model across all 5 repositories, trained on 1,196 rows with model selection,
163
  temperature and thresholds fitted on the 300-issue validation split. B2 is the base checkpoint
164
  without fine-tuning. B3 publishes no probabilities, so it has no ECE.
165
  - Per class (M1), precision / recall: bug 0.8326 / 0.7760; feature 0.8317 / 0.8400; question 0.7467 / 0.7900.
 
227
  - **Thresholds fitted on validation do not transfer exactly.** The `bug` threshold lost
228
  7.3 points of precision from validation to test.
229
  - **English only.** Trained and evaluated on English issues.
230
+ - Exactly three classes (bug, feature, question); adding a class requires new labelled data, retraining and recalibrating.
231
  - **Latency was measured on a local CPU, not on a GitHub runner** (local Apple M4 Pro CPU, fp32, model preloaded):
232
  p50 178 ms and p95 436 ms per issue at 8 threads
233
  (test run); p50 288 ms and p95 690 ms at 2 threads (100 validation issues).