aquiro1994 commited on
Commit
87a94ee
·
verified ·
1 Parent(s): d0b38c1

Correct the cross-lingual figures: they came from a different checkpoint

Browse files
Files changed (1) hide show
  1. README.md +14 -11
README.md CHANGED
@@ -46,7 +46,7 @@ training data, same recipe; the difference is the encoder underneath.
46
  | Test accuracy | **86.95%** | 86.72% |
47
  | Weighted F1 | **86.39%** | 86.33% |
48
  | Macro F1 | 83.06% | 82.95% |
49
- | Same sector for a README and its own English translation | **73%** | 19% |
50
  | Calibrated as trained | yes, `T` = 1.04 | yes, `T` = 1.01 |
51
 
52
  On English text the two are indistinguishable: a 0.2-point difference on a
@@ -63,22 +63,24 @@ brings is a shared multilingual representation space from its own pre-training,
63
  so a Spanish or Chinese README lands near its English equivalent and the
64
  classification head, trained on English, still applies.
65
 
66
- That transfer is real but imperfect, and it is what the 73% above measures: 93
67
  non-English repositories were classified twice, once as written and once from a
68
- human-quality English translation, and the two answers agreed 73% of the time.
69
  RoBERTa-large agrees with itself 19% of the time on the same test, which is what
70
  having no vocabulary for the text looks like.
71
 
72
  | | FR (31) | ES (28) | RU (14) | PT (11) | ZH (9) | all (93) |
73
  |---|---:|---:|---:|---:|---:|---:|
74
- | this model | 84% | 57% | 79% | 91% | 56% | **73%** |
75
  | RoBERTa-large | 16% | 11% | 21% | 18% | 56% | 19% |
76
 
77
- Two caveats worth stating plainly. **Agreement is not accuracy**: when the two
78
- readings differ, neither is known to be right. And the figure varies between
79
- training runs more than the English metrics do — a second run of the same recipe
80
- measured 81%, against an English F1 of 85.2. Treat 73% as one measurement, not a
81
- constant.
 
 
82
 
83
  Training on multilingual labelled data would be the real upgrade, and nothing
84
  here measures what it would add.
@@ -202,8 +204,9 @@ split or sequence length.
202
  ## Limitations
203
 
204
  - **The training data is English.** See the section above.
205
- - **Agreement across languages is not accuracy**, and it varies between runs
206
- (73% here, 81% in a second run of the same recipe).
 
207
  - **The labels are GPT-4.1 judgements**, not human annotation.
208
  - **The model cannot abstain.** The training data contains only positives, so
209
  every repository is assigned some sector. On a production corpus of public
 
46
  | Test accuracy | **86.95%** | 86.72% |
47
  | Weighted F1 | **86.39%** | 86.33% |
48
  | Macro F1 | 83.06% | 82.95% |
49
+ | Same sector for a README and its own English translation | **77%** | 19% |
50
  | Calibrated as trained | yes, `T` = 1.04 | yes, `T` = 1.01 |
51
 
52
  On English text the two are indistinguishable: a 0.2-point difference on a
 
63
  so a Spanish or Chinese README lands near its English equivalent and the
64
  classification head, trained on English, still applies.
65
 
66
+ That transfer is real but imperfect, and it is what the 77% above measures: 93
67
  non-English repositories were classified twice, once as written and once from a
68
+ human-quality English translation, and the two answers agreed 77% of the time.
69
  RoBERTa-large agrees with itself 19% of the time on the same test, which is what
70
  having no vocabulary for the text looks like.
71
 
72
  | | FR (31) | ES (28) | RU (14) | PT (11) | ZH (9) | all (93) |
73
  |---|---:|---:|---:|---:|---:|---:|
74
+ | this model | 84% | 68% | 100% | 73% | 56% | **77%** |
75
  | RoBERTa-large | 16% | 11% | 21% | 18% | 56% | 19% |
76
 
77
+ Three caveats worth stating plainly. **Agreement is not accuracy**: when the two
78
+ readings differ, neither is known to be right. **The per-language cells are tiny**
79
+ — 9 to 31 repositories each, so a single flipped repository moves Chinese by 11
80
+ points and the Russian 100% rests on 14 cases. Only the 93-repository total is
81
+ worth quoting. And the figure varies between training runs more than the English
82
+ metrics do — a second run of the same recipe measured 81%, against an English F1
83
+ of 85.2. Treat 77% as one measurement, not a constant.
84
 
85
  Training on multilingual labelled data would be the real upgrade, and nothing
86
  here measures what it would add.
 
204
  ## Limitations
205
 
206
  - **The training data is English.** See the section above.
207
+ - **Agreement across languages is not accuracy**, it rests on 93 repositories,
208
+ and it varies between runs (77% here, 81% in a second run of the same recipe).
209
+ The per-language breakdown has 9 to 31 repositories per cell.
210
  - **The labels are GPT-4.1 judgements**, not human annotation.
211
  - **The model cannot abstain.** The training data contains only positives, so
212
  every repository is assigned some sector. On a production corpus of public