Text Classification
Transformers
Safetensors
xlm-roberta
naics
industry-classification
github
bge-m3
text-embeddings-inference
Instructions to use aquiro1994/naics-github-classifier-multilingual with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use aquiro1994/naics-github-classifier-multilingual with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="aquiro1994/naics-github-classifier-multilingual")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("aquiro1994/naics-github-classifier-multilingual") model = AutoModelForSequenceClassification.from_pretrained("aquiro1994/naics-github-classifier-multilingual", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Correct the cross-lingual figures: they came from a different checkpoint
Browse files
README.md
CHANGED
|
@@ -46,7 +46,7 @@ training data, same recipe; the difference is the encoder underneath.
|
|
| 46 |
| Test accuracy | **86.95%** | 86.72% |
|
| 47 |
| Weighted F1 | **86.39%** | 86.33% |
|
| 48 |
| Macro F1 | 83.06% | 82.95% |
|
| 49 |
-
| Same sector for a README and its own English translation | **
|
| 50 |
| Calibrated as trained | yes, `T` = 1.04 | yes, `T` = 1.01 |
|
| 51 |
|
| 52 |
On English text the two are indistinguishable: a 0.2-point difference on a
|
|
@@ -63,22 +63,24 @@ brings is a shared multilingual representation space from its own pre-training,
|
|
| 63 |
so a Spanish or Chinese README lands near its English equivalent and the
|
| 64 |
classification head, trained on English, still applies.
|
| 65 |
|
| 66 |
-
That transfer is real but imperfect, and it is what the
|
| 67 |
non-English repositories were classified twice, once as written and once from a
|
| 68 |
-
human-quality English translation, and the two answers agreed
|
| 69 |
RoBERTa-large agrees with itself 19% of the time on the same test, which is what
|
| 70 |
having no vocabulary for the text looks like.
|
| 71 |
|
| 72 |
| | FR (31) | ES (28) | RU (14) | PT (11) | ZH (9) | all (93) |
|
| 73 |
|---|---:|---:|---:|---:|---:|---:|
|
| 74 |
-
| this model | 84% |
|
| 75 |
| RoBERTa-large | 16% | 11% | 21% | 18% | 56% | 19% |
|
| 76 |
|
| 77 |
-
|
| 78 |
-
readings differ, neither is known to be right.
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
|
|
|
|
|
|
| 82 |
|
| 83 |
Training on multilingual labelled data would be the real upgrade, and nothing
|
| 84 |
here measures what it would add.
|
|
@@ -202,8 +204,9 @@ split or sequence length.
|
|
| 202 |
## Limitations
|
| 203 |
|
| 204 |
- **The training data is English.** See the section above.
|
| 205 |
-
- **Agreement across languages is not accuracy**,
|
| 206 |
-
(
|
|
|
|
| 207 |
- **The labels are GPT-4.1 judgements**, not human annotation.
|
| 208 |
- **The model cannot abstain.** The training data contains only positives, so
|
| 209 |
every repository is assigned some sector. On a production corpus of public
|
|
|
|
| 46 |
| Test accuracy | **86.95%** | 86.72% |
|
| 47 |
| Weighted F1 | **86.39%** | 86.33% |
|
| 48 |
| Macro F1 | 83.06% | 82.95% |
|
| 49 |
+
| Same sector for a README and its own English translation | **77%** | 19% |
|
| 50 |
| Calibrated as trained | yes, `T` = 1.04 | yes, `T` = 1.01 |
|
| 51 |
|
| 52 |
On English text the two are indistinguishable: a 0.2-point difference on a
|
|
|
|
| 63 |
so a Spanish or Chinese README lands near its English equivalent and the
|
| 64 |
classification head, trained on English, still applies.
|
| 65 |
|
| 66 |
+
That transfer is real but imperfect, and it is what the 77% above measures: 93
|
| 67 |
non-English repositories were classified twice, once as written and once from a
|
| 68 |
+
human-quality English translation, and the two answers agreed 77% of the time.
|
| 69 |
RoBERTa-large agrees with itself 19% of the time on the same test, which is what
|
| 70 |
having no vocabulary for the text looks like.
|
| 71 |
|
| 72 |
| | FR (31) | ES (28) | RU (14) | PT (11) | ZH (9) | all (93) |
|
| 73 |
|---|---:|---:|---:|---:|---:|---:|
|
| 74 |
+
| this model | 84% | 68% | 100% | 73% | 56% | **77%** |
|
| 75 |
| RoBERTa-large | 16% | 11% | 21% | 18% | 56% | 19% |
|
| 76 |
|
| 77 |
+
Three caveats worth stating plainly. **Agreement is not accuracy**: when the two
|
| 78 |
+
readings differ, neither is known to be right. **The per-language cells are tiny**
|
| 79 |
+
— 9 to 31 repositories each, so a single flipped repository moves Chinese by 11
|
| 80 |
+
points and the Russian 100% rests on 14 cases. Only the 93-repository total is
|
| 81 |
+
worth quoting. And the figure varies between training runs more than the English
|
| 82 |
+
metrics do — a second run of the same recipe measured 81%, against an English F1
|
| 83 |
+
of 85.2. Treat 77% as one measurement, not a constant.
|
| 84 |
|
| 85 |
Training on multilingual labelled data would be the real upgrade, and nothing
|
| 86 |
here measures what it would add.
|
|
|
|
| 204 |
## Limitations
|
| 205 |
|
| 206 |
- **The training data is English.** See the section above.
|
| 207 |
+
- **Agreement across languages is not accuracy**, it rests on 93 repositories,
|
| 208 |
+
and it varies between runs (77% here, 81% in a second run of the same recipe).
|
| 209 |
+
The per-language breakdown has 9 to 31 repositories per cell.
|
| 210 |
- **The labels are GPT-4.1 judgements**, not human annotation.
|
| 211 |
- **The model cannot abstain.** The training data contains only positives, so
|
| 212 |
every repository is assigned some sector. On a production corpus of public
|