Text Classification
Transformers
Safetensors
xlm-roberta
naics
industry-classification
github
bge-m3
text-embeddings-inference
Instructions to use aquiro1994/naics-github-classifier-multilingual with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use aquiro1994/naics-github-classifier-multilingual with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="aquiro1994/naics-github-classifier-multilingual")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("aquiro1994/naics-github-classifier-multilingual") model = AutoModelForSequenceClassification.from_pretrained("aquiro1994/naics-github-classifier-multilingual", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Correct the RoBERTa temperature and the training runtime
Browse files
README.md
CHANGED
|
@@ -47,13 +47,14 @@ training data, same recipe; the difference is the encoder underneath.
|
|
| 47 |
| Weighted F1 | **86.39%** | 86.33% |
|
| 48 |
| Macro F1 | 83.06% | 82.95% |
|
| 49 |
| Same sector for a README and its own English translation | **77%** | 19% |
|
| 50 |
-
| Calibrated as trained | yes, `T` = 1.04 | yes, `T` = 1.
|
| 51 |
|
| 52 |
On English text the two are indistinguishable: a 0.2-point difference on a
|
| 53 |
1,318-row test set is well inside run-to-run noise. **Use this one when the
|
| 54 |
corpus is not all English**, which for public GitHub is about 12% of
|
| 55 |
-
repositories
|
| 56 |
-
|
|
|
|
| 57 |
|
| 58 |
## Important: the training data is English
|
| 59 |
|
|
@@ -192,10 +193,11 @@ python scripts/train.py --model bge-m3 \
|
|
| 192 |
| Learning rate | 1.5e-5, polynomial decay, 15% warm-up |
|
| 193 |
| Weight decay | 0.02 |
|
| 194 |
| Optimizer | AdamW |
|
| 195 |
-
| Hardware | Apple M5 Max, fp32,
|
| 196 |
|
| 197 |
-
A 1,024-token
|
| 198 |
-
|
|
|
|
| 199 |
|
| 200 |
The settings above are recorded in `training_config.json` in this repository, and
|
| 201 |
`scripts/evaluate.py` reads them so evaluation cannot silently use a different
|
|
|
|
| 47 |
| Weighted F1 | **86.39%** | 86.33% |
|
| 48 |
| Macro F1 | 83.06% | 82.95% |
|
| 49 |
| Same sector for a README and its own English translation | **77%** | 19% |
|
| 50 |
+
| Calibrated as trained | yes, `T` = 1.04 | yes, `T` = 1.10 |
|
| 51 |
|
| 52 |
On English text the two are indistinguishable: a 0.2-point difference on a
|
| 53 |
1,318-row test set is well inside run-to-run noise. **Use this one when the
|
| 54 |
corpus is not all English**, which for public GitHub is about 12% of
|
| 55 |
+
repositories: on a sample of 821,929, that is the share whose language is
|
| 56 |
+
confidently detected as something other than English. Use the RoBERTa model
|
| 57 |
+
when the corpus is English and you want the smaller, faster network.
|
| 58 |
|
| 59 |
## Important: the training data is English
|
| 60 |
|
|
|
|
| 193 |
| Learning rate | 1.5e-5, polynomial decay, 15% warm-up |
|
| 194 |
| Weight decay | 0.02 |
|
| 195 |
| Optimizer | AdamW |
|
| 196 |
+
| Hardware | Apple M5 Max, fp32, 151 minutes |
|
| 197 |
|
| 198 |
+
A 1,024-token window was tried under an earlier recipe and scored 0.5 points
|
| 199 |
+
lower than its 512-token counterpart: the signal is in the opening of the
|
| 200 |
+
README, not in its length.
|
| 201 |
|
| 202 |
The settings above are recorded in `training_config.json` in this repository, and
|
| 203 |
`scripts/evaluate.py` reads them so evaluation cannot silently use a different
|