Token Classification
Transformers
PyTorch
English
bert
fill-mask
bert-base-cased
biodiversity
sequence-classification
Instructions to use NoYo25/BiodivBERT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use NoYo25/BiodivBERT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="NoYo25/BiodivBERT")# Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("NoYo25/BiodivBERT") model = AutoModelForMaskedLM.from_pretrained("NoYo25/BiodivBERT", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Add pre-training hyperparams
Browse files
README.md
CHANGED
|
@@ -36,6 +36,13 @@ training_data:
|
|
| 36 |
- corpora:
|
| 37 |
- (+Abs) Springer and Elsevier abstracts in the duration of 1990-2020
|
| 38 |
- (+Abs+Full) Springer and Elsevier abstracts and open access full publication text in the duration of 1990-2020
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
---
|
| 40 |
|
| 41 |
# BiodivBERT
|
|
|
|
| 36 |
- corpora:
|
| 37 |
- (+Abs) Springer and Elsevier abstracts in the duration of 1990-2020
|
| 38 |
- (+Abs+Full) Springer and Elsevier abstracts and open access full publication text in the duration of 1990-2020
|
| 39 |
+
pre-training-hyperparams:
|
| 40 |
+
- MAX_LEN = 512 # Default of BERT Tokenizer
|
| 41 |
+
- MLM_PROP = 0.15 # Data Collator
|
| 42 |
+
- num_train_epochs = 3 # the minimum sufficient epochs found on many articles && default of trainer here
|
| 43 |
+
- per_device_train_batch_size = 16 # the maximumn that could be held by V100 on Ara with 512 MAX_LEN was 8 in the old run
|
| 44 |
+
- per_device_eval_batch_size = 16 # usually as above
|
| 45 |
+
- gradient_accumulation_steps = 4 # this will grant a minim batch size 16 * 4 * nGPUs.
|
| 46 |
---
|
| 47 |
|
| 48 |
# BiodivBERT
|