Token Classification
SpanMarker
TensorBoard
Safetensors
Hebrew
ner
named-entity-recognition
generated_from_span_marker_trainer
Eval Results (legacy)
Instructions to use iahlt/span-marker-xlm-roberta-base-nemo-mt-he with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- SpanMarker
How to use iahlt/span-marker-xlm-roberta-base-nemo-mt-he with SpanMarker:
from span_marker import SpanMarkerModel model = SpanMarkerModel.from_pretrained("iahlt/span-marker-xlm-roberta-base-nemo-mt-he") entities = model.predict("Amelia Earhart flew her single engine Lockheed Vega 5B across the Atlantic to Paris.") print(entities) - Notebooks
- Google Colab
- Kaggle
|
Download README.md from iahlt/span-marker-xlm-roberta-base-nemo-mt-he: direct link, hf CLI and curl.
- Browser
- Download file 7.58 kB
-
https://huggingface.co/iahlt/span-marker-xlm-roberta-base-nemo-mt-he/resolve/a12a6aeed278255cd3eeea7510e3d9bc7fe3103d/README.md
- Command line
-
hf download hf://iahlt/span-marker-xlm-roberta-base-nemo-mt-he@a12a6aeed278255cd3eeea7510e3d9bc7fe3103d/README.md
-
curl -L -o README.md https://huggingface.co/iahlt/span-marker-xlm-roberta-base-nemo-mt-he/resolve/a12a6aeed278255cd3eeea7510e3d9bc7fe3103d/README.md
7.58 kB
metadata
library_name: span-marker
tags:
- span-marker
- token-classification
- ner
- named-entity-recognition
- generated_from_span_marker_trainer
datasets:
- imvladikon/nemo_corpus
metrics:
- precision
- recall
- f1
widget:
- text: >-
אלי ויזל, פרופסור ב אוניברסיטת בוסטון, ש סילבר התאמץ הרבה למען זכייתו ב
פרס נובל ל שלום, תמך בגלוי ב מועמדותו ל משרת ה מושל.
- text: >-
מאמרו של תום שגב, " ה קרב על סן סימון היה או לא היה " (" ה ארץ " 105),
הגיע ל ידי רק ב ימים אלה.
- text: >-
רק ב דבריו של ה רב אברהם טולדאנו, משגיח ב ישיבת ה רעיון ה יהודי ו מספר 4 ב
רשימת כך ל ה כנסת, היו כבר הוראות מעשיות: " אלוקים ייקום דמו ו אנו ניקום
את הוא.
- text: >-
מרכז ה מידע ל זכויות ה אדם ב ה שטחים, " בצלם ", מפרסם מ פעם ל פעם דפי מידע
ו ב המ פרטים על ה נעשה ב ה שטחים ב תחומים שונים.
- text: >-
גרוסבורד נהג לבדו ב ה מכונית, ב דרכו מ ה עיר מיניאפוליס ב אינדיאנה ל נמל ה
תעופה של היא.
pipeline_tag: token-classification
model-index:
- name: SpanMarker
results:
- task:
type: token-classification
name: Named Entity Recognition
dataset:
name: Unknown
type: imvladikon/nemo_corpus
split: test
metrics:
- type: f1
value: 0.7757111597374179
name: F1
- type: precision
value: 0.7912946428571429
name: Precision
- type: recall
value: 0.7607296137339056
name: Recall
SpanMarker
This is a SpanMarker model trained on the imvladikon/nemo_corpus dataset that can be used for Named Entity Recognition.
Model Details
Model Description
- Model Type: SpanMarker
- Maximum Sequence Length: 512 tokens
- Maximum Entity Length: 100 words
- Training Dataset: imvladikon/nemo_corpus
Model Sources
- Repository: SpanMarker on GitHub
- Thesis: SpanMarker For Named Entity Recognition
Model Labels
| Label | Examples |
|---|---|
| ANG | "יידיש", "אנגלית", "גרמנית" |
| DUC | "סובארו", "מרצדס", "דינמיט" |
| EVE | "מצדה", "הצהרת בלפור", "ה שואה" |
| FAC | "ברזילי", "תל - ה שומר", "כלא עזה" |
| GPE | "שפרעם", "רצועת עזה", "ה שטחים" |
| LOC | "חאן יונס", "גיבאליה", "שייח רדואן" |
| ORG | "ה ארץ", "מרחב ה גליל", "כך" |
| PER | "נימר חוסיין", "איברהים נימר חוסיין", "רמי רהב" |
| WOA | "ה ארץ", "קדיש", "קיטש ו מוות" |
Evaluation
Metrics
| Label | Precision | Recall | F1 |
|---|---|---|---|
| all | 0.7913 | 0.7607 | 0.7757 |
| ANG | 0.0 | 0.0 | 0.0 |
| DUC | 0.0 | 0.0 | 0.0 |
| FAC | 0.3571 | 0.4545 | 0.4 |
| GPE | 0.7817 | 0.7897 | 0.7857 |
| LOC | 0.5263 | 0.4878 | 0.5063 |
| ORG | 0.7854 | 0.7623 | 0.7736 |
| PER | 0.8725 | 0.8202 | 0.8456 |
| WOA | 0.0 | 0.0 | 0.0 |
Uses
Direct Use for Inference
from span_marker import SpanMarkerModel
# Download from the 🤗 Hub
model = SpanMarkerModel.from_pretrained("span_marker_model_id")
# Run inference
entities = model.predict("גרוסבורד נהג לבדו ב ה מכונית, ב דרכו מ ה עיר מיניאפוליס ב אינדיאנה ל נמל ה תעופה של היא.")
Downstream Use
You can finetune this model on your own dataset.
Click to expand
from span_marker import SpanMarkerModel, Trainer
# Download from the 🤗 Hub
model = SpanMarkerModel.from_pretrained("span_marker_model_id")
# Specify a Dataset with "tokens" and "ner_tag" columns
dataset = load_dataset("conll2003") # For example CoNLL2003
# Initialize a Trainer using the pretrained model & dataset
trainer = Trainer(
model=model,
train_dataset=dataset["train"],
eval_dataset=dataset["validation"],
)
trainer.train()
trainer.save_model("span_marker_model_id-finetuned")
Training Details
Training Set Metrics
| Training set | Min | Median | Max |
|---|---|---|---|
| Sentence length | 0 | 25.7252 | 117 |
| Entities per sentence | 0 | 1.2722 | 20 |
Training Hyperparameters
- learning_rate: 1e-05
- train_batch_size: 2
- eval_batch_size: 2
- seed: 42
- gradient_accumulation_steps: 2
- total_train_batch_size: 4
- optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
- lr_scheduler_type: linear
- lr_scheduler_warmup_ratio: 0.1
- num_epochs: 2
- mixed_precision_training: Native AMP
Training Results
| Epoch | Step | Validation Loss | Validation Precision | Validation Recall | Validation F1 | Validation Accuracy |
|---|---|---|---|---|---|---|
| 0.4393 | 1000 | 0.0083 | 0.7632 | 0.5812 | 0.6598 | 0.9477 |
| 0.8785 | 2000 | 0.0056 | 0.8366 | 0.6774 | 0.7486 | 0.9609 |
| 1.3178 | 3000 | 0.0052 | 0.8322 | 0.7655 | 0.7975 | 0.9714 |
| 1.7571 | 4000 | 0.0053 | 0.8008 | 0.7735 | 0.7870 | 0.9712 |
Framework Versions
- Python: 3.10.12
- SpanMarker: 1.5.0
- Transformers: 4.35.2
- PyTorch: 2.1.0+cu118
- Datasets: 2.15.0
- Tokenizers: 0.15.0
Citation
BibTeX
@software{Aarsen_SpanMarker,
author = {Aarsen, Tom},
license = {Apache-2.0},
title = {{SpanMarker for Named Entity Recognition}},
url = {https://github.com/tomaarsen/SpanMarkerNER}
}