--- language: - bn license: mit tags: - text-classification - bangla - bengali - emergency - social-media - bert - xlm-roberta pipeline_tag: text-classification --- # Bangla Emergency Post Classification Models that classify Bangla social media posts into nine emergency categories, so that posts needing attention from emergency services, local authorities or law enforcement can be surfaced automatically. Trained for *Bangla Emergency Post Classification on Social Media using Transformer Based BERT Models*, 6th International Conference on Electrical Information and Communication Technology (EICT), Khulna, Bangladesh, December 2023. ## Labels Nine classes, encoded in alphabetical order: | id | label | id | label | id | label | |----|-------|----|-------|----|-------| | 0 | accident | 3 | fire | 6 | suicide | | 1 | blood | 4 | natural_disaster | 7 | war | | 2 | crime | 5 | pandemic | 8 | weather | `blood` covers urgent blood-donation appeals, which are a common and distinct category of emergency post in Bangla social media. ## Data 5,836 Bangla posts collected from Facebook, Twitter and daily newspapers, labelled by hand by native speakers. Posts frequently mix English words into Bangla script. | split | posts | |---|---:| | train | 3,267 | | validation | 819 | | test | 1,750 | The classes are heavily imbalanced — `crime` is 42.7% of the data and `pandemic` 2.5%: | label | count | | label | count | |---|---:|---|---|---:| | crime | 2,492 | | suicide | 208 | | accident | 1,003 | | war | 175 | | weather | 598 | | pandemic | 147 | | fire | 499 | | | | | natural_disaster | 494 | | **total** | **5,836** | | blood | 220 | | | | ## What is in this repository ``` Transformer/ BanglaBERT/social-media_weights.pt sagorsarker/bangla-bert-base XLM-RoBERTa/social-media_weights.pt xlm-roberta-base mBERT/social-media_weights.pt bert-base-multilingual-uncased dnn/ BiLSTM.model/ BiLSTM_CNN.model/ LSTM.model/ TensorFlow SavedModel machine_learning/ logistic/ knn/ mnb/ scikit-learn, joblib ``` ## Usage The transformer checkpoints are `state_dict`s for the wrapper classes in the project's `modelss/` package, **not** plain Hugging Face models — loading them with `AutoModelForSequenceClassification.from_pretrained` will not work. Use the code from the project repository: ```python from huggingface_hub import hf_hub_download import torch from modelss.text_xlm_roberta import TextRoBERTa model = TextRoBERTa(pretrained_model='xlm-roberta-base', num_class=9, fine_tune=True) weights = hf_hub_download( repo_id="NightRaven/bangla-emergency-post-classification", filename="Transformer/XLM-RoBERTa/social-media_weights.pt", ) model.load_state_dict(torch.load(weights, map_location='cpu')) model.eval() ``` Inputs are tokenised at a maximum sequence length of 100. ## Citation ```bibtex @inproceedings{nabil2023bangla, title = {Bangla Emergency Post Classification on Social Media using Transformer Based BERT Models}, author = {Nabil, Alvi Ahmmed and Arifeen, Shamsul and Das, Dola and Salim, Md. Shahidul and Fattah, H. M. Abdul}, booktitle = {6th International Conference on Electrical Information and Communication Technology (EICT)}, address = {Khulna, Bangladesh}, year = {2023} } ```