NightRaven commited on
Commit
ecd72e0
·
verified ·
1 Parent(s): 327306b

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +0 -64
README.md CHANGED
@@ -95,64 +95,6 @@ model.eval()
95
 
96
  Inputs are tokenised at a maximum sequence length of 100.
97
 
98
- ## Reported results
99
-
100
- From the paper, on the test split as it stood at publication:
101
-
102
- | Model | Precision (%) | Recall (%) | F1 (%) |
103
- |---|---:|---:|---:|
104
- | **XLM-RoBERTa** | **95.22** | **95.38** | **95.25** |
105
- | BanglaBERT | 94.73 | 94.67 | 94.63 |
106
- | mBERT | 93.77 | 93.66 | 93.59 |
107
- | Logistic Regression | 86.31 | 85.02 | 84.04 |
108
- | K-Nearest Neighbors | 83.80 | 82.83 | 82.01 |
109
- | Multinomial Naive Bayes | 74.17 | 66.07 | 60.76 |
110
- | BiLSTM + CNN | 62.42 | 64.65 | 60.37 |
111
- | LSTM | 20.24 | 44.99 | 27.92 |
112
-
113
- Best results for XLM-RoBERTa came from batch size 16 and learning rate 1e-5;
114
- mBERT and BanglaBERT used batch size 16 and learning rate 3e-5.
115
-
116
- ## Limitations
117
-
118
- Please read these before relying on any of the non-transformer models.
119
-
120
- **The published metrics were measured on an earlier version of the dataset.**
121
- That version contained 5,629 posts and included only 13 `blood` examples — 7 in
122
- train, 1 in validation, 5 in test. The `blood` column of the published confusion
123
- matrix therefore rests on five test samples. The dataset has since been corrected
124
- to the full 220 `blood` posts described above. These checkpoints were trained
125
- before that correction, so their real-world `blood` performance is far weaker
126
- than the aggregate scores suggest, and the table above does not describe
127
- performance on the current data.
128
-
129
- **The DNN and scikit-learn models are not self-contained.** They consume token
130
- ids from a Keras `Tokenizer` and TF-IDF features fitted at import time on the
131
- training CSV, rather than a saved vocabulary. Their embedding indices and feature
132
- columns are only meaningful against that exact fitted vocabulary. Reconstructing
133
- it from corrected data shifts the indices and the predictions become unreliable
134
- while still looking confident. Treat these six as historical baselines.
135
-
136
- **The scikit-learn models were pickled under scikit-learn 1.0.2.** Loading them
137
- on a later version raises `InconsistentVersionWarning`, which scikit-learn notes
138
- may lead to invalid results.
139
-
140
- **The Multinomial Naive Bayes model predicts string labels** (`'crime'`), while
141
- the logistic regression and KNN models predict integer ids. Callers must handle
142
- both.
143
-
144
- **Loading pickles executes arbitrary code.** The `.pkl` files and `torch.load`
145
- both unpickle. Only load them if you trust this repository.
146
-
147
- **Class imbalance and annotation noise.** `crime` dominates at 42.7%. The source
148
- spreadsheet also contains a small number of duplicate posts, four of which carry
149
- conflicting labels (for example the same fire report labelled both `fire` and
150
- `natural_disaster`) — a reminder that some categories genuinely overlap.
151
-
152
- **Scope.** Trained on Bangla posts about events in Bangladesh and neighbouring
153
- India. Behaviour on other varieties of Bangla, or on non-emergency text, is
154
- untested. This is a research artifact and is not fit to be a sole trigger for
155
- any real emergency response.
156
 
157
  ## Citation
158
 
@@ -168,9 +110,3 @@ any real emergency response.
168
  year = {2023}
169
  }
170
  ```
171
-
172
- ## Acknowledgements
173
-
174
- Department of Computer Science and Engineering, Khulna University of Engineering
175
- & Technology (KUET). The training code is derived from
176
- [xashru/bangla-text-classification](https://github.com/xashru/bangla-text-classification) (MIT).
 
95
 
96
  Inputs are tokenised at a maximum sequence length of 100.
97
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
98
 
99
  ## Citation
100
 
 
110
  year = {2023}
111
  }
112
  ```