Tanor commited on
Commit
8991938
·
verified ·
1 Parent(s): 6099840

Document Serbian WordNet sentiment classifier and research paper

Browse files

Expand model card with task, labels, original base model, paper citation, paired classifiers, inference example, archived results and provenance. Preserve historical training records and declared license. README-only update; no model artifacts or repository visibility changes.

Files changed (1) hide show
  1. README.md +181 -19
README.md CHANGED
@@ -1,39 +1,141 @@
1
  ---
2
- base_model: Tanor/SRGPTSENTNEG2
3
  tags:
4
- - generated_from_trainer
 
 
 
 
5
  metrics:
6
  - f1
7
  model-index:
8
  - name: SRGPTSENTNEG2
9
  results: []
 
 
 
 
 
 
 
 
10
  ---
11
 
12
- <!-- This model card has been generated automatically according to the information the Trainer had access to. You
13
- should probably proofread and complete it, then remove this comment. -->
14
-
15
  # SRGPTSENTNEG2
16
 
17
- This model is a fine-tuned version of [Tanor/SRGPTSENTNEG2](https://huggingface.co/Tanor/SRGPTSENTNEG2) on the None dataset.
18
- It achieves the following results on the evaluation set:
19
- - Loss: 0.1820
20
- - F1: 0.3235
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
21
 
22
- ## Model description
 
 
 
 
23
 
24
- More information needed
25
 
26
- ## Intended uses & limitations
 
 
 
 
 
27
 
28
- More information needed
29
 
30
- ## Training and evaluation data
31
 
32
- More information needed
 
 
33
 
34
- ## Training procedure
 
 
 
 
 
 
35
 
36
- ### Training hyperparameters
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
37
 
38
  The following hyperparameters were used during training:
39
  - learning_rate: 2e-05
@@ -46,7 +148,7 @@ The following hyperparameters were used during training:
46
  - lr_scheduler_type: linear
47
  - num_epochs: 32
48
 
49
- ### Training results
50
 
51
  | Training Loss | Epoch | Step | Validation Loss | F1 |
52
  |:-------------:|:-----:|:-----:|:---------------:|:------:|
@@ -57,9 +159,69 @@ The following hyperparameters were used during training:
57
  | 0.0403 | 5.0 | 13485 | 0.1820 | 0.3235 |
58
 
59
 
60
- ### Framework versions
61
 
62
  - Transformers 4.31.0
63
  - Pytorch 2.1.0.dev20230801
64
  - Datasets 2.14.2
65
  - Tokenizers 0.13.3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ base_model: jerteh/gpt2-orao
3
  tags:
4
+ - serbian
5
+ - sentiment-analysis
6
+ - wordnet
7
+ - sentiwordnet
8
+ - lexicon-induction
9
  metrics:
10
  - f1
11
  model-index:
12
  - name: SRGPTSENTNEG2
13
  results: []
14
+ language:
15
+ - sr
16
+ library_name: transformers
17
+ pipeline_tag: text-classification
18
+ base_model_relation: finetune
19
+ widget:
20
+ - text: koji oseća radost i zadovoljstvo
21
+ example_title: Illustrative Serbian gloss
22
  ---
23
 
 
 
 
24
  # SRGPTSENTNEG2
25
 
26
+ This is a **negative-polarity classifier for Serbian WordNet synset glosses**, fine-tuned from the GPT2-Orao model family. It is a component of the **S6 sentiment-lexicon construction method** described in:
27
+
28
+ Saša Petalinkar, Ranka M. Stanković, and Milica Ikonić Nešić (2025). **Comparative analysis of methods for creating a sentiment lexicon of the Serbian WordNet.** *The Electronic Library*, 43(4), 547–577. [Paper and DOI](https://doi.org/10.1108/EL-08-2024-0253).
29
+
30
+ [Companion code and sentiment lexicons](https://github.com/sasa5linkar/Serbian-WordNet-Sentiment-Lexicon-Analysis) · [Paired classifier](https://huggingface.co/Tanor/SRGPTSENTPOS2) · [Original base model](https://huggingface.co/jerteh/gpt2-orao)
31
+
32
+ ## Model identity and task
33
+
34
+ | Field | Value |
35
+ |---|---|
36
+ | Model ID | `Tanor/SRGPTSENTNEG2` |
37
+ | Architecture | `GPT2ForSequenceClassification` |
38
+ | Original base model | [`jerteh/gpt2-orao`](https://huggingface.co/jerteh/gpt2-orao) |
39
+ | Input | A Serbian synset gloss, representing one lexical meaning |
40
+ | Target | Negative vs non-negative polarity |
41
+ | Dataset expansion iteration | **2** (`T2`) |
42
+ | Label 0 | `NON-NEGATIVE` |
43
+ | Label 1 | `NEGATIVE` |
44
+ | Derived lexicon family | **S6** |
45
+ | Paired classifier, same iteration | [`Tanor/SRGPTSENTPOS2`](https://huggingface.co/Tanor/SRGPTSENTPOS2) |
46
+
47
+ The linked training script initializes the classifier from `jerteh/gpt2-orao`. `SRGPT` is the project naming convention for the GPT2-Orao sentiment classifiers.
48
+
49
+ A non-negative label is the complement of the target class. It does not by itself mean that the gloss has the opposite polarity. A separate classifier handles that polarity.
50
+
51
+ ## Training data and construction method
52
+
53
+ The paper constructs polarity-labeled synsets from Serbian WordNet, starting from curated positive, negative, and objective seed sets and expanding through semantic relations. The initial sets reported in the paper contain 149 positive, 219 negative, and 19,475 objective synsets. Polarity-preserving relations expand the corresponding set; antonymy contributes to the opposite polarity. The selected datasets are `T0`, `T2`, `T4`, and `T6`, after zero, two, four, and six expansion iterations.
54
+
55
+ This checkpoint is associated with **T2 and the NEG classification task**. Its numerical suffix is a dataset-expansion iteration, not an epoch count or a lexicon identifier.
56
+
57
+ The [training script](https://github.com/sasa5linkar/Serbian-WordNet-Sentiment-Lexicon-Analysis/blob/833582dbbf561a902fc5b872248db12bf529b3d9/trainSRGPT.py) reads `Sysnet` from `X_train_UPNEG2.csv` and the target `NEG` from `y_train_UPNEG2.csv`, replaces missing text with an empty string, and creates a stratified validation subset of **20% of that training CSV**, with `random_state=42`. The `UP` inputs are the non-lemmatized gloss variant in the [dataset-generation script](https://github.com/sasa5linkar/Serbian-WordNet-Sentiment-Lexicon-Analysis/blob/833582dbbf561a902fc5b872248db12bf529b3d9/create_sets.py).
58
+
59
+ The paper and code snapshot have different split descriptions: the paper describes an 80/10/10 split for neural models, while the scripts split existing training CSV files and `create_sets.py` leaves the initial split size at the library default. Exact checkpoint-specific sample assignments are not supplied in the public revision. The `train_sets/` files referenced by the scripts are absent from that revision. Reconstructing the experiment requires the relevant lexical resources and saved preparation/split information; the paper's proportions alone do not establish this checkpoint's split.
60
+
61
+ ## Use in sentiment lexicon S6
62
+
63
+ For each iteration, the POS model estimates positive-class probability `p_pos` and the NEG model estimates negative-class probability `p_neg`. The [lexicon-building code](https://github.com/sasa5linkar/Serbian-WordNet-Sentiment-Lexicon-Analysis/blob/833582dbbf561a902fc5b872248db12bf529b3d9/sentiwordnet_calculator.py) combines them as:
64
 
65
+ ```text
66
+ POS = p_pos * (1 - p_neg)
67
+ NEG = p_neg * (1 - p_pos)
68
+ OBJ = 1 - POS - NEG
69
+ ```
70
 
71
+ It averages each score across the four iteration-specific pairs. A single checkpoint is one component of this construction; its two class probabilities are not the final three lexicon scores.
72
 
73
+ | Iteration | POS classifier | NEG classifier |
74
+ |---|---|---|
75
+ | 0 | [SRGPTSENTPOS0](https://huggingface.co/Tanor/SRGPTSENTPOS0) | [SRGPTSENTNEG0](https://huggingface.co/Tanor/SRGPTSENTNEG0) |
76
+ | 2 | [SRGPTSENTPOS2](https://huggingface.co/Tanor/SRGPTSENTPOS2) | [SRGPTSENTNEG2](https://huggingface.co/Tanor/SRGPTSENTNEG2) |
77
+ | 4 | [SRGPTSENTPOS4](https://huggingface.co/Tanor/SRGPTSENTPOS4) | [SRGPTSENTNEG4](https://huggingface.co/Tanor/SRGPTSENTNEG4) |
78
+ | 6 | [SRGPTSENTPOS6](https://huggingface.co/Tanor/SRGPTSENTPOS6) | [SRGPTSENTNEG6](https://huggingface.co/Tanor/SRGPTSENTNEG6) |
79
 
80
+ ## Usage
81
 
82
+ This example reads one gloss and returns both class probabilities. It pins the model to the weights revision present before the documentation update.
83
 
84
+ ```python
85
+ import torch
86
+ from transformers import AutoModelForSequenceClassification, AutoTokenizer
87
 
88
+ MODEL_ID = "Tanor/SRGPTSENTNEG2"
89
+ WEIGHTS_REVISION = "609984070541e85cc2f51e0bafdfeeb42653c25c"
90
+ tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, revision=WEIGHTS_REVISION)
91
+ model = AutoModelForSequenceClassification.from_pretrained(
92
+ MODEL_ID, revision=WEIGHTS_REVISION
93
+ )
94
+ model.eval()
95
 
96
+ # An illustrative gloss, not a benchmark item or a claimed prediction.
97
+ inputs = tokenizer(
98
+ "koji oseća radost i zadovoljstvo",
99
+ return_tensors="pt", truncation=True, max_length=300,
100
+ )
101
+ with torch.inference_mode():
102
+ probabilities = model(**inputs).logits.softmax(dim=-1)[0]
103
+ scores = {model.config.id2label[i]: float(p) for i, p in enumerate(probabilities)}
104
+ print(scores)
105
+ ```
106
+
107
+ The model uses standard PyTorch/Transformers sequence-classification classes, without custom remote code. The example requires PyTorch and Transformers. Where a Trainer record is available below, it includes software versions from the original run. No example prediction or benchmark score was generated for this documentation update.
108
+
109
+ ## Evaluation and provenance
110
+
111
+ ### Archived test report
112
+
113
+ The companion repository contains an [evaluation report for SRGPT, NEG, T2](https://github.com/sasa5linkar/Serbian-WordNet-Sentiment-Lexicon-Analysis/blob/833582dbbf561a902fc5b872248db12bf529b3d9/reports/SRGPT/report_UPNEG2.csv.txt). It is preserved below, with class 0 = `NON-NEGATIVE` and class 1 = `NEGATIVE`. Metric values retain the source's rounding; support values are sample counts. The confusion matrix uses that class order.
114
+
115
+ ```text
116
+ [[4404 16]
117
+ [ 59 17]]
118
+
119
+ precision recall f1-score support
120
+
121
+ 0 0.99 1.00 0.99 4420
122
+ 1 0.52 0.22 0.31 76
123
+
124
+ accuracy 0.98 4496
125
+ macro avg 0.75 0.61 0.65 4496
126
+ weighted avg 0.98 0.98 0.98 4496
127
+ ```
128
+
129
+ This is an archived experiment artifact, not a new evaluation. The report does not record a model-weight SHA, so its exact correspondence to the currently hosted weights has not been independently re-established. These binary-classification metrics are separate from evaluation of the derived lexicon and from sentiment evaluation of full sentences.
130
+
131
+ ### Preserved Trainer record
132
+
133
+ The following validation summary and training log are retained from the [previous model card](https://huggingface.co/Tanor/SRGPTSENTNEG2/blob/609984070541e85cc2f51e0bafdfeeb42653c25c/README.md), without recomputing them. The recorded F1 is a training-validation metric, separate from the archived test report above and from lexicon or sentence-level sentiment evaluation. The source training code uses binary F1 for target label 1 when `eval="f1"` is selected.
134
+
135
+ - Loss: 0.1820
136
+ - F1: 0.3235
137
+
138
+ #### Training hyperparameters
139
 
140
  The following hyperparameters were used during training:
141
  - learning_rate: 2e-05
 
148
  - lr_scheduler_type: linear
149
  - num_epochs: 32
150
 
151
+ #### Training results
152
 
153
  | Training Loss | Epoch | Step | Validation Loss | F1 |
154
  |:-------------:|:-----:|:-----:|:---------------:|:------:|
 
159
  | 0.0403 | 5.0 | 13485 | 0.1820 | 0.3235 |
160
 
161
 
162
+ #### Framework versions
163
 
164
  - Transformers 4.31.0
165
  - Pytorch 2.1.0.dev20230801
166
  - Datasets 2.14.2
167
  - Tokenizers 0.13.3
168
+
169
+ The Trainer record and the linked source script are separate provenance sources. In particular, the recorded optimizer may differ from the script's `adafactor` setting; the historical record is retained without asserting that the linked script reproduces that exact run.
170
+
171
+ ## Settings in the archived training source
172
+
173
+ These settings describe the linked code revision, not a replacement for the per-run Trainer record.
174
+
175
+ | Setting | Source-code value |
176
+ |---|---|
177
+ | Maximum tokenized input length | 300 |
178
+ | Learning rate | 2e-5 |
179
+ | Training batch size per device | 1 |
180
+ | Evaluation batch size per device | 1 |
181
+ | Gradient accumulation | 4 steps |
182
+ | Optimizer | Adafactor |
183
+ | Weight decay | 0.01 |
184
+ | Validation split seed | 42 |
185
+ | Evaluation and saving | Each epoch |
186
+ | Early stopping patience | 3 evaluation calls |
187
+ | Epoch budget in experiment notebooks | Up to 32 |
188
+
189
+ The paper reports early-stopping patience of 3. The linked family script also uses 3.
190
+
191
+ ## Intended use and limitations
192
+
193
+ - Intended for research on polarity of Serbian WordNet meanings and construction or analysis of Serbian sentiment lexicons.
194
+ - Training inputs are glosses. Performance on reviews, news, social media, and documents requires separate evaluation; this checkpoint is not documented as a general sentence-sentiment benchmark model.
195
+ - Class imbalance is visible in the archived report. Read accuracy and weighted averages alongside target-class precision, recall, F1, and support.
196
+ - Labels depend on seed selection and semantic-relation expansion. Polysemy, domain-specific polarity, and propagation errors can affect predictions. English-to-Serbian alignment alone does not establish polarity in Serbian.
197
+ - Softmax outputs are model scores; probability calibration is not established by this card.
198
+ - The model performs no sense selection for a word in context. Applying the lexicon to text requires a separate choice or aggregation of meanings.
199
+
200
+ ## License
201
+
202
+ No license is declared in this model repository's model-card metadata. A new license has not been inferred from the code repository or the base model. Contact the model authors to clarify reuse terms; code and external lexical resources have separate terms.
203
+
204
+ ## Citation
205
+
206
+ When using this model family or the resulting lexicon-construction method, cite the paper and record the model ID and revision used.
207
+
208
+ ```bibtex
209
+ @article{petalinkar2025sentimentlexicon,
210
+ author = {Petalinkar, Saša and Stanković, Ranka M. and Ikonić Nešić, Milica},
211
+ title = {Comparative analysis of methods for creating a sentiment lexicon of the Serbian WordNet},
212
+ journal = {The Electronic Library},
213
+ year = {2025},
214
+ volume = {43},
215
+ number = {4},
216
+ pages = {547--577},
217
+ doi = {10.1108/EL-08-2024-0253},
218
+ url = {https://doi.org/10.1108/EL-08-2024-0253}
219
+ }
220
+ ```
221
+
222
+ ## Version and documentation sources
223
+
224
+ - Weights/configuration revision documented here: [`609984070541e85cc2f51e0bafdfeeb42653c25c`](https://huggingface.co/Tanor/SRGPTSENTNEG2/tree/609984070541e85cc2f51e0bafdfeeb42653c25c). This pins the hosted checkpoint; it does not prove which weights produced every table in the paper.
225
+ - Companion code revision: [`833582dbbf561a902fc5b872248db12bf529b3d9`](https://github.com/sasa5linkar/Serbian-WordNet-Sentiment-Lexicon-Analysis/tree/833582dbbf561a902fc5b872248db12bf529b3d9).
226
+ - BERTić initialization is recorded in [Prepare and load models.ipynb](https://github.com/sasa5linkar/Serbian-WordNet-Sentiment-Lexicon-Analysis/blob/833582dbbf561a902fc5b872248db12bf529b3d9/Prepare%20and%20load%20models.ipynb); GPT2-Orao and Jerteh-355 initialization is recorded in their training scripts.
227
+ - Documentation expanded on 23 September 2026 using the paper, model configuration, repository source, archived evaluation report, and any pre-existing Trainer record. Weights, tokenizer files, and model configuration were not changed by this documentation update.