anggars commited on
Commit
7647126
·
verified ·
1 Parent(s): a5a2b3d

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +28 -23
README.md CHANGED
@@ -11,6 +11,7 @@ tags:
11
  - hybrid-corpus
12
  metrics:
13
  - accuracy
 
14
  datasets:
15
  - anggars/mbti-emotion
16
  model-index:
@@ -24,8 +25,23 @@ model-index:
24
  type: anggars/mbti-emotion
25
  metrics:
26
  - type: accuracy
27
- value: 0.8796
28
  name: Accuracy
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
29
  ---
30
 
31
  # XLM-RoBERTa Emotion (Domain-Adapted for Midwest Emo/Math Rock)
@@ -36,27 +52,25 @@ This model is a fine-tuned version of [xlm-roberta-base](https://huggingface.co/
36
 
37
  - **Model Type:** XLM-RoBERTa Base (Sequence Classification Head with 28 Nodes)
38
  - **Labels:** 28 Emotion Categories (e.g., sadness, grief, admiration, anger, joy)
39
- - **Dataset:** `anggars/mbti-emotion` (Hybrid Corpus: 112,351 synthetic narrative rows + 129 organic scraped lyrics)
40
  - **Language:** English & Indonesian (Multilingual)
41
- - **Training Environment:** Kaggle Compute (Dual Tesla T4 GPU, fp16 Mixed Precision)
42
 
43
  ## Architectural Innovations: Overcoming Domain Shift
44
 
45
- Initial iterations of this model were trained purely on synthetic narrative data, which caused severe **Domain Shift** when predicting real-world music lyrics (e.g., misclassifying depressive metaphors like *"drowning"* or *"heal this soul"* as *Admiration*).
46
 
47
- To mitigate this blind spot, a **Hybrid Corpus Integration** was executed. The model was forced to adapt to organic lyrics scraped directly from Genius.com. The integration process successfully recalibrated the latent space, forcing the model to understand poetic contexts and lyrical structures.
48
-
49
- *Note: The slight reduction in absolute accuracy (from previous baselines) and the increased validation loss are expected phenomena known as **Strategic Accuracy Drop** and **Softmax Calibration**. The aggressive weight decay (0.05) ensures the model does not overconfidently hallucinate on ambiguous lyrics, resulting in highly generalized, real-world zero-shot capabilities.*
50
 
51
  ## Training Results
52
 
53
- The following results were achieved on the evaluation set during the 3-epoch training process:
54
 
55
- | Epoch | Step | Validation Loss | Accuracy |
56
- |:-----:|:-----:|:---------------:|:--------:|
57
- | 1.0 | 5624 | 1.0018 | 0.8422 |
58
- | 2.0 | 11248 | 0.8367 | 0.8660 |
59
- | 3.0 | 16872 | 0.7716 | 0.8796 |
60
 
61
  ## Intended Uses & Limitations
62
 
@@ -67,8 +81,6 @@ This model is explicitly designed for the backend NLP engine of music analytics
67
 
68
  ### Training Hyperparameters
69
 
70
- To ensure stable convergence on the complex organic lyrics and prevent overfitting on the synthetic data, the following hyperparameters were utilized:
71
-
72
  - **learning_rate:** 1.5e-05
73
  - **train_batch_size:** 16
74
  - **eval_batch_size:** 16
@@ -77,11 +89,4 @@ To ensure stable convergence on the complex organic lyrics and prevent overfitti
77
  - **optimizer:** AdamW with betas=(0.9,0.999) and epsilon=1e-08
78
  - **lr_scheduler_type:** linear
79
  - **num_epochs:** 3
80
- - **mixed_precision_training:** Native AMP (fp16)
81
-
82
- ### Framework Versions
83
-
84
- - Transformers 4.44.2
85
- - Pytorch 2.5.1+cu124
86
- - Datasets 3.1.0
87
- - Tokenizers 0.20.3
 
11
  - hybrid-corpus
12
  metrics:
13
  - accuracy
14
+ - f1
15
  datasets:
16
  - anggars/mbti-emotion
17
  model-index:
 
25
  type: anggars/mbti-emotion
26
  metrics:
27
  - type: accuracy
28
+ value: 0.9546
29
  name: Accuracy
30
+ - type: f1
31
+ value: 0.9544
32
+ name: F1 Macro
33
+ language:
34
+ - en
35
+ - id
36
+ - ja
37
+ - ko
38
+ - es
39
+ - fr
40
+ - de
41
+ - zh
42
+ - pt
43
+ - ru
44
+ - it
45
  ---
46
 
47
  # XLM-RoBERTa Emotion (Domain-Adapted for Midwest Emo/Math Rock)
 
52
 
53
  - **Model Type:** XLM-RoBERTa Base (Sequence Classification Head with 28 Nodes)
54
  - **Labels:** 28 Emotion Categories (e.g., sadness, grief, admiration, anger, joy)
55
+ - **Dataset:** `anggars/mbti-emotion` (Hybrid Corpus: 120,060 total rows. Stratified split: 96,048 train / 24,012 eval)
56
  - **Language:** English & Indonesian (Multilingual)
57
+ - **Training Environment:** Kaggle Compute (Dual NVIDIA Tesla T4 GPU, fp16 Mixed Precision)
58
 
59
  ## Architectural Innovations: Overcoming Domain Shift
60
 
61
+ Initial iterations of this model were trained purely on synthetic narrative data, which caused severe **Domain Shift** when predicting real-world music lyrics. To mitigate this blind spot, a **Hybrid Corpus Integration** was executed. The model was forced to adapt to organic lyrics scraped directly from Genius.com and augmented with high-quality, balanced synthetic data generated via Gemma-2B-IT.
62
 
63
+ The integration process successfully recalibrated the latent space, forcing the model to understand poetic contexts and lyrical structures. The aggressive weight decay (0.05) ensures the model does not overconfidently hallucinate on ambiguous lyrics, resulting in highly generalized, real-world zero-shot capabilities.
 
 
64
 
65
  ## Training Results
66
 
67
+ The following results were achieved on the evaluation set (24,012 rows) during the 3-epoch training process:
68
 
69
+ | Epoch | Training Loss | Validation Loss | Accuracy | F1 Macro |
70
+ |:-----:|:-------------:|:---------------:|:--------:|:--------:|
71
+ | 1.0 | 0.2403 | 0.2086 | 0.9347 | 0.9339 |
72
+ | 2.0 | 0.1403 | 0.1837 | 0.9480 | 0.9475 |
73
+ | 3.0 | 0.1005 | 0.1888 | 0.9546 | 0.9544 |
74
 
75
  ## Intended Uses & Limitations
76
 
 
81
 
82
  ### Training Hyperparameters
83
 
 
 
84
  - **learning_rate:** 1.5e-05
85
  - **train_batch_size:** 16
86
  - **eval_batch_size:** 16
 
89
  - **optimizer:** AdamW with betas=(0.9,0.999) and epsilon=1e-08
90
  - **lr_scheduler_type:** linear
91
  - **num_epochs:** 3
92
+ - **mixed_precision_training:** Native AMP (fp16)