Ar4ikov commited on
Commit
309cbe2
·
verified ·
1 Parent(s): ebf9ba6

Add the crosslingual release with its quantized builds

Browse files
README.md ADDED
@@ -0,0 +1,182 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - ru
4
+ - en
5
+ - de
6
+ - pl
7
+ - it
8
+ - fr
9
+ - es
10
+ - bn
11
+ license: mit
12
+ library_name: transformers
13
+ pipeline_tag: audio-classification
14
+ base_model: Aniemore/unispeech-sat-emotion-russian-resd
15
+ base_model_relation: finetune
16
+ datasets:
17
+ - Aniemore/resd
18
+ - Aniemore/resd_annotated
19
+ - amu-cai/CAMEO
20
+ tags:
21
+ - audio-classification
22
+ - emotion-recognition
23
+ - speech-emotion-recognition
24
+ - speech
25
+ - multilingual
26
+ - russian
27
+ - quantized
28
+ - compressed-tensors
29
+ - int8
30
+ - fp8
31
+ - int4
32
+ metrics:
33
+ - f1
34
+ - accuracy
35
+ - recall
36
+ model-index:
37
+ - name: unispeech-sat-emotion-v1-crosslingual
38
+ results:
39
+ - task:
40
+ name: Speech Emotion Recognition
41
+ type: audio-classification
42
+ dataset:
43
+ name: RESD test
44
+ type: Aniemore/resd
45
+ metrics:
46
+ - name: Macro F1
47
+ type: f1
48
+ value: 0.5697
49
+ - name: Unweighted accuracy
50
+ type: recall
51
+ value: 0.6158
52
+ - task:
53
+ name: Speech Emotion Recognition
54
+ type: audio-classification
55
+ dataset:
56
+ name: Dusha podcast test
57
+ type: dusha
58
+ metrics:
59
+ - name: Macro F1
60
+ type: f1
61
+ value: 0.3705
62
+ - name: Unweighted accuracy
63
+ type: recall
64
+ value: 0.6992
65
+ - task:
66
+ name: Speech Emotion Recognition
67
+ type: audio-classification
68
+ dataset:
69
+ name: CAMEO test
70
+ type: amu-cai/CAMEO
71
+ metrics:
72
+ - name: Macro F1
73
+ type: f1
74
+ value: 0.5689
75
+ - name: Unweighted accuracy
76
+ type: recall
77
+ value: 0.5706
78
+ ---
79
+
80
+ <img src="assets/banner.svg" alt="unispeech-sat-emotion-v1-crosslingual" width="100%">
81
+
82
+ # unispeech-sat-emotion-v1-crosslingual
83
+
84
+ Speech emotion recognition over seven classes &mdash; `anger`, `disgust`, `enthusiasm`, `fear`, `happiness`, `neutral`, `sadness`.
85
+
86
+ Same architecture as [`Aniemore/unispeech-sat-emotion-russian-resd`](https://huggingface.co/Aniemore/unispeech-sat-emotion-russian-resd), retrained on a mix of 27,939 clips spanning eight languages and three speaking registers instead of one acted Russian corpus. The point of the change is spontaneous speech: the previous release was trained only on acted dialogue, where every class is equally frequent and every utterance is performed, and real speech is neither.
87
+
88
+ Quantized builds ship in the same repository under `int8/`, `fp8/` and `int4/`.
89
+
90
+ ## Results
91
+
92
+ | test set | what it is | macro-F1 | UA | WA | previous release |
93
+ |---|---|---:|---:|---:|---:|
94
+ | RESD test | acted Russian, 7 balanced classes | **0.5697** | 0.6158 | 0.6179 | 0.7027 |
95
+ | Dusha podcast test | spontaneous Russian, majority neutral | **0.3705** | 0.6992 | 0.6993 | 0.1000 |
96
+ | CAMEO test | 7 non-Russian languages | **0.5689** | 0.5706 | 0.6832 | 0.1946 |
97
+
98
+ <img src="assets/panel.svg" alt="Panel results" width="760">
99
+
100
+ Read the first two rows together. The acted score goes down and the spontaneous score goes up; both follow from the same change, and which one matters is a deployment question. If your audio is read or performed speech, the previous release may still suit you better.
101
+
102
+ <details><summary>About the CAMEO row</summary>
103
+
104
+ CAMEO ships no train/test partition, and the usual way to make one &mdash; a random split over clips &mdash; puts nearly every test speaker into training as well: three of its twelve constituent corpora contain a single speaker each, so no clip-level split of them can be speaker-disjoint even in principle. The number above is reported for completeness. Treat it as an in-domain figure, not as evidence of cross-lingual transfer.
105
+
106
+ </details>
107
+
108
+ ### Per-class recall on spontaneous speech
109
+
110
+ | class | this model | previous release |
111
+ |---|---:|---:|
112
+ | `angry` | 0.5689 | 0.5629 |
113
+ | `neutral` | 0.6958 | 0.0583 |
114
+ | `positive` | 0.8138 | 0.5029 |
115
+ | `sad` | 0.7184 | 0.2718 |
116
+
117
+ `neutral` carries most of real speech and is the class the previous release missed.
118
+
119
+ ## Quantized variants
120
+
121
+ | subfolder | scheme | weights | vs fp32 | macro-F1 | UA | WA |
122
+ |---|---|---:|---:|---:|---:|---:|
123
+ | _(root)_ | fp32 | 1206 MiB | 1.0x | 0.5697 | 0.6158 | 0.6179 |
124
+ | `int8` | W8A16 | 351 MiB | 3.4x smaller | 0.5697 | 0.6158 | 0.6179 |
125
+ | `fp8` | W8A16-float | 343 MiB | 3.5x smaller | 0.5663 | 0.6126 | 0.6143 |
126
+ | `int4` | W4A16_ASYM | 209 MiB | 5.8x smaller | 0.5714 | 0.6163 | 0.6179 |
127
+
128
+ <img src="assets/quality.svg" alt="Quality after quantization" width="760">
129
+
130
+ <img src="assets/size.svg" alt="Weights on disk" width="760">
131
+
132
+ Weight-only, round-to-nearest, no calibration. Every variant lands within the seed spread of the fp32 parent on RESD test, so the choice is about download size rather than about quality.
133
+
134
+ ## Training data
135
+
136
+ | corpus | clips | language | register |
137
+ |---|---:|---|---|
138
+ | RESD | 948 | Russian | acted dialogue |
139
+ | Dusha crowd | 6,800 | Russian | acted, crowd-sourced |
140
+ | CAMEO | 6,800 | 7 languages | 12 corpora, no Russian |
141
+ | Dusha podcast | 6,060 | Russian | spontaneous podcast speech |
142
+ | IEMOCAP | 4,735 | English | elicited dyadic sessions |
143
+ | ASVP-ESD | 2,596 | multilingual | mixed register |
144
+ | **total** | **27,939** | 8 languages | 3 registers |
145
+
146
+ A slice is held out of every corpus in the mix, in the same proportion, and model selection is on that held-out split &mdash; never on any of the test sets above. Labels are unified to seven classes; four-class corpora are mapped upward and scored on the classes they actually contain.
147
+
148
+ ## Usage
149
+
150
+ ```python
151
+ import torch, librosa
152
+ from transformers import AutoModelForAudioClassification, AutoFeatureExtractor
153
+
154
+ repo = "Aniemore/unispeech-sat-emotion-v1-crosslingual"
155
+ model = AutoModelForAudioClassification.from_pretrained(repo).eval()
156
+ fe = AutoFeatureExtractor.from_pretrained(repo)
157
+
158
+ # Resample to 16 kHz. Do not skip it: RESD itself ships at 44.1 kHz,
159
+ # and handing the model 44.1 kHz audio while telling the extractor it
160
+ # is 16 kHz stretches time 2.8x and silently changes the answer.
161
+ wav, _ = librosa.load("clip.wav", sr=16000, mono=True)
162
+ x = fe(wav, sampling_rate=16000, return_tensors="pt", padding=True)
163
+ with torch.no_grad():
164
+ probs = model(**x).logits.softmax(-1)[0]
165
+ print({model.config.id2label[i]: round(p.item(), 3) for i, p in enumerate(probs)})
166
+ ```
167
+
168
+ For a quantized build, name the subfolder &mdash; only that subfolder is downloaded:
169
+
170
+ ```python
171
+ model = AutoModelForAudioClassification.from_pretrained(
172
+ repo, subfolder="int8").eval() # or "fp8", "int4"
173
+ fe = AutoFeatureExtractor.from_pretrained(repo, subfolder="int8")
174
+ ```
175
+
176
+ ## Limitations
177
+
178
+ - Scores are the mean of two seeds; the seed spread on the 280-clip RESD split is ±0.03&ndash;0.05, so differences smaller than that are not differences.
179
+ - The spontaneous set has four classes where the model has seven, so its numbers are computed over a mapped label space and are not comparable to seven-class figures.
180
+ - Spontaneous scores are at zero decision bias. Calibrating the `neutral` threshold on your own development split will move them.
181
+ - Inherited from [`microsoft/unispeech-sat-large`](https://huggingface.co/microsoft/unispeech-sat-large); the licence follows the base model.
182
+
assets/banner.svg ADDED
assets/panel.svg ADDED
assets/quality.svg ADDED
assets/size.svg ADDED
config.json ADDED
@@ -0,0 +1,126 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "activation_dropout": 0.05,
3
+ "apply_spec_augment": true,
4
+ "architectures": [
5
+ "UniSpeechSatForSequenceClassification"
6
+ ],
7
+ "attention_dropout": 0.05,
8
+ "bos_token_id": 1,
9
+ "classifier_proj_size": 768,
10
+ "codevector_dim": 768,
11
+ "contrastive_logits_temperature": 0.1,
12
+ "conv_bias": false,
13
+ "conv_dim": [
14
+ 512,
15
+ 512,
16
+ 512,
17
+ 512,
18
+ 512,
19
+ 512,
20
+ 512
21
+ ],
22
+ "conv_kernel": [
23
+ 10,
24
+ 3,
25
+ 3,
26
+ 3,
27
+ 3,
28
+ 2,
29
+ 2
30
+ ],
31
+ "conv_stride": [
32
+ 5,
33
+ 2,
34
+ 2,
35
+ 2,
36
+ 2,
37
+ 2,
38
+ 2
39
+ ],
40
+ "ctc_loss_reduction": "sum",
41
+ "ctc_zero_infinity": false,
42
+ "diversity_loss_weight": 0.1,
43
+ "do_stable_layer_norm": true,
44
+ "dtype": "float32",
45
+ "eos_token_id": 2,
46
+ "feat_extract_activation": "gelu",
47
+ "feat_extract_dropout": 0.0,
48
+ "feat_extract_norm": "layer",
49
+ "feat_proj_dropout": 0.05,
50
+ "feat_quantizer_dropout": 0.0,
51
+ "final_dropout": 0.05,
52
+ "finetuning_task": [
53
+ "unispeech_sat_classification"
54
+ ],
55
+ "hidden_act": "gelu",
56
+ "hidden_dropout": 0.05,
57
+ "hidden_size": 1024,
58
+ "id2label": {
59
+ "0": "anger",
60
+ "1": "disgust",
61
+ "2": "enthusiasm",
62
+ "3": "fear",
63
+ "4": "happiness",
64
+ "5": "neutral",
65
+ "6": "sadness"
66
+ },
67
+ "initializer_range": 0.02,
68
+ "intermediate_size": 4096,
69
+ "label2id": {
70
+ "anger": 0,
71
+ "disgust": 1,
72
+ "enthusiasm": 2,
73
+ "fear": 3,
74
+ "happiness": 4,
75
+ "neutral": 5,
76
+ "sadness": 6
77
+ },
78
+ "layer_norm_eps": 1e-05,
79
+ "layerdrop": 0.05,
80
+ "mask_feature_length": 10,
81
+ "mask_feature_min_masks": 0,
82
+ "mask_feature_prob": 0.0,
83
+ "mask_time_length": 10,
84
+ "mask_time_min_masks": 2,
85
+ "mask_time_prob": 0.05,
86
+ "model_type": "unispeech-sat",
87
+ "num_attention_heads": 16,
88
+ "num_clusters": 504,
89
+ "num_codevector_groups": 2,
90
+ "num_codevectors_per_group": 320,
91
+ "num_conv_pos_embedding_groups": 16,
92
+ "num_conv_pos_embeddings": 128,
93
+ "num_feat_extract_layers": 7,
94
+ "num_hidden_layers": 24,
95
+ "num_negatives": 100,
96
+ "pad_token_id": 0,
97
+ "pooling_mode": "mean",
98
+ "problem_type": "single_label_classification",
99
+ "proj_codevector_dim": 768,
100
+ "tdnn_dilation": [
101
+ 1,
102
+ 2,
103
+ 3,
104
+ 1,
105
+ 1
106
+ ],
107
+ "tdnn_dim": [
108
+ 512,
109
+ 512,
110
+ 512,
111
+ 512,
112
+ 1500
113
+ ],
114
+ "tdnn_kernel": [
115
+ 5,
116
+ 3,
117
+ 3,
118
+ 1,
119
+ 1
120
+ ],
121
+ "tokenizer_class": "Wav2Vec2CTCTokenizer",
122
+ "transformers_version": "5.14.1",
123
+ "use_weighted_layer_sum": false,
124
+ "vocab_size": 40,
125
+ "xvector_output_dim": 512
126
+ }
fp8/config.json ADDED
@@ -0,0 +1,165 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "activation_dropout": 0.05,
3
+ "apply_spec_augment": true,
4
+ "architectures": [
5
+ "UniSpeechSatForSequenceClassification"
6
+ ],
7
+ "attention_dropout": 0.05,
8
+ "bos_token_id": 1,
9
+ "classifier_proj_size": 768,
10
+ "codevector_dim": 768,
11
+ "contrastive_logits_temperature": 0.1,
12
+ "conv_bias": false,
13
+ "conv_dim": [
14
+ 512,
15
+ 512,
16
+ 512,
17
+ 512,
18
+ 512,
19
+ 512,
20
+ 512
21
+ ],
22
+ "conv_kernel": [
23
+ 10,
24
+ 3,
25
+ 3,
26
+ 3,
27
+ 3,
28
+ 2,
29
+ 2
30
+ ],
31
+ "conv_stride": [
32
+ 5,
33
+ 2,
34
+ 2,
35
+ 2,
36
+ 2,
37
+ 2,
38
+ 2
39
+ ],
40
+ "ctc_loss_reduction": "sum",
41
+ "ctc_zero_infinity": false,
42
+ "diversity_loss_weight": 0.1,
43
+ "do_stable_layer_norm": true,
44
+ "dtype": "float32",
45
+ "eos_token_id": 2,
46
+ "feat_extract_activation": "gelu",
47
+ "feat_extract_dropout": 0.0,
48
+ "feat_extract_norm": "layer",
49
+ "feat_proj_dropout": 0.05,
50
+ "feat_quantizer_dropout": 0.0,
51
+ "final_dropout": 0.05,
52
+ "finetuning_task": [
53
+ "unispeech_sat_classification"
54
+ ],
55
+ "hidden_act": "gelu",
56
+ "hidden_dropout": 0.05,
57
+ "hidden_size": 1024,
58
+ "id2label": {
59
+ "0": "anger",
60
+ "1": "disgust",
61
+ "2": "enthusiasm",
62
+ "3": "fear",
63
+ "4": "happiness",
64
+ "5": "neutral",
65
+ "6": "sadness"
66
+ },
67
+ "initializer_range": 0.02,
68
+ "intermediate_size": 4096,
69
+ "label2id": {
70
+ "anger": 0,
71
+ "disgust": 1,
72
+ "enthusiasm": 2,
73
+ "fear": 3,
74
+ "happiness": 4,
75
+ "neutral": 5,
76
+ "sadness": 6
77
+ },
78
+ "layer_norm_eps": 1e-05,
79
+ "layerdrop": 0.05,
80
+ "mask_feature_length": 10,
81
+ "mask_feature_min_masks": 0,
82
+ "mask_feature_prob": 0.0,
83
+ "mask_time_length": 10,
84
+ "mask_time_min_masks": 2,
85
+ "mask_time_prob": 0.05,
86
+ "model_type": "unispeech-sat",
87
+ "num_attention_heads": 16,
88
+ "num_clusters": 504,
89
+ "num_codevector_groups": 2,
90
+ "num_codevectors_per_group": 320,
91
+ "num_conv_pos_embedding_groups": 16,
92
+ "num_conv_pos_embeddings": 128,
93
+ "num_feat_extract_layers": 7,
94
+ "num_hidden_layers": 24,
95
+ "num_negatives": 100,
96
+ "pad_token_id": 0,
97
+ "pooling_mode": "mean",
98
+ "problem_type": "single_label_classification",
99
+ "proj_codevector_dim": 768,
100
+ "quantization_config": {
101
+ "config_groups": {
102
+ "group_0": {
103
+ "format": "naive-quantized",
104
+ "input_activations": null,
105
+ "output_activations": null,
106
+ "targets": [
107
+ "Linear"
108
+ ],
109
+ "weights": {
110
+ "actorder": null,
111
+ "block_structure": null,
112
+ "dynamic": false,
113
+ "group_size": null,
114
+ "num_bits": 8,
115
+ "observer": "memoryless_minmax",
116
+ "observer_kwargs": {},
117
+ "scale_dtype": null,
118
+ "strategy": "channel",
119
+ "symmetric": true,
120
+ "type": "float",
121
+ "zp_dtype": null
122
+ }
123
+ }
124
+ },
125
+ "format": "naive-quantized",
126
+ "global_compression_ratio": null,
127
+ "ignore": [
128
+ "unispeech_sat.feature_projection.projection",
129
+ "projector",
130
+ "classifier"
131
+ ],
132
+ "kv_cache_scheme": null,
133
+ "quant_method": "compressed-tensors",
134
+ "quantization_status": "compressed",
135
+ "sparsity_config": {},
136
+ "transform_config": {},
137
+ "version": "0.17.1"
138
+ },
139
+ "tdnn_dilation": [
140
+ 1,
141
+ 2,
142
+ 3,
143
+ 1,
144
+ 1
145
+ ],
146
+ "tdnn_dim": [
147
+ 512,
148
+ 512,
149
+ 512,
150
+ 512,
151
+ 1500
152
+ ],
153
+ "tdnn_kernel": [
154
+ 5,
155
+ 3,
156
+ 3,
157
+ 1,
158
+ 1
159
+ ],
160
+ "tokenizer_class": "Wav2Vec2CTCTokenizer",
161
+ "transformers_version": "5.10.1",
162
+ "use_weighted_layer_sum": false,
163
+ "vocab_size": 40,
164
+ "xvector_output_dim": 512
165
+ }
fp8/model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c4451d1cae24923bffa3bdd10c2d34d68fb7dbc0365eeda53e5638c6c484d74b
3
+ size 359898548
fp8/preprocessor_config.json ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "do_normalize": true,
3
+ "feature_extractor_type": "Wav2Vec2FeatureExtractor",
4
+ "feature_size": 1,
5
+ "padding_side": "right",
6
+ "padding_value": 0,
7
+ "return_attention_mask": true,
8
+ "sampling_rate": 16000
9
+ }
fp8/recipe.yaml ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ default_stage:
2
+ default_modifiers:
3
+ QuantizationModifier:
4
+ config_groups:
5
+ group_0:
6
+ targets: [Linear]
7
+ weights:
8
+ num_bits: 8
9
+ type: float
10
+ symmetric: true
11
+ group_size: null
12
+ strategy: channel
13
+ block_structure: null
14
+ dynamic: false
15
+ actorder: null
16
+ scale_dtype: null
17
+ zp_dtype: null
18
+ observer: memoryless_minmax
19
+ observer_kwargs: {}
20
+ input_activations: null
21
+ output_activations: null
22
+ format: null
23
+ targets: [Linear]
24
+ ignore: ['re:.*feature_extractor.*', 're:.*feature_projection.*', 're:.*pos_conv_embed.*',
25
+ 're:.*gru_rel_pos_linear$', projector, classifier, 're:.*audio_projector$', 're:.*text_projector$',
26
+ 're:.*fusion_module.*', 're:.*text_model\.pooler.*']
27
+ bypass_divisibility_checks: false
fp8/special_tokens_map.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": "<s>",
3
+ "eos_token": "</s>",
4
+ "pad_token": "<pad>",
5
+ "unk_token": "<unk>"
6
+ }
fp8/tokenizer_config.json ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": "<s>",
3
+ "do_lower_case": false,
4
+ "eos_token": "</s>",
5
+ "model_max_length": 1000000000000000019884624838656,
6
+ "name_or_path": "Ar4ikov/unispeech-sat-projector-russian-resd",
7
+ "pad_token": "<pad>",
8
+ "processor_class": "Wav2Vec2Processor",
9
+ "replace_word_delimiter_char": " ",
10
+ "special_tokens_map_file": null,
11
+ "tokenizer_class": "Wav2Vec2CTCTokenizer",
12
+ "tokenizer_file": null,
13
+ "unk_token": "<unk>",
14
+ "word_delimiter_token": "|"
15
+ }
fp8/vocab.json ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "'": 5,
3
+ "-": 6,
4
+ "</s>": 2,
5
+ "<pad>": 0,
6
+ "<s>": 1,
7
+ "<unk>": 3,
8
+ "|": 4,
9
+ "а": 7,
10
+ "б": 8,
11
+ "в": 9,
12
+ "г": 10,
13
+ "д": 11,
14
+ "е": 12,
15
+ "ж": 13,
16
+ "з": 14,
17
+ "и": 15,
18
+ "й": 16,
19
+ "к": 17,
20
+ "л": 18,
21
+ "м": 19,
22
+ "н": 20,
23
+ "о": 21,
24
+ "п": 22,
25
+ "р": 23,
26
+ "с": 24,
27
+ "т": 25,
28
+ "у": 26,
29
+ "ф": 27,
30
+ "х": 28,
31
+ "ц": 29,
32
+ "ч": 30,
33
+ "ш": 31,
34
+ "щ": 32,
35
+ "ъ": 33,
36
+ "ы": 34,
37
+ "ь": 35,
38
+ "э": 36,
39
+ "ю": 37,
40
+ "я": 38,
41
+ "ё": 39
42
+ }
int4/config.json ADDED
@@ -0,0 +1,165 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "activation_dropout": 0.05,
3
+ "apply_spec_augment": true,
4
+ "architectures": [
5
+ "UniSpeechSatForSequenceClassification"
6
+ ],
7
+ "attention_dropout": 0.05,
8
+ "bos_token_id": 1,
9
+ "classifier_proj_size": 768,
10
+ "codevector_dim": 768,
11
+ "contrastive_logits_temperature": 0.1,
12
+ "conv_bias": false,
13
+ "conv_dim": [
14
+ 512,
15
+ 512,
16
+ 512,
17
+ 512,
18
+ 512,
19
+ 512,
20
+ 512
21
+ ],
22
+ "conv_kernel": [
23
+ 10,
24
+ 3,
25
+ 3,
26
+ 3,
27
+ 3,
28
+ 2,
29
+ 2
30
+ ],
31
+ "conv_stride": [
32
+ 5,
33
+ 2,
34
+ 2,
35
+ 2,
36
+ 2,
37
+ 2,
38
+ 2
39
+ ],
40
+ "ctc_loss_reduction": "sum",
41
+ "ctc_zero_infinity": false,
42
+ "diversity_loss_weight": 0.1,
43
+ "do_stable_layer_norm": true,
44
+ "dtype": "float32",
45
+ "eos_token_id": 2,
46
+ "feat_extract_activation": "gelu",
47
+ "feat_extract_dropout": 0.0,
48
+ "feat_extract_norm": "layer",
49
+ "feat_proj_dropout": 0.05,
50
+ "feat_quantizer_dropout": 0.0,
51
+ "final_dropout": 0.05,
52
+ "finetuning_task": [
53
+ "unispeech_sat_classification"
54
+ ],
55
+ "hidden_act": "gelu",
56
+ "hidden_dropout": 0.05,
57
+ "hidden_size": 1024,
58
+ "id2label": {
59
+ "0": "anger",
60
+ "1": "disgust",
61
+ "2": "enthusiasm",
62
+ "3": "fear",
63
+ "4": "happiness",
64
+ "5": "neutral",
65
+ "6": "sadness"
66
+ },
67
+ "initializer_range": 0.02,
68
+ "intermediate_size": 4096,
69
+ "label2id": {
70
+ "anger": 0,
71
+ "disgust": 1,
72
+ "enthusiasm": 2,
73
+ "fear": 3,
74
+ "happiness": 4,
75
+ "neutral": 5,
76
+ "sadness": 6
77
+ },
78
+ "layer_norm_eps": 1e-05,
79
+ "layerdrop": 0.05,
80
+ "mask_feature_length": 10,
81
+ "mask_feature_min_masks": 0,
82
+ "mask_feature_prob": 0.0,
83
+ "mask_time_length": 10,
84
+ "mask_time_min_masks": 2,
85
+ "mask_time_prob": 0.05,
86
+ "model_type": "unispeech-sat",
87
+ "num_attention_heads": 16,
88
+ "num_clusters": 504,
89
+ "num_codevector_groups": 2,
90
+ "num_codevectors_per_group": 320,
91
+ "num_conv_pos_embedding_groups": 16,
92
+ "num_conv_pos_embeddings": 128,
93
+ "num_feat_extract_layers": 7,
94
+ "num_hidden_layers": 24,
95
+ "num_negatives": 100,
96
+ "pad_token_id": 0,
97
+ "pooling_mode": "mean",
98
+ "problem_type": "single_label_classification",
99
+ "proj_codevector_dim": 768,
100
+ "quantization_config": {
101
+ "config_groups": {
102
+ "group_0": {
103
+ "format": "pack-quantized",
104
+ "input_activations": null,
105
+ "output_activations": null,
106
+ "targets": [
107
+ "Linear"
108
+ ],
109
+ "weights": {
110
+ "actorder": null,
111
+ "block_structure": null,
112
+ "dynamic": false,
113
+ "group_size": 128,
114
+ "num_bits": 4,
115
+ "observer": "memoryless_minmax",
116
+ "observer_kwargs": {},
117
+ "scale_dtype": null,
118
+ "strategy": "group",
119
+ "symmetric": false,
120
+ "type": "int",
121
+ "zp_dtype": "torch.int8"
122
+ }
123
+ }
124
+ },
125
+ "format": "pack-quantized",
126
+ "global_compression_ratio": null,
127
+ "ignore": [
128
+ "unispeech_sat.feature_projection.projection",
129
+ "projector",
130
+ "classifier"
131
+ ],
132
+ "kv_cache_scheme": null,
133
+ "quant_method": "compressed-tensors",
134
+ "quantization_status": "compressed",
135
+ "sparsity_config": {},
136
+ "transform_config": {},
137
+ "version": "0.17.1"
138
+ },
139
+ "tdnn_dilation": [
140
+ 1,
141
+ 2,
142
+ 3,
143
+ 1,
144
+ 1
145
+ ],
146
+ "tdnn_dim": [
147
+ 512,
148
+ 512,
149
+ 512,
150
+ 512,
151
+ 1500
152
+ ],
153
+ "tdnn_kernel": [
154
+ 5,
155
+ 3,
156
+ 3,
157
+ 1,
158
+ 1
159
+ ],
160
+ "tokenizer_class": "Wav2Vec2CTCTokenizer",
161
+ "transformers_version": "5.10.1",
162
+ "use_weighted_layer_sum": false,
163
+ "vocab_size": 40,
164
+ "xvector_output_dim": 512
165
+ }
int4/model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1796bf3fc97b4dab77d52b621a66a6148da582f7f2cf89d0e5226865c63d79b5
3
+ size 218676540
int4/preprocessor_config.json ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "do_normalize": true,
3
+ "feature_extractor_type": "Wav2Vec2FeatureExtractor",
4
+ "feature_size": 1,
5
+ "padding_side": "right",
6
+ "padding_value": 0,
7
+ "return_attention_mask": true,
8
+ "sampling_rate": 16000
9
+ }
int4/recipe.yaml ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ default_stage:
2
+ default_modifiers:
3
+ QuantizationModifier:
4
+ targets: [Linear]
5
+ ignore: ['re:.*feature_extractor.*', 're:.*feature_projection.*', 're:.*pos_conv_embed.*',
6
+ 're:.*gru_rel_pos_linear$', projector, classifier, 're:.*audio_projector$', 're:.*text_projector$',
7
+ 're:.*fusion_module.*', 're:.*text_model\.pooler.*']
8
+ scheme: W4A16_ASYM
9
+ bypass_divisibility_checks: false
int4/special_tokens_map.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": "<s>",
3
+ "eos_token": "</s>",
4
+ "pad_token": "<pad>",
5
+ "unk_token": "<unk>"
6
+ }
int4/tokenizer_config.json ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": "<s>",
3
+ "do_lower_case": false,
4
+ "eos_token": "</s>",
5
+ "model_max_length": 1000000000000000019884624838656,
6
+ "name_or_path": "Ar4ikov/unispeech-sat-projector-russian-resd",
7
+ "pad_token": "<pad>",
8
+ "processor_class": "Wav2Vec2Processor",
9
+ "replace_word_delimiter_char": " ",
10
+ "special_tokens_map_file": null,
11
+ "tokenizer_class": "Wav2Vec2CTCTokenizer",
12
+ "tokenizer_file": null,
13
+ "unk_token": "<unk>",
14
+ "word_delimiter_token": "|"
15
+ }
int4/vocab.json ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "'": 5,
3
+ "-": 6,
4
+ "</s>": 2,
5
+ "<pad>": 0,
6
+ "<s>": 1,
7
+ "<unk>": 3,
8
+ "|": 4,
9
+ "а": 7,
10
+ "б": 8,
11
+ "в": 9,
12
+ "г": 10,
13
+ "д": 11,
14
+ "е": 12,
15
+ "ж": 13,
16
+ "з": 14,
17
+ "и": 15,
18
+ "й": 16,
19
+ "к": 17,
20
+ "л": 18,
21
+ "м": 19,
22
+ "н": 20,
23
+ "о": 21,
24
+ "п": 22,
25
+ "р": 23,
26
+ "с": 24,
27
+ "т": 25,
28
+ "у": 26,
29
+ "ф": 27,
30
+ "х": 28,
31
+ "ц": 29,
32
+ "ч": 30,
33
+ "ш": 31,
34
+ "щ": 32,
35
+ "ъ": 33,
36
+ "ы": 34,
37
+ "ь": 35,
38
+ "э": 36,
39
+ "ю": 37,
40
+ "я": 38,
41
+ "ё": 39
42
+ }
int8/config.json ADDED
@@ -0,0 +1,165 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "activation_dropout": 0.05,
3
+ "apply_spec_augment": true,
4
+ "architectures": [
5
+ "UniSpeechSatForSequenceClassification"
6
+ ],
7
+ "attention_dropout": 0.05,
8
+ "bos_token_id": 1,
9
+ "classifier_proj_size": 768,
10
+ "codevector_dim": 768,
11
+ "contrastive_logits_temperature": 0.1,
12
+ "conv_bias": false,
13
+ "conv_dim": [
14
+ 512,
15
+ 512,
16
+ 512,
17
+ 512,
18
+ 512,
19
+ 512,
20
+ 512
21
+ ],
22
+ "conv_kernel": [
23
+ 10,
24
+ 3,
25
+ 3,
26
+ 3,
27
+ 3,
28
+ 2,
29
+ 2
30
+ ],
31
+ "conv_stride": [
32
+ 5,
33
+ 2,
34
+ 2,
35
+ 2,
36
+ 2,
37
+ 2,
38
+ 2
39
+ ],
40
+ "ctc_loss_reduction": "sum",
41
+ "ctc_zero_infinity": false,
42
+ "diversity_loss_weight": 0.1,
43
+ "do_stable_layer_norm": true,
44
+ "dtype": "float32",
45
+ "eos_token_id": 2,
46
+ "feat_extract_activation": "gelu",
47
+ "feat_extract_dropout": 0.0,
48
+ "feat_extract_norm": "layer",
49
+ "feat_proj_dropout": 0.05,
50
+ "feat_quantizer_dropout": 0.0,
51
+ "final_dropout": 0.05,
52
+ "finetuning_task": [
53
+ "unispeech_sat_classification"
54
+ ],
55
+ "hidden_act": "gelu",
56
+ "hidden_dropout": 0.05,
57
+ "hidden_size": 1024,
58
+ "id2label": {
59
+ "0": "anger",
60
+ "1": "disgust",
61
+ "2": "enthusiasm",
62
+ "3": "fear",
63
+ "4": "happiness",
64
+ "5": "neutral",
65
+ "6": "sadness"
66
+ },
67
+ "initializer_range": 0.02,
68
+ "intermediate_size": 4096,
69
+ "label2id": {
70
+ "anger": 0,
71
+ "disgust": 1,
72
+ "enthusiasm": 2,
73
+ "fear": 3,
74
+ "happiness": 4,
75
+ "neutral": 5,
76
+ "sadness": 6
77
+ },
78
+ "layer_norm_eps": 1e-05,
79
+ "layerdrop": 0.05,
80
+ "mask_feature_length": 10,
81
+ "mask_feature_min_masks": 0,
82
+ "mask_feature_prob": 0.0,
83
+ "mask_time_length": 10,
84
+ "mask_time_min_masks": 2,
85
+ "mask_time_prob": 0.05,
86
+ "model_type": "unispeech-sat",
87
+ "num_attention_heads": 16,
88
+ "num_clusters": 504,
89
+ "num_codevector_groups": 2,
90
+ "num_codevectors_per_group": 320,
91
+ "num_conv_pos_embedding_groups": 16,
92
+ "num_conv_pos_embeddings": 128,
93
+ "num_feat_extract_layers": 7,
94
+ "num_hidden_layers": 24,
95
+ "num_negatives": 100,
96
+ "pad_token_id": 0,
97
+ "pooling_mode": "mean",
98
+ "problem_type": "single_label_classification",
99
+ "proj_codevector_dim": 768,
100
+ "quantization_config": {
101
+ "config_groups": {
102
+ "group_0": {
103
+ "format": "pack-quantized",
104
+ "input_activations": null,
105
+ "output_activations": null,
106
+ "targets": [
107
+ "Linear"
108
+ ],
109
+ "weights": {
110
+ "actorder": null,
111
+ "block_structure": null,
112
+ "dynamic": false,
113
+ "group_size": 128,
114
+ "num_bits": 8,
115
+ "observer": "memoryless_minmax",
116
+ "observer_kwargs": {},
117
+ "scale_dtype": null,
118
+ "strategy": "group",
119
+ "symmetric": true,
120
+ "type": "int",
121
+ "zp_dtype": null
122
+ }
123
+ }
124
+ },
125
+ "format": "pack-quantized",
126
+ "global_compression_ratio": null,
127
+ "ignore": [
128
+ "unispeech_sat.feature_projection.projection",
129
+ "projector",
130
+ "classifier"
131
+ ],
132
+ "kv_cache_scheme": null,
133
+ "quant_method": "compressed-tensors",
134
+ "quantization_status": "compressed",
135
+ "sparsity_config": {},
136
+ "transform_config": {},
137
+ "version": "0.17.1"
138
+ },
139
+ "tdnn_dilation": [
140
+ 1,
141
+ 2,
142
+ 3,
143
+ 1,
144
+ 1
145
+ ],
146
+ "tdnn_dim": [
147
+ 512,
148
+ 512,
149
+ 512,
150
+ 512,
151
+ 1500
152
+ ],
153
+ "tdnn_kernel": [
154
+ 5,
155
+ 3,
156
+ 3,
157
+ 1,
158
+ 1
159
+ ],
160
+ "tokenizer_class": "Wav2Vec2CTCTokenizer",
161
+ "transformers_version": "5.10.1",
162
+ "use_weighted_layer_sum": false,
163
+ "vocab_size": 40,
164
+ "xvector_output_dim": 512
165
+ }
int8/model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9456bed8f72193adc478396a1cc7acbf35c3bab05b40083c0510b6cf1c5b8492
3
+ size 368471492
int8/preprocessor_config.json ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "do_normalize": true,
3
+ "feature_extractor_type": "Wav2Vec2FeatureExtractor",
4
+ "feature_size": 1,
5
+ "padding_side": "right",
6
+ "padding_value": 0,
7
+ "return_attention_mask": true,
8
+ "sampling_rate": 16000
9
+ }
int8/recipe.yaml ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ default_stage:
2
+ default_modifiers:
3
+ QuantizationModifier:
4
+ targets: [Linear]
5
+ ignore: ['re:.*feature_extractor.*', 're:.*feature_projection.*', 're:.*pos_conv_embed.*',
6
+ 're:.*gru_rel_pos_linear$', projector, classifier, 're:.*audio_projector$', 're:.*text_projector$',
7
+ 're:.*fusion_module.*', 're:.*text_model\.pooler.*']
8
+ scheme: W8A16
9
+ bypass_divisibility_checks: false
int8/special_tokens_map.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": "<s>",
3
+ "eos_token": "</s>",
4
+ "pad_token": "<pad>",
5
+ "unk_token": "<unk>"
6
+ }
int8/tokenizer_config.json ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": "<s>",
3
+ "do_lower_case": false,
4
+ "eos_token": "</s>",
5
+ "model_max_length": 1000000000000000019884624838656,
6
+ "name_or_path": "Ar4ikov/unispeech-sat-projector-russian-resd",
7
+ "pad_token": "<pad>",
8
+ "processor_class": "Wav2Vec2Processor",
9
+ "replace_word_delimiter_char": " ",
10
+ "special_tokens_map_file": null,
11
+ "tokenizer_class": "Wav2Vec2CTCTokenizer",
12
+ "tokenizer_file": null,
13
+ "unk_token": "<unk>",
14
+ "word_delimiter_token": "|"
15
+ }
int8/vocab.json ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "'": 5,
3
+ "-": 6,
4
+ "</s>": 2,
5
+ "<pad>": 0,
6
+ "<s>": 1,
7
+ "<unk>": 3,
8
+ "|": 4,
9
+ "а": 7,
10
+ "б": 8,
11
+ "в": 9,
12
+ "г": 10,
13
+ "д": 11,
14
+ "е": 12,
15
+ "ж": 13,
16
+ "з": 14,
17
+ "и": 15,
18
+ "й": 16,
19
+ "к": 17,
20
+ "л": 18,
21
+ "м": 19,
22
+ "н": 20,
23
+ "о": 21,
24
+ "п": 22,
25
+ "р": 23,
26
+ "с": 24,
27
+ "т": 25,
28
+ "у": 26,
29
+ "ф": 27,
30
+ "х": 28,
31
+ "ц": 29,
32
+ "ч": 30,
33
+ "ш": 31,
34
+ "щ": 32,
35
+ "ъ": 33,
36
+ "ы": 34,
37
+ "ь": 35,
38
+ "э": 36,
39
+ "ю": 37,
40
+ "я": 38,
41
+ "ё": 39
42
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4974bb17eb3f6e3e464448c05cf5ca9ef3715c5961063e68a5a9c92d183e7372
3
+ size 1264964820
preprocessor_config.json ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "do_normalize": true,
3
+ "feature_extractor_type": "Wav2Vec2FeatureExtractor",
4
+ "feature_size": 1,
5
+ "padding_side": "right",
6
+ "padding_value": 0,
7
+ "return_attention_mask": true,
8
+ "sampling_rate": 16000
9
+ }
special_tokens_map.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": "<s>",
3
+ "eos_token": "</s>",
4
+ "pad_token": "<pad>",
5
+ "unk_token": "<unk>"
6
+ }
tokenizer_config.json ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": "<s>",
3
+ "do_lower_case": false,
4
+ "eos_token": "</s>",
5
+ "model_max_length": 1000000000000000019884624838656,
6
+ "name_or_path": "Ar4ikov/unispeech-sat-projector-russian-resd",
7
+ "pad_token": "<pad>",
8
+ "processor_class": "Wav2Vec2Processor",
9
+ "replace_word_delimiter_char": " ",
10
+ "special_tokens_map_file": null,
11
+ "tokenizer_class": "Wav2Vec2CTCTokenizer",
12
+ "tokenizer_file": null,
13
+ "unk_token": "<unk>",
14
+ "word_delimiter_token": "|"
15
+ }
vocab.json ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "'": 5,
3
+ "-": 6,
4
+ "</s>": 2,
5
+ "<pad>": 0,
6
+ "<s>": 1,
7
+ "<unk>": 3,
8
+ "|": 4,
9
+ "а": 7,
10
+ "б": 8,
11
+ "в": 9,
12
+ "г": 10,
13
+ "д": 11,
14
+ "е": 12,
15
+ "ж": 13,
16
+ "з": 14,
17
+ "и": 15,
18
+ "й": 16,
19
+ "к": 17,
20
+ "л": 18,
21
+ "м": 19,
22
+ "н": 20,
23
+ "о": 21,
24
+ "п": 22,
25
+ "р": 23,
26
+ "с": 24,
27
+ "т": 25,
28
+ "у": 26,
29
+ "ф": 27,
30
+ "х": 28,
31
+ "ц": 29,
32
+ "ч": 30,
33
+ "ш": 31,
34
+ "щ": 32,
35
+ "ъ": 33,
36
+ "ы": 34,
37
+ "ь": 35,
38
+ "э": 36,
39
+ "ю": 37,
40
+ "я": 38,
41
+ "ё": 39
42
+ }