André Pedro Ribeiro commited on
Commit
ddbfe9d
·
verified ·
1 Parent(s): 95b5804

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,579 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ tags:
3
+ - sentence-transformers
4
+ - cross-encoder
5
+ - reranker
6
+ - generated_from_trainer
7
+ - dataset_size:257341
8
+ - loss:BinaryCrossEntropyLoss
9
+ pipeline_tag: text-ranking
10
+ library_name: sentence-transformers
11
+ metrics:
12
+ - map
13
+ - mrr@10
14
+ - ndcg@10
15
+ model-index:
16
+ - name: CrossEncoder
17
+ results:
18
+ - task:
19
+ type: cross-encoder-reranking
20
+ name: Cross Encoder Reranking
21
+ dataset:
22
+ name: mmarco pt dev
23
+ type: mmarco-pt-dev
24
+ metrics:
25
+ - type: map
26
+ value: 0.9332
27
+ name: Map
28
+ - type: mrr@10
29
+ value: 0.933
30
+ name: Mrr@10
31
+ - type: ndcg@10
32
+ value: 0.9449
33
+ name: Ndcg@10
34
+ ---
35
+
36
+ # CrossEncoder
37
+
38
+ This is a [Cross Encoder](https://www.sbert.net/docs/cross_encoder/usage/usage.html) model trained using the [sentence-transformers](https://www.SBERT.net) library. It computes scores for pairs of texts, which can be used for text reranking and semantic search.
39
+
40
+ ## Model Details
41
+
42
+ ### Model Description
43
+ - **Model Type:** Cross Encoder
44
+ <!-- - **Base model:** [Unknown](https://huggingface.co/unknown) -->
45
+ - **Maximum Sequence Length:** 512 tokens
46
+ - **Number of Output Labels:** 1 label
47
+ <!-- - **Training Dataset:** Unknown -->
48
+ <!-- - **Language:** Unknown -->
49
+ <!-- - **License:** Unknown -->
50
+
51
+ ### Model Sources
52
+
53
+ - **Documentation:** [Sentence Transformers Documentation](https://sbert.net)
54
+ - **Documentation:** [Cross Encoder Documentation](https://www.sbert.net/docs/cross_encoder/usage/usage.html)
55
+ - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers)
56
+ - **Hugging Face:** [Cross Encoders on Hugging Face](https://huggingface.co/models?library=sentence-transformers&other=cross-encoder)
57
+
58
+ ## Usage
59
+
60
+ ### Direct Usage (Sentence Transformers)
61
+
62
+ First install the Sentence Transformers library:
63
+
64
+ ```bash
65
+ pip install -U sentence-transformers
66
+ ```
67
+
68
+ Then you can load this model and run inference.
69
+ ```python
70
+ from sentence_transformers import CrossEncoder
71
+
72
+ # Download from the 🤗 Hub
73
+ model = CrossEncoder("cross_encoder_model_id")
74
+ # Get scores for pairs of texts
75
+ pairs = [
76
+ ['altura média para mulheres de 8 anos', 'Obtenha ajuda de um médico agora ࢠ¬Âº. Detalhes sobre sua altura: Supondo que uma menina de 8 anos de idade média de 50 polegadas (50º percentil), o peso médio é 57 libras (um índice de massa corporal (IMM) de 50º percentil). Obtenha ajuda de um médico agora à¢ à ‚ €Âº. Detalhes sobre sua altura: Supondo que uma menina de 8 anos de idade média de 50 polegadas (50º percentil), o peso médio é 57 libras (um índice de massa corporal (IMM) de 50º percentil). O IMM não é uma medida perfeita, portanto, existem variações com base no tipo de corpo, especialmente a massa muscular.'],
77
+ ['altura média para mulheres de 8 anos', 'a altura média de um homem asiático pode ser de 5,2 a 5,9 dependendo de onde ele é, por exemplo, um homem chinês tem uma média de 5,7 ou menos. Uma pequena edição - a altura média do homem persa é entre 5 pés e 5 polegadas de altura e 5 pés e 8 polegadas de altura. Os homens persas são conhecidos por serem de estatura média.'],
78
+ ['altura média para mulheres de 8 anos', "Estudos conduzidos pelo National Center for Health Statistics mostram que a altura média de um homem americano (20-74 anos) aumentou de 5'8 para 5'9Ã⠀ šÃ‚½ nos últimos quarenta anos. O índice de massa corporal (IMC) médio também aumentou entre os adultos americanos, de aproximadamente 25 em 1960 para 28 em 2002."],
79
+ ['altura média para mulheres de 8 anos', "A altura média de um homem americano caucasiano é 5 '10. Eu me sinto muito baixo. Quando eu era mais jovem do que agora, sonhava ter pelo menos 5 '8. Fiquei me perguntando com que idade as placas de crescimento nos homens se fecham e se alguém sabe alguma coisa que posso fazer para aumentar minha altura."],
80
+ ['altura média para mulheres de 8 anos', 'A meia-idade tecnicamente se refere a atingir uma idade em que você viveu a metade da expectativa de vida média para o seu gênero. O ponto médio da vida é agora cerca de 40 para as mulheres e 38 para os homens (os homens tendem a morrer 6 a 8 anos antes das mulheres). Em comparação, há 100 anos as mulheres chegavam à meia-idade aos 22, principalmente porque muitas morriam no parto.'],
81
+ ]
82
+ scores = model.predict(pairs)
83
+ print(scores.shape)
84
+ # (5,)
85
+
86
+ # Or rank different texts based on similarity to a single text
87
+ ranks = model.rank(
88
+ 'altura média para mulheres de 8 anos',
89
+ [
90
+ 'Obtenha ajuda de um médico agora ࢠ¬Âº. Detalhes sobre sua altura: Supondo que uma menina de 8 anos de idade média de 50 polegadas (50º percentil), o peso médio é 57 libras (um índice de massa corporal (IMM) de 50º percentil). Obtenha ajuda de um médico agora à¢ à ‚ €Âº. Detalhes sobre sua altura: Supondo que uma menina de 8 anos de idade média de 50 polegadas (50º percentil), o peso médio é 57 libras (um índice de massa corporal (IMM) de 50º percentil). O IMM não é uma medida perfeita, portanto, existem variações com base no tipo de corpo, especialmente a massa muscular.',
91
+ 'a altura média de um homem asiático pode ser de 5,2 a 5,9 dependendo de onde ele é, por exemplo, um homem chinês tem uma média de 5,7 ou menos. Uma pequena edição - a altura média do homem persa é entre 5 pés e 5 polegadas de altura e 5 pés e 8 polegadas de altura. Os homens persas são conhecidos por serem de estatura média.',
92
+ "Estudos conduzidos pelo National Center for Health Statistics mostram que a altura média de um homem americano (20-74 anos) aumentou de 5'8 para 5'9Ã⠀ šÃ‚½ nos últimos quarenta anos. O índice de massa corporal (IMC) médio também aumentou entre os adultos americanos, de aproximadamente 25 em 1960 para 28 em 2002.",
93
+ "A altura média de um homem americano caucasiano é 5 '10. Eu me sinto muito baixo. Quando eu era mais jovem do que agora, sonhava ter pelo menos 5 '8. Fiquei me perguntando com que idade as placas de crescimento nos homens se fecham e se alguém sabe alguma coisa que posso fazer para aumentar minha altura.",
94
+ 'A meia-idade tecnicamente se refere a atingir uma idade em que você viveu a metade da expectativa de vida média para o seu gênero. O ponto médio da vida é agora cerca de 40 para as mulheres e 38 para os homens (os homens tendem a morrer 6 a 8 anos antes das mulheres). Em comparação, há 100 anos as mulheres chegavam à meia-idade aos 22, principalmente porque muitas morriam no parto.',
95
+ ]
96
+ )
97
+ # [{'corpus_id': ..., 'score': ...}, {'corpus_id': ..., 'score': ...}, ...]
98
+ ```
99
+
100
+ <!--
101
+ ### Direct Usage (Transformers)
102
+
103
+ <details><summary>Click to see the direct usage in Transformers</summary>
104
+
105
+ </details>
106
+ -->
107
+
108
+ <!--
109
+ ### Downstream Usage (Sentence Transformers)
110
+
111
+ You can finetune this model on your own dataset.
112
+
113
+ <details><summary>Click to expand</summary>
114
+
115
+ </details>
116
+ -->
117
+
118
+ <!--
119
+ ### Out-of-Scope Use
120
+
121
+ *List how the model may foreseeably be misused and address what users ought not to do with the model.*
122
+ -->
123
+
124
+ ## Evaluation
125
+
126
+ ### Metrics
127
+
128
+ #### Cross Encoder Reranking
129
+
130
+ * Dataset: `mmarco-pt-dev`
131
+ * Evaluated with [<code>CrossEncoderRerankingEvaluator</code>](https://sbert.net/docs/package_reference/cross_encoder/evaluation.html#sentence_transformers.cross_encoder.evaluation.CrossEncoderRerankingEvaluator) with these parameters:
132
+ ```json
133
+ {
134
+ "at_k": 10,
135
+ "always_rerank_positives": false
136
+ }
137
+ ```
138
+
139
+ | Metric | Value |
140
+ |:------------|:---------------------|
141
+ | map | 0.9332 (+0.0407) |
142
+ | mrr@10 | 0.9330 (+0.0413) |
143
+ | **ndcg@10** | **0.9449 (+0.0334)** |
144
+
145
+ <!--
146
+ ## Bias, Risks and Limitations
147
+
148
+ *What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.*
149
+ -->
150
+
151
+ <!--
152
+ ### Recommendations
153
+
154
+ *What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.*
155
+ -->
156
+
157
+ ## Training Details
158
+
159
+ ### Training Dataset
160
+
161
+ #### Unnamed Dataset
162
+
163
+ * Size: 257,341 training samples
164
+ * Columns: <code>query</code>, <code>positive</code>, and <code>label</code>
165
+ * Approximate statistics based on the first 1000 samples:
166
+ | | query | positive | label |
167
+ |:--------|:------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------------|:------------------------------------------------|
168
+ | type | string | string | int |
169
+ | details | <ul><li>min: 16 characters</li><li>mean: 37.63 characters</li><li>max: 126 characters</li></ul> | <ul><li>min: 65 characters</li><li>mean: 401.06 characters</li><li>max: 1195 characters</li></ul> | <ul><li>0: ~81.10%</li><li>1: ~18.90%</li></ul> |
170
+ * Samples:
171
+ | query | positive | label |
172
+ |:--------------------------------------------------|:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:---------------|
173
+ | <code>altura média para mulheres de 8 anos</code> | <code>Obtenha ajuda de um médico agora ࢠ¬Âº. Detalhes sobre sua altura: Supondo que uma menina de 8 anos de idade média de 50 polegadas (50º percentil), o peso médio é 57 libras (um índice de massa corporal (IMM) de 50º percentil). Obtenha ajuda de um médico agora à¢ à ‚ €Âº. Detalhes sobre sua altura: Supondo que uma menina de 8 anos de idade média de 50 polegadas (50º percentil), o peso médio é 57 libras (um índice de massa corporal (IMM) de 50º percentil). O IMM não é uma medida perfeita, portanto, existem variações com base no tipo de corpo, especialmente a massa muscular.</code> | <code>1</code> |
174
+ | <code>altura média para mulheres de 8 anos</code> | <code>a altura média de um homem asiático pode ser de 5,2 a 5,9 dependendo de onde ele é, por exemplo, um homem chinês tem uma média de 5,7 ou menos. Uma pequena edição - a altura média do homem persa é entre 5 pés e 5 polegadas de altura e 5 pés e 8 polegadas de altura. Os homens persas são conhecidos por serem de estatura média.</code> | <code>0</code> |
175
+ | <code>altura média para mulheres de 8 anos</code> | <code>Estudos conduzidos pelo National Center for Health Statistics mostram que a altura média de um homem americano (20-74 anos) aumentou de 5'8 para 5'9Ã⠀ šÃ‚½ nos últimos quarenta anos. O índice de massa corporal (IMC) médio também aumentou entre os adultos americanos, de aproximadamente 25 em 1960 para 28 em 2002.</code> | <code>0</code> |
176
+ * Loss: [<code>BinaryCrossEntropyLoss</code>](https://sbert.net/docs/package_reference/cross_encoder/losses.html#binarycrossentropyloss) with these parameters:
177
+ ```json
178
+ {
179
+ "activation_fn": "torch.nn.modules.linear.Identity",
180
+ "pos_weight": 5.0
181
+ }
182
+ ```
183
+
184
+ ### Training Hyperparameters
185
+ #### Non-Default Hyperparameters
186
+
187
+ - `eval_strategy`: steps
188
+ - `per_device_train_batch_size`: 256
189
+ - `per_device_eval_batch_size`: 256
190
+ - `learning_rate`: 2e-05
191
+ - `weight_decay`: 0.01
192
+ - `num_train_epochs`: 2
193
+ - `warmup_ratio`: 0.1
194
+ - `bf16`: True
195
+ - `load_best_model_at_end`: True
196
+ - `optim`: adamw_torch
197
+ - `gradient_checkpointing`: True
198
+
199
+ #### All Hyperparameters
200
+ <details><summary>Click to expand</summary>
201
+
202
+ - `overwrite_output_dir`: False
203
+ - `do_predict`: False
204
+ - `eval_strategy`: steps
205
+ - `prediction_loss_only`: True
206
+ - `per_device_train_batch_size`: 256
207
+ - `per_device_eval_batch_size`: 256
208
+ - `per_gpu_train_batch_size`: None
209
+ - `per_gpu_eval_batch_size`: None
210
+ - `gradient_accumulation_steps`: 1
211
+ - `eval_accumulation_steps`: None
212
+ - `torch_empty_cache_steps`: None
213
+ - `learning_rate`: 2e-05
214
+ - `weight_decay`: 0.01
215
+ - `adam_beta1`: 0.9
216
+ - `adam_beta2`: 0.999
217
+ - `adam_epsilon`: 1e-08
218
+ - `max_grad_norm`: 1.0
219
+ - `num_train_epochs`: 2
220
+ - `max_steps`: -1
221
+ - `lr_scheduler_type`: linear
222
+ - `lr_scheduler_kwargs`: {}
223
+ - `warmup_ratio`: 0.1
224
+ - `warmup_steps`: 0
225
+ - `log_level`: passive
226
+ - `log_level_replica`: warning
227
+ - `log_on_each_node`: True
228
+ - `logging_nan_inf_filter`: True
229
+ - `save_safetensors`: True
230
+ - `save_on_each_node`: False
231
+ - `save_only_model`: False
232
+ - `restore_callback_states_from_checkpoint`: False
233
+ - `no_cuda`: False
234
+ - `use_cpu`: False
235
+ - `use_mps_device`: False
236
+ - `seed`: 42
237
+ - `data_seed`: None
238
+ - `jit_mode_eval`: False
239
+ - `bf16`: True
240
+ - `fp16`: False
241
+ - `fp16_opt_level`: O1
242
+ - `half_precision_backend`: auto
243
+ - `bf16_full_eval`: False
244
+ - `fp16_full_eval`: False
245
+ - `tf32`: None
246
+ - `local_rank`: 0
247
+ - `ddp_backend`: None
248
+ - `tpu_num_cores`: None
249
+ - `tpu_metrics_debug`: False
250
+ - `debug`: []
251
+ - `dataloader_drop_last`: False
252
+ - `dataloader_num_workers`: 0
253
+ - `dataloader_prefetch_factor`: None
254
+ - `past_index`: -1
255
+ - `disable_tqdm`: False
256
+ - `remove_unused_columns`: True
257
+ - `label_names`: None
258
+ - `load_best_model_at_end`: True
259
+ - `ignore_data_skip`: False
260
+ - `fsdp`: []
261
+ - `fsdp_min_num_params`: 0
262
+ - `fsdp_config`: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
263
+ - `fsdp_transformer_layer_cls_to_wrap`: None
264
+ - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
265
+ - `parallelism_config`: None
266
+ - `deepspeed`: None
267
+ - `label_smoothing_factor`: 0.0
268
+ - `optim`: adamw_torch
269
+ - `optim_args`: None
270
+ - `adafactor`: False
271
+ - `group_by_length`: False
272
+ - `length_column_name`: length
273
+ - `project`: huggingface
274
+ - `trackio_space_id`: trackio
275
+ - `ddp_find_unused_parameters`: None
276
+ - `ddp_bucket_cap_mb`: None
277
+ - `ddp_broadcast_buffers`: False
278
+ - `dataloader_pin_memory`: True
279
+ - `dataloader_persistent_workers`: False
280
+ - `skip_memory_metrics`: True
281
+ - `use_legacy_prediction_loop`: False
282
+ - `push_to_hub`: False
283
+ - `resume_from_checkpoint`: None
284
+ - `hub_model_id`: None
285
+ - `hub_strategy`: every_save
286
+ - `hub_private_repo`: None
287
+ - `hub_always_push`: False
288
+ - `hub_revision`: None
289
+ - `gradient_checkpointing`: True
290
+ - `gradient_checkpointing_kwargs`: None
291
+ - `include_inputs_for_metrics`: False
292
+ - `include_for_metrics`: []
293
+ - `eval_do_concat_batches`: True
294
+ - `fp16_backend`: auto
295
+ - `push_to_hub_model_id`: None
296
+ - `push_to_hub_organization`: None
297
+ - `mp_parameters`:
298
+ - `auto_find_batch_size`: False
299
+ - `full_determinism`: False
300
+ - `torchdynamo`: None
301
+ - `ray_scope`: last
302
+ - `ddp_timeout`: 1800
303
+ - `torch_compile`: False
304
+ - `torch_compile_backend`: None
305
+ - `torch_compile_mode`: None
306
+ - `include_tokens_per_second`: False
307
+ - `include_num_input_tokens_seen`: no
308
+ - `neftune_noise_alpha`: None
309
+ - `optim_target_modules`: None
310
+ - `batch_eval_metrics`: False
311
+ - `eval_on_start`: False
312
+ - `use_liger_kernel`: False
313
+ - `liger_kernel_config`: None
314
+ - `eval_use_gather_object`: False
315
+ - `average_tokens_across_devices`: True
316
+ - `prompts`: None
317
+ - `batch_sampler`: batch_sampler
318
+ - `multi_dataset_batch_sampler`: proportional
319
+ - `router_mapping`: {}
320
+ - `learning_rate_mapping`: {}
321
+
322
+ </details>
323
+
324
+ ### Training Logs
325
+ <details><summary>Click to expand</summary>
326
+
327
+ | Epoch | Step | Training Loss | mmarco-pt-dev_ndcg@10 |
328
+ |:----------:|:--------:|:-------------:|:---------------------:|
329
+ | -1 | -1 | - | 0.9291 (+0.0176) |
330
+ | 0.0010 | 1 | 0.9329 | - |
331
+ | 0.0099 | 10 | 0.7359 | - |
332
+ | 0.0199 | 20 | 0.758 | - |
333
+ | 0.0298 | 30 | 0.6144 | - |
334
+ | 0.0398 | 40 | 0.4337 | - |
335
+ | 0.0497 | 50 | 0.4198 | - |
336
+ | 0.0596 | 60 | 0.4157 | - |
337
+ | 0.0696 | 70 | 0.3955 | - |
338
+ | 0.0795 | 80 | 0.3569 | - |
339
+ | 0.0895 | 90 | 0.4447 | - |
340
+ | 0.0994 | 100 | 0.4593 | 0.9337 (+0.0222) |
341
+ | 0.1093 | 110 | 0.3853 | - |
342
+ | 0.1193 | 120 | 0.3984 | - |
343
+ | 0.1292 | 130 | 0.3537 | - |
344
+ | 0.1392 | 140 | 0.3597 | - |
345
+ | 0.1491 | 150 | 0.3911 | - |
346
+ | 0.1590 | 160 | 0.3823 | - |
347
+ | 0.1690 | 170 | 0.3785 | - |
348
+ | 0.1789 | 180 | 0.3466 | - |
349
+ | 0.1889 | 190 | 0.3503 | - |
350
+ | 0.1988 | 200 | 0.3502 | 0.9379 (+0.0264) |
351
+ | 0.2087 | 210 | 0.371 | - |
352
+ | 0.2187 | 220 | 0.3245 | - |
353
+ | 0.2286 | 230 | 0.3623 | - |
354
+ | 0.2386 | 240 | 0.3386 | - |
355
+ | 0.2485 | 250 | 0.3121 | - |
356
+ | 0.2584 | 260 | 0.3467 | - |
357
+ | 0.2684 | 270 | 0.3601 | - |
358
+ | 0.2783 | 280 | 0.3275 | - |
359
+ | 0.2883 | 290 | 0.3654 | - |
360
+ | 0.2982 | 300 | 0.4039 | 0.9408 (+0.0293) |
361
+ | 0.3082 | 310 | 0.3472 | - |
362
+ | 0.3181 | 320 | 0.3181 | - |
363
+ | 0.3280 | 330 | 0.3162 | - |
364
+ | 0.3380 | 340 | 0.3502 | - |
365
+ | 0.3479 | 350 | 0.3764 | - |
366
+ | 0.3579 | 360 | 0.304 | - |
367
+ | 0.3678 | 370 | 0.3384 | - |
368
+ | 0.3777 | 380 | 0.3125 | - |
369
+ | 0.3877 | 390 | 0.3166 | - |
370
+ | 0.3976 | 400 | 0.3199 | 0.9422 (+0.0307) |
371
+ | 0.4076 | 410 | 0.3078 | - |
372
+ | 0.4175 | 420 | 0.3694 | - |
373
+ | 0.4274 | 430 | 0.3401 | - |
374
+ | 0.4374 | 440 | 0.2958 | - |
375
+ | 0.4473 | 450 | 0.3127 | - |
376
+ | 0.4573 | 460 | 0.2986 | - |
377
+ | 0.4672 | 470 | 0.308 | - |
378
+ | 0.4771 | 480 | 0.2923 | - |
379
+ | 0.4871 | 490 | 0.3191 | - |
380
+ | 0.4970 | 500 | 0.3216 | 0.9433 (+0.0318) |
381
+ | 0.5070 | 510 | 0.3182 | - |
382
+ | 0.5169 | 520 | 0.3221 | - |
383
+ | 0.5268 | 530 | 0.2945 | - |
384
+ | 0.5368 | 540 | 0.3415 | - |
385
+ | 0.5467 | 550 | 0.2455 | - |
386
+ | 0.5567 | 560 | 0.3509 | - |
387
+ | 0.5666 | 570 | 0.3215 | - |
388
+ | 0.5765 | 580 | 0.2832 | - |
389
+ | 0.5865 | 590 | 0.2988 | - |
390
+ | 0.5964 | 600 | 0.3078 | 0.9415 (+0.0300) |
391
+ | 0.6064 | 610 | 0.3054 | - |
392
+ | 0.6163 | 620 | 0.298 | - |
393
+ | 0.6262 | 630 | 0.3364 | - |
394
+ | 0.6362 | 640 | 0.2635 | - |
395
+ | 0.6461 | 650 | 0.3286 | - |
396
+ | 0.6561 | 660 | 0.2909 | - |
397
+ | 0.6660 | 670 | 0.3172 | - |
398
+ | 0.6759 | 680 | 0.3081 | - |
399
+ | 0.6859 | 690 | 0.28 | - |
400
+ | 0.6958 | 700 | 0.3372 | 0.9423 (+0.0308) |
401
+ | 0.7058 | 710 | 0.3123 | - |
402
+ | 0.7157 | 720 | 0.3315 | - |
403
+ | 0.7256 | 730 | 0.2531 | - |
404
+ | 0.7356 | 740 | 0.3111 | - |
405
+ | 0.7455 | 750 | 0.265 | - |
406
+ | 0.7555 | 760 | 0.289 | - |
407
+ | 0.7654 | 770 | 0.2935 | - |
408
+ | 0.7753 | 780 | 0.2625 | - |
409
+ | 0.7853 | 790 | 0.3538 | - |
410
+ | 0.7952 | 800 | 0.2851 | 0.9434 (+0.0319) |
411
+ | 0.8052 | 810 | 0.2869 | - |
412
+ | 0.8151 | 820 | 0.2983 | - |
413
+ | 0.8250 | 830 | 0.3163 | - |
414
+ | 0.8350 | 840 | 0.2942 | - |
415
+ | 0.8449 | 850 | 0.2479 | - |
416
+ | 0.8549 | 860 | 0.2932 | - |
417
+ | 0.8648 | 870 | 0.2712 | - |
418
+ | 0.8748 | 880 | 0.2732 | - |
419
+ | 0.8847 | 890 | 0.2903 | - |
420
+ | 0.8946 | 900 | 0.3099 | 0.9434 (+0.0319) |
421
+ | 0.9046 | 910 | 0.37 | - |
422
+ | 0.9145 | 920 | 0.3039 | - |
423
+ | 0.9245 | 930 | 0.2882 | - |
424
+ | 0.9344 | 940 | 0.3288 | - |
425
+ | 0.9443 | 950 | 0.3212 | - |
426
+ | 0.9543 | 960 | 0.2967 | - |
427
+ | 0.9642 | 970 | 0.2976 | - |
428
+ | 0.9742 | 980 | 0.3474 | - |
429
+ | 0.9841 | 990 | 0.2698 | - |
430
+ | 0.9940 | 1000 | 0.2718 | 0.9446 (+0.0331) |
431
+ | 1.0040 | 1010 | 0.3122 | - |
432
+ | 1.0139 | 1020 | 0.2486 | - |
433
+ | 1.0239 | 1030 | 0.233 | - |
434
+ | 1.0338 | 1040 | 0.2538 | - |
435
+ | 1.0437 | 1050 | 0.277 | - |
436
+ | 1.0537 | 1060 | 0.2661 | - |
437
+ | 1.0636 | 1070 | 0.2525 | - |
438
+ | 1.0736 | 1080 | 0.2607 | - |
439
+ | 1.0835 | 1090 | 0.2509 | - |
440
+ | 1.0934 | 1100 | 0.2417 | 0.9436 (+0.0321) |
441
+ | 1.1034 | 1110 | 0.2782 | - |
442
+ | 1.1133 | 1120 | 0.2566 | - |
443
+ | 1.1233 | 1130 | 0.2352 | - |
444
+ | 1.1332 | 1140 | 0.2301 | - |
445
+ | 1.1431 | 1150 | 0.2302 | - |
446
+ | 1.1531 | 1160 | 0.26 | - |
447
+ | 1.1630 | 1170 | 0.2607 | - |
448
+ | 1.1730 | 1180 | 0.2355 | - |
449
+ | 1.1829 | 1190 | 0.3052 | - |
450
+ | 1.1928 | 1200 | 0.3072 | 0.9442 (+0.0327) |
451
+ | 1.2028 | 1210 | 0.2847 | - |
452
+ | 1.2127 | 1220 | 0.2387 | - |
453
+ | 1.2227 | 1230 | 0.2332 | - |
454
+ | 1.2326 | 1240 | 0.2687 | - |
455
+ | 1.2425 | 1250 | 0.2339 | - |
456
+ | 1.2525 | 1260 | 0.273 | - |
457
+ | 1.2624 | 1270 | 0.2124 | - |
458
+ | 1.2724 | 1280 | 0.2491 | - |
459
+ | 1.2823 | 1290 | 0.2629 | - |
460
+ | 1.2922 | 1300 | 0.2262 | 0.9443 (+0.0329) |
461
+ | 1.3022 | 1310 | 0.2502 | - |
462
+ | 1.3121 | 1320 | 0.277 | - |
463
+ | 1.3221 | 1330 | 0.2526 | - |
464
+ | 1.3320 | 1340 | 0.2484 | - |
465
+ | 1.3419 | 1350 | 0.2675 | - |
466
+ | 1.3519 | 1360 | 0.2658 | - |
467
+ | 1.3618 | 1370 | 0.2352 | - |
468
+ | 1.3718 | 1380 | 0.2476 | - |
469
+ | 1.3817 | 1390 | 0.2042 | - |
470
+ | 1.3917 | 1400 | 0.2693 | 0.9435 (+0.0320) |
471
+ | 1.4016 | 1410 | 0.2817 | - |
472
+ | 1.4115 | 1420 | 0.254 | - |
473
+ | 1.4215 | 1430 | 0.2424 | - |
474
+ | 1.4314 | 1440 | 0.2244 | - |
475
+ | 1.4414 | 1450 | 0.254 | - |
476
+ | 1.4513 | 1460 | 0.2624 | - |
477
+ | 1.4612 | 1470 | 0.2767 | - |
478
+ | 1.4712 | 1480 | 0.211 | - |
479
+ | 1.4811 | 1490 | 0.2861 | - |
480
+ | 1.4911 | 1500 | 0.2372 | 0.9433 (+0.0318) |
481
+ | 1.5010 | 1510 | 0.2428 | - |
482
+ | 1.5109 | 1520 | 0.2109 | - |
483
+ | 1.5209 | 1530 | 0.2658 | - |
484
+ | 1.5308 | 1540 | 0.2577 | - |
485
+ | 1.5408 | 1550 | 0.2243 | - |
486
+ | 1.5507 | 1560 | 0.2454 | - |
487
+ | 1.5606 | 1570 | 0.2256 | - |
488
+ | 1.5706 | 1580 | 0.2567 | - |
489
+ | 1.5805 | 1590 | 0.2729 | - |
490
+ | 1.5905 | 1600 | 0.2506 | 0.9431 (+0.0316) |
491
+ | 1.6004 | 1610 | 0.2112 | - |
492
+ | 1.6103 | 1620 | 0.2903 | - |
493
+ | 1.6203 | 1630 | 0.2383 | - |
494
+ | 1.6302 | 1640 | 0.2839 | - |
495
+ | 1.6402 | 1650 | 0.2695 | - |
496
+ | 1.6501 | 1660 | 0.2147 | - |
497
+ | 1.6600 | 1670 | 0.2577 | - |
498
+ | 1.6700 | 1680 | 0.2107 | - |
499
+ | 1.6799 | 1690 | 0.2474 | - |
500
+ | 1.6899 | 1700 | 0.2756 | 0.9448 (+0.0333) |
501
+ | 1.6998 | 1710 | 0.2365 | - |
502
+ | 1.7097 | 1720 | 0.2816 | - |
503
+ | 1.7197 | 1730 | 0.2686 | - |
504
+ | 1.7296 | 1740 | 0.2443 | - |
505
+ | 1.7396 | 1750 | 0.2574 | - |
506
+ | 1.7495 | 1760 | 0.2448 | - |
507
+ | 1.7594 | 1770 | 0.2536 | - |
508
+ | 1.7694 | 1780 | 0.242 | - |
509
+ | 1.7793 | 1790 | 0.2563 | - |
510
+ | **1.7893** | **1800** | **0.2169** | **0.9449 (+0.0334)** |
511
+ | 1.7992 | 1810 | 0.2194 | - |
512
+ | 1.8091 | 1820 | 0.2208 | - |
513
+ | 1.8191 | 1830 | 0.2335 | - |
514
+ | 1.8290 | 1840 | 0.2224 | - |
515
+ | 1.8390 | 1850 | 0.2698 | - |
516
+ | 1.8489 | 1860 | 0.2369 | - |
517
+ | 1.8588 | 1870 | 0.2244 | - |
518
+ | 1.8688 | 1880 | 0.2525 | - |
519
+ | 1.8787 | 1890 | 0.2286 | - |
520
+ | 1.8887 | 1900 | 0.2632 | 0.9448 (+0.0333) |
521
+ | 1.8986 | 1910 | 0.2173 | - |
522
+ | 1.9085 | 1920 | 0.2541 | - |
523
+ | 1.9185 | 1930 | 0.3137 | - |
524
+ | 1.9284 | 1940 | 0.2454 | - |
525
+ | 1.9384 | 1950 | 0.2299 | - |
526
+ | 1.9483 | 1960 | 0.2564 | - |
527
+ | 1.9583 | 1970 | 0.217 | - |
528
+ | 1.9682 | 1980 | 0.2135 | - |
529
+ | 1.9781 | 1990 | 0.2399 | - |
530
+ | 1.9881 | 2000 | 0.2174 | 0.9448 (+0.0333) |
531
+ | 1.9980 | 2010 | 0.2734 | - |
532
+ | -1 | -1 | - | 0.9449 (+0.0334) |
533
+
534
+ * The bold row denotes the saved checkpoint.
535
+ </details>
536
+
537
+ ### Framework Versions
538
+ - Python: 3.14.0
539
+ - Sentence Transformers: 5.2.0
540
+ - Transformers: 4.57.3
541
+ - PyTorch: 2.9.1+cu128
542
+ - Accelerate: 1.12.0
543
+ - Datasets: 4.4.1
544
+ - Tokenizers: 0.22.1
545
+
546
+ ## Citation
547
+
548
+ ### BibTeX
549
+
550
+ #### Sentence Transformers
551
+ ```bibtex
552
+ @inproceedings{reimers-2019-sentence-bert,
553
+ title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
554
+ author = "Reimers, Nils and Gurevych, Iryna",
555
+ booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
556
+ month = "11",
557
+ year = "2019",
558
+ publisher = "Association for Computational Linguistics",
559
+ url = "https://arxiv.org/abs/1908.10084",
560
+ }
561
+ ```
562
+
563
+ <!--
564
+ ## Glossary
565
+
566
+ *Clearly define terms in order to be accessible across audiences.*
567
+ -->
568
+
569
+ <!--
570
+ ## Model Card Authors
571
+
572
+ *Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*
573
+ -->
574
+
575
+ <!--
576
+ ## Model Card Contact
577
+
578
+ *Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*
579
+ -->
config.json ADDED
@@ -0,0 +1,37 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "BertForSequenceClassification"
4
+ ],
5
+ "attention_probs_dropout_prob": 0.1,
6
+ "bos_token_id": 0,
7
+ "classifier_dropout": null,
8
+ "dtype": "float32",
9
+ "eos_token_id": 2,
10
+ "hidden_act": "gelu",
11
+ "hidden_dropout_prob": 0.1,
12
+ "hidden_size": 384,
13
+ "id2label": {
14
+ "0": "LABEL_0"
15
+ },
16
+ "initializer_range": 0.02,
17
+ "intermediate_size": 1536,
18
+ "label2id": {
19
+ "LABEL_0": 0
20
+ },
21
+ "layer_norm_eps": 1e-12,
22
+ "max_position_embeddings": 512,
23
+ "model_type": "bert",
24
+ "num_attention_heads": 12,
25
+ "num_hidden_layers": 12,
26
+ "pad_token_id": 1,
27
+ "position_embedding_type": "absolute",
28
+ "sentence_transformers": {
29
+ "activation_fn": "torch.nn.modules.activation.Sigmoid",
30
+ "version": "5.2.0"
31
+ },
32
+ "tokenizer_class": "XLMRobertaTokenizer",
33
+ "transformers_version": "4.57.3",
34
+ "type_vocab_size": 2,
35
+ "use_cache": true,
36
+ "vocab_size": 250037
37
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:356e13c7cca6e5fd47dd7d0c43d4208acdacb8be2abe685e680a7527ae9e47db
3
+ size 470640124
special_tokens_map.json ADDED
@@ -0,0 +1,51 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": {
3
+ "content": "<s>",
4
+ "lstrip": false,
5
+ "normalized": false,
6
+ "rstrip": false,
7
+ "single_word": false
8
+ },
9
+ "cls_token": {
10
+ "content": "<s>",
11
+ "lstrip": false,
12
+ "normalized": false,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "eos_token": {
17
+ "content": "</s>",
18
+ "lstrip": false,
19
+ "normalized": false,
20
+ "rstrip": false,
21
+ "single_word": false
22
+ },
23
+ "mask_token": {
24
+ "content": "<mask>",
25
+ "lstrip": true,
26
+ "normalized": false,
27
+ "rstrip": false,
28
+ "single_word": false
29
+ },
30
+ "pad_token": {
31
+ "content": "<pad>",
32
+ "lstrip": false,
33
+ "normalized": false,
34
+ "rstrip": false,
35
+ "single_word": false
36
+ },
37
+ "sep_token": {
38
+ "content": "</s>",
39
+ "lstrip": false,
40
+ "normalized": false,
41
+ "rstrip": false,
42
+ "single_word": false
43
+ },
44
+ "unk_token": {
45
+ "content": "<unk>",
46
+ "lstrip": false,
47
+ "normalized": false,
48
+ "rstrip": false,
49
+ "single_word": false
50
+ }
51
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cf44dabfaa82b1276a7af64a2ea2c76c047d560cf7bfb5711d6135382372c93d
3
+ size 17083153
tokenizer_config.json ADDED
@@ -0,0 +1,63 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "added_tokens_decoder": {
3
+ "0": {
4
+ "content": "<s>",
5
+ "lstrip": false,
6
+ "normalized": false,
7
+ "rstrip": false,
8
+ "single_word": false,
9
+ "special": true
10
+ },
11
+ "1": {
12
+ "content": "<pad>",
13
+ "lstrip": false,
14
+ "normalized": false,
15
+ "rstrip": false,
16
+ "single_word": false,
17
+ "special": true
18
+ },
19
+ "2": {
20
+ "content": "</s>",
21
+ "lstrip": false,
22
+ "normalized": false,
23
+ "rstrip": false,
24
+ "single_word": false,
25
+ "special": true
26
+ },
27
+ "3": {
28
+ "content": "<unk>",
29
+ "lstrip": false,
30
+ "normalized": false,
31
+ "rstrip": false,
32
+ "single_word": false,
33
+ "special": true
34
+ },
35
+ "250001": {
36
+ "content": "<mask>",
37
+ "lstrip": true,
38
+ "normalized": false,
39
+ "rstrip": false,
40
+ "single_word": false,
41
+ "special": true
42
+ }
43
+ },
44
+ "bos_token": "<s>",
45
+ "clean_up_tokenization_spaces": false,
46
+ "cls_token": "<s>",
47
+ "eos_token": "</s>",
48
+ "extra_special_tokens": {},
49
+ "mask_token": "<mask>",
50
+ "max_length": 512,
51
+ "model_max_length": 512,
52
+ "pad_to_multiple_of": null,
53
+ "pad_token": "<pad>",
54
+ "pad_token_type_id": 0,
55
+ "padding_side": "right",
56
+ "sep_token": "</s>",
57
+ "sp_model_kwargs": {},
58
+ "stride": 0,
59
+ "tokenizer_class": "XLMRobertaTokenizer",
60
+ "truncation_side": "right",
61
+ "truncation_strategy": "longest_first",
62
+ "unk_token": "<unk>"
63
+ }