Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Paper • 1908.10084 • Published • 17
How to use Nessrine9/Finetune2-MiniLM-L12-v2 with sentence-transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("Nessrine9/Finetune2-MiniLM-L12-v2")
sentences = [
"A woman wearing a yellow shirt is holding a plate which contains a piece of cake.",
"The woman in the yellow shirt might have cut the cake and placed it on the plate.",
"Male bicyclists compete in the Tour de France.",
"The man is walking"
]
embeddings = model.encode(sentences)
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [4, 4]This is a sentence-transformers model finetuned from sentence-transformers/all-MiniLM-L12-v2. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
SentenceTransformer(
(0): Transformer({'max_seq_length': 128, 'do_lower_case': False}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("Nessrine9/Finetune2-MiniLM-L12-v2")
# Run inference
sentences = [
'A man fishing in a pointy blue boat on a river lined with palm trees.',
'The man is with friends.',
'A man rubs his bald head.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]
snli-devEmbeddingSimilarityEvaluator| Metric | Value |
|---|---|
| pearson_cosine | 0.5003 |
| spearman_cosine | 0.4919 |
| pearson_manhattan | 0.4752 |
| spearman_manhattan | 0.4917 |
| pearson_euclidean | 0.476 |
| spearman_euclidean | 0.4919 |
| pearson_dot | 0.5003 |
| spearman_dot | 0.4919 |
| pearson_max | 0.5003 |
| spearman_max | 0.4919 |
sentence_0, sentence_1, and label| sentence_0 | sentence_1 | label | |
|---|---|---|---|
| type | string | string | float |
| details |
|
|
|
| sentence_0 | sentence_1 | label |
|---|---|---|
Three men in an art gallery posing for the camera. |
Paintings are nearby. |
0.5 |
A shirtless man wearing a vest walks on a stage with his arms up. |
The man is about to perform. |
0.5 |
The man is walking outside near a rocky river. |
The man is walking |
0.0 |
CosineSimilarityLoss with these parameters:{
"loss_fct": "torch.nn.modules.loss.MSELoss"
}
eval_strategy: stepsper_device_train_batch_size: 16per_device_eval_batch_size: 16num_train_epochs: 4fp16: Truemulti_dataset_batch_sampler: round_robinoverwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 16per_device_eval_batch_size: 16per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 5e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1num_train_epochs: 4max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.0warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falseuse_ipex: Falsebf16: Falsefp16: Truefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Falseignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}deepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torchoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Falsehub_always_push: Falsegradient_checkpointing: Falsegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseeval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters: auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Nonedispatch_batches: Nonesplit_batches: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: Falseneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseeval_use_gather_object: Falsebatch_sampler: batch_samplermulti_dataset_batch_sampler: round_robin| Epoch | Step | Training Loss | snli-dev_spearman_max |
|---|---|---|---|
| 0.08 | 500 | 0.1842 | 0.3333 |
| 0.16 | 1000 | 0.1489 | 0.3449 |
| 0.24 | 1500 | 0.1427 | 0.3633 |
| 0.32 | 2000 | 0.1391 | 0.3854 |
| 0.4 | 2500 | 0.1401 | 0.4015 |
| 0.48 | 3000 | 0.139 | 0.3982 |
| 0.56 | 3500 | 0.1352 | 0.4327 |
| 0.64 | 4000 | 0.1319 | 0.4262 |
| 0.72 | 4500 | 0.1336 | 0.4034 |
| 0.8 | 5000 | 0.1321 | 0.4021 |
| 0.88 | 5500 | 0.1309 | 0.4294 |
| 0.96 | 6000 | 0.1271 | 0.4198 |
| 1.0 | 6250 | - | 0.4317 |
| 1.04 | 6500 | 0.132 | 0.4445 |
| 1.12 | 7000 | 0.1296 | 0.4509 |
| 1.2 | 7500 | 0.1236 | 0.4559 |
| 1.28 | 8000 | 0.1257 | 0.4542 |
| 1.3600 | 8500 | 0.1236 | 0.4507 |
| 1.44 | 9000 | 0.1277 | 0.4540 |
| 1.52 | 9500 | 0.1249 | 0.4664 |
| 1.6 | 10000 | 0.1208 | 0.4418 |
| 1.6800 | 10500 | 0.1228 | 0.4457 |
| 1.76 | 11000 | 0.1212 | 0.4222 |
| 1.8400 | 11500 | 0.1203 | 0.4507 |
| 1.92 | 12000 | 0.119 | 0.4572 |
| 2.0 | 12500 | 0.1196 | 0.4667 |
| 2.08 | 13000 | 0.1194 | 0.4733 |
| 2.16 | 13500 | 0.1172 | 0.4786 |
| 2.24 | 14000 | 0.1172 | 0.4765 |
| 2.32 | 14500 | 0.1145 | 0.4717 |
| 2.4 | 15000 | 0.1167 | 0.4803 |
| 2.48 | 15500 | 0.1177 | 0.4678 |
| 2.56 | 16000 | 0.1162 | 0.4805 |
| 2.64 | 16500 | 0.1137 | 0.4780 |
| 2.7200 | 17000 | 0.1153 | 0.4788 |
| 2.8 | 17500 | 0.115 | 0.4784 |
| 2.88 | 18000 | 0.1128 | 0.4864 |
| 2.96 | 18500 | 0.11 | 0.4812 |
| 3.0 | 18750 | - | 0.4823 |
| 3.04 | 19000 | 0.1136 | 0.4900 |
| 3.12 | 19500 | 0.1135 | 0.4897 |
| 3.2 | 20000 | 0.1094 | 0.4856 |
| 3.2800 | 20500 | 0.1108 | 0.4889 |
| 3.36 | 21000 | 0.1083 | 0.4909 |
| 3.44 | 21500 | 0.1133 | 0.4892 |
| 3.52 | 22000 | 0.1106 | 0.4910 |
| 3.6 | 22500 | 0.1079 | 0.4888 |
| 3.68 | 23000 | 0.1091 | 0.4890 |
| 3.76 | 23500 | 0.1079 | 0.4822 |
| 3.84 | 24000 | 0.1087 | 0.4887 |
| 3.92 | 24500 | 0.1066 | 0.4926 |
| 4.0 | 25000 | 0.1069 | 0.4919 |
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
Base model
microsoft/MiniLM-L12-H384-uncased