SentenceTransformer based on google-bert/bert-base-uncased

This is a sentence-transformers model finetuned from google-bert/bert-base-uncased. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for retrieval.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: google-bert/bert-base-uncased
  • Maximum Sequence Length: 128 tokens
  • Output Dimensionality: 768 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("tomaarsen/bert-base-uncased-qqp-cross-domain")
# Run inference
sentences = [
    "What is the best code troll you've ever seen?",
    "What is the best favicon you've ever seen?",
    'What is it feel like to die?',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.4664, 0.0131],
#         [0.4664, 1.0000, 0.0492],
#         [0.0131, 0.0492, 1.0000]])

Evaluation

Metrics

Binary Classification

Metric Value
cosine_accuracy 0.753
cosine_accuracy_threshold 0.7228
cosine_f1 0.6999
cosine_f1_threshold 0.636
cosine_precision 0.5904
cosine_recall 0.8593
cosine_ap 0.7111
cosine_mcc 0.5009

Training Details

Training Dataset

Unnamed Dataset

  • Size: 394,290 training samples
  • Columns: sentence1, sentence2, and score
  • Approximate statistics based on the first 100 samples:
    sentence1 sentence2 score
    type string string float
    modality text text
    details
    • min: 7 tokens
    • mean: 17.88 tokens
    • max: 62 tokens
    • min: 6 tokens
    • mean: 16.96 tokens
    • max: 50 tokens
    • min: 0.03
    • mean: 0.57
    • max: 0.93
  • Samples:
    sentence1 sentence2 score
    According to garuda purana, how can a child conceived through IVF be differentiated from an inauspicious soul? What differentiates a 10x doctor from the rest? 0.10419108718633652
    Does diabetes cause erectile dysfunction? Does stress is one of the cause erectile dysfunction? 0.4059196412563324
    What are the pros and cons of using Python vs. Java? What are the advantages and disadvantages of Python over Java? 0.7368191480636597
  • Loss: CosineSimilarityLoss with these parameters:
    {
        "loss_fct": "torch.nn.modules.loss.MSELoss",
        "cos_score_transformation": "torch.nn.modules.linear.Identity"
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 16
  • num_train_epochs: 1
  • warmup_steps: 0.1
  • per_device_eval_batch_size: 16

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 16
  • num_train_epochs: 1
  • max_steps: -1
  • learning_rate: 5e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0.1
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1.0
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: False
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 16
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: []
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • warmup_ratio: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss cosine_ap
0.0500 1233 0.0152 -
0.1000 2465 - 0.7075
0.1001 2466 0.0070 -
0.1501 3699 0.0074 -
0.2000 4930 - 0.7079
0.2001 4932 0.0066 -
0.2502 6165 0.0065 -
0.3001 7395 - 0.7102
0.3002 7398 0.0061 -
0.3502 8631 0.0059 -
0.4001 9860 - 0.7060
0.4003 9864 0.0057 -
0.4503 11097 0.0055 -
0.5001 12325 - 0.7086
0.5003 12330 0.0053 -
0.5504 13563 0.0050 -
0.6001 14790 - 0.7113
0.6004 14796 0.0049 -
0.6504 16029 0.0047 -
0.7002 17255 - 0.7087
0.7005 17262 0.0046 -
0.7505 18495 0.0044 -
0.8002 19720 - 0.7082
0.8005 19728 0.0042 -
0.8506 20961 0.0041 -
0.9002 22185 - 0.7103
0.9006 22194 0.0039 -
0.9506 23427 0.0039 -
1.0 24644 - 0.7111

Training Time

  • Training: 46.4 minutes
  • Evaluation: 59.1 seconds
  • Total: 47.4 minutes

Framework Versions

  • Python: 3.11.6
  • Sentence Transformers: 5.6.0.dev0
  • Transformers: 5.8.0.dev0
  • PyTorch: 2.10.0+cu128
  • Accelerate: 1.13.0.dev0
  • Datasets: 4.8.4
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}
Downloads last month
74
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tomaarsen/bert-base-uncased-qqp-cross-domain

Finetuned
(6877)
this model

Paper for tomaarsen/bert-base-uncased-qqp-cross-domain

Evaluation results