SentenceTransformer based on nomic-ai/nomic-embed-text-v1.5

This is a sentence-transformers model finetuned from nomic-ai/nomic-embed-text-v1.5. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: nomic-ai/nomic-embed-text-v1.5
  • Maximum Sequence Length: 384 tokens
  • Output Dimensionality: 768 dimensions
  • Similarity Function: Cosine Similarity

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 384, 'do_lower_case': False, 'architecture': 'NomicBertModel'})
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    'Procedure for the detection of organic UV filters. Title: Procedure for the detection of organic UV filters\n\nAbstract: \n\nBusiness Description: The invention concerns an optimized analytical procedure for the detection and quantification of organic UV filters, emerging contaminants with potential harmful effects on aquatic ecosystems. The aim is to develop a simple, rapid, portable, and cost-effective voltammetric method, as an alternative to traditional chromatographic techniques, for environmental water monitoring and the characterization of sunscreen products.\n\nTech Features: The technology is based on an innovative method to detect and quantify UV filters such as octocrylene, oxybenzone, and octinoxate, applicable to aqueous matrices and sunscreen products. After a simple sample treatment (concentration for water samples or extraction for creams), the analysis is performed using an electrochemical sensor that measures specific electrical signals: each substance produces a characteristic “peak,” allowing it to be identified and quantified simultaneously. Compared to traditional laboratory methods, such as liquid chromatography or gas chromatography, the proposed solution is more cost-effective, portable, rapid, uses fewer solvents, and enables on-site analysis while maintaining high sensitivity and repeatability.\n\nApplications: Environmental monitoring of marine, lake, and river waters; Control of wastewater and treatment plants; Rapid on-site analysis of emerging contaminants; Characterization of sunscreens and oils; Quality control in the cosmetic sector.\n\nAdvantages: Portable method suitable for field analysis; Lower costs compared to chromatographic techniques; Reduced use of organic solvents; Fast and simple analyses; Simultaneous quantification of multiple UV filters; Good correlation with conventional analytical methods.',
    'Procedure for the detection of organic UV filters. Title: Procedure for the detection of organic UV filters\n\nAbstract: \n\nBusiness Description: The invention concerns an optimized analytical procedure for the detection and quantification of organic UV filters, emerging contaminants with potential harmful effects on aquatic ecosystems. The aim is to develop a simple, rapid, portable, and cost-effective voltammetric method, as an alternative to traditional chromatographic techniques, for environmental water monitoring and the characterization of sunscreen products.\n\nTech Features: The technology is based on an innovative method to detect and quantify UV filters such as octocrylene, oxybenzone, and octinoxate, applicable to aqueous matrices and sunscreen products. After a simple sample treatment (concentration for water samples or extraction for creams), the analysis is performed using an electrochemical sensor that measures specific electrical signals: each substance produces a characteristic “peak,” allowing it to be identified and quantified simultaneously. Compared to traditional laboratory methods, such as liquid chromatography or gas chromatography, the proposed solution is more cost-effective, portable, rapid, uses fewer solvents, and enables on-site analysis while maintaining high sensitivity and repeatability.\n\nApplications: Environmental monitoring of marine, lake, and river waters; Control of wastewater and treatment plants; Rapid on-site analysis of emerging contaminants; Characterization of sunscreens and oils; Quality control in the cosmetic sector.\n\nAdvantages: Portable method suitable for field analysis; Lower costs compared to chromatographic techniques; Reduced use of organic solvents; Fast and simple analyses; Simultaneous quantification of multiple UV filters; Good correlation with conventional analytical methods.',
    'PRODUCTION OF HIGH ORGANOLEPTIC AND NUTRITIONAL VALUE OLIVE OIL. Title: PRODUCTION OF HIGH ORGANOLEPTIC AND NUTRITIONAL VALUE OLIVE OIL\n\nAbstract: \n\nBusiness Description: The procedure involves the use of a non-toxic and organoleptically inert cryogen in the process of extracting oil from olives. In addition to high yields, it guarantees the production of olive oils, especially extra virgin, enriched in cellular compounds extracted from the fruit and, in particular, in components with aromatic and antioxidant activity, with a consequent significant increase in their organoleptic and nutritional quality. The unmistakable characteristics of the oils most recognizable by the consumer are closely linked to the raw material, the type of olives processed and their production area.\n\nTech Features: The proposed procedure provides the use of "carbonic snow", the carbon dioxide (CO 2 ) in a solid state for the extraction of olive oil. The solid CO 2 causes the formation of ice crystals in the freezing fruit, which in turn determines the collapse of the cellular structure of the pulp. This facilitates the release of substances and their transfer into the oil, which results rich in cellular metabolites with high biological value. The gaseous CO 2 , being heavier than air, tends to remain above the olive paste, creating a gaseous layer able to avoid direct contact with the air oxygen and to preserve the cellular constituents from oxidative degradation. The method makes economically sustainable early harvesting of the olives: the olives less mature will be richer in water and bioactive components (polyphenols, tocopherols); then, the early harvesting limits the damage caused by attacks of Bactrocera oleae (the olive fly), one of the most feared adversities by producers of the sector, able to significantly affect both the yield and the quality of the oil produced.\n\nApplications: Use in mills.\n\nAdvantages: Higher yield, on average 9% more (17.4 kg of product instead of 16 kg per quintal of olives); Better nutritional quality (e.g. 6% more vitamin E); Greater resistance to oxidative processes; Production of oil richer in antioxidants and aromatic components; Longer shelf life than that of oil obtained using conventional technologies.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 1.0000, 0.2489],
#         [1.0000, 1.0000, 0.2489],
#         [0.2489, 0.2489, 1.0000]])

Evaluation

Metrics

Information Retrieval

Metric Value
cosine_accuracy@1 0.0433
cosine_accuracy@3 0.0967
cosine_accuracy@5 0.1412
cosine_accuracy@10 0.2087
cosine_precision@1 0.0433
cosine_precision@3 0.0322
cosine_precision@5 0.0282
cosine_precision@10 0.0209
cosine_recall@1 0.0433
cosine_recall@3 0.0967
cosine_recall@5 0.1412
cosine_recall@10 0.2087
cosine_ndcg@10 0.1138
cosine_mrr@10 0.0849
cosine_map@100 0.0953

Training Details

Training Dataset

Unnamed Dataset

  • Size: 6,288 training samples
  • Columns: sentence_0 and sentence_1
  • Approximate statistics based on the first 1000 samples:
    sentence_0 sentence_1
    type string string
    details
    • min: 4 tokens
    • mean: 12.3 tokens
    • max: 35 tokens
    • min: 8 tokens
    • mean: 12.08 tokens
    • max: 36 tokens
  • Samples:
    sentence_0 sentence_1
    Epigenetic regulation system for the control of target gene expression Applicant/Organization: FONDAZIONE ISTITUTO ITALIANO DI TECNOLOGIA
    System integrating a membrane humidifier and an adsorption-based storage for polymer membrane hydrogen fuel cell applications. Technical Classification: H01M
    Superconducting bipolar thermoelectric memory Technical Classification: G11C_11
  • Loss: MultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 16
  • num_train_epochs: 10
  • eval_strategy: steps
  • per_device_eval_batch_size: 16
  • multi_dataset_batch_sampler: round_robin

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 16
  • num_train_epochs: 10
  • max_steps: -1
  • learning_rate: 5e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: False
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: trackio
  • eval_strategy: steps
  • per_device_eval_batch_size: 16
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: []
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • warmup_ratio: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: round_robin
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss mnrl-val_cosine_ndcg@10
1.0 393 - 0.0695
1.0178 400 - 0.0697
1.2723 500 2.2361 -
2.0 786 - 0.1032
2.0356 800 - 0.1079
2.5445 1000 1.5522 -
3.0 1179 - 0.1043
3.0534 1200 - 0.0946
3.8168 1500 1.0458 -
4.0 1572 - 0.1138

Framework Versions

  • Python: 3.12.3
  • Sentence Transformers: 5.3.0
  • Transformers: 5.3.0
  • PyTorch: 2.10.0+cu128
  • Accelerate: 1.13.0
  • Datasets: 4.7.0
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}
Downloads last month
30
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GiacomoSignorile/Nomic-v1.5-FineTuned-for-Patent

Finetuned
(41)
this model

Papers for GiacomoSignorile/Nomic-v1.5-FineTuned-for-Patent

Evaluation results