CrossEncoder based on nreimers/MiniLM-L3-H384-uncased

This is a Cross Encoder model finetuned from nreimers/MiniLM-L3-H384-uncased using the sentence-transformers library. It computes scores for pairs of texts, which can be used for text reranking and semantic search.

Model Details

Model Description

Model Sources

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import CrossEncoder

# Download from the 🤗 Hub
model = CrossEncoder("lucasflins/CE-MiniLM-L3-H384-uncased-RankNet-20260207-052807")
# Get scores for pairs of texts
pairs = [
    ['escada 5', 'escada mor alumínio 5 degraus uso doméstico'],
    ['escada 5', 'escada alumínio 5 degraus mor'],
    ['escada 5', 'escada doméstica 05 degraus alumínio - real'],
    ['escada 5', 'escada alumínio 5 degraus - mor'],
    ['escada 5', 'escada de alumínio 5 degraus com fita de segurança mor'],
]
scores = model.predict(pairs)
print(scores.shape)
# (5,)

# Or rank different texts based on similarity to a single text
ranks = model.rank(
    'escada 5',
    [
        'escada mor alumínio 5 degraus uso doméstico',
        'escada alumínio 5 degraus mor',
        'escada doméstica 05 degraus alumínio - real',
        'escada alumínio 5 degraus - mor',
        'escada de alumínio 5 degraus com fita de segurança mor',
    ]
)
# [{'corpus_id': ..., 'score': ...}, {'corpus_id': ..., 'score': ...}, ...]

Training Details

Training Dataset

Unnamed Dataset

  • Size: 164,849 training samples
  • Columns: query, text, and label
  • Approximate statistics based on the first 1000 samples:
    query text label
    type string list list
    details
    • min: 3 characters
    • mean: 24.5 characters
    • max: 97 characters
    • min: 3 elements
    • mean: 92.48 elements
    • max: 450 elements
    • min: 3 elements
    • mean: 92.48 elements
    • max: 450 elements
  • Samples:
    query text label
    blusas femininas de verao ['blusa de malha feminina com decote v - manga curta - moda verão 2024', 'kit 2 regata feminina verão soltinhas confortáveis plus', 'blusa feminina manga princesa crepe verão básico em tendência 2025.', 'regata feminina malha fria senhoras plus verão', 'blusa cropped feminino tricot modal decote cruzado com alça verão', ...] [1.0, 0.9726344517752807, 0.8374873988889217, 0.8008303826878321, 0.7915794516443119, ...]
    boneca grande bebe reborn ['bebê reborn boneca princesa grande com vários acessórios', 'boneca reborn bebê realista grande com vestido florido', 'boneca bebe reborn grande 60cm corpo de pano com girafinha de pelúcia', 'bebe reborn loira realista grande 100% silicone 48cm banho', 'bebe reborn realista princesa grande 100% silicone 20 itens', ...] [1.0, 0.9116494092065375, 0.8928073970303235, 0.8879649368952169, 0.8731986703319667, ...]
    porta tempero de inox ['kit porta temperos e condimentos vidro tampa inox 6un 200ml', 'kit potes de vidro hermético 150ml porta tempero tampa inox 3 un', 'porta tempero condimento inox magnético imã cozinha 3 potes', 'kit 6 porta tempero condimento vidro suporte giratório inox cozinha pote organizador', 'porta tempero condimento giratório redondo 20 peças inox', ...] [0.9655845721090657, 0.8396654217445704, 0.7738607871762022, 0.6343641152653602, 0.566489008709851, ...]
  • Loss: RankNetLoss with these parameters:
    {
        "k": null,
        "sigma": 1.0,
        "eps": 1e-10,
        "reduction_log": "binary",
        "activation_fn": "torch.nn.modules.linear.Identity",
        "mini_batch_size": null
    }
    

Evaluation Dataset

Unnamed Dataset

  • Size: 54,809 evaluation samples
  • Columns: query, text, and label
  • Approximate statistics based on the first 1000 samples:
    query text label
    type string list list
    details
    • min: 3 characters
    • mean: 25.32 characters
    • max: 90 characters
    • min: 3 elements
    • mean: 91.67 elements
    • max: 511 elements
    • min: 3 elements
    • mean: 91.67 elements
    • max: 511 elements
  • Samples:
    query text label
    escada 5 ['escada mor alumínio 5 degraus uso doméstico', 'escada alumínio 5 degraus mor', 'escada doméstica 05 degraus alumínio - real', 'escada alumínio 5 degraus - mor', 'escada de alumínio 5 degraus com fita de segurança mor', ...] [1.0, 0.7096172820312041, 0.6436050498616379, 0.6233409190446045, 0.5943286560299376, ...]
    mascara lola ['lola cosmetics danos vorazes máscara de reparação intensiva', 'máscara disciplinante xapadinha 450g lola', 'máscara nutritiva lola bossa crespos e cachos 450g', 'máscara lola morte súbita 450gr', 'kit argan lola cosmetics - shampoo + máscara', ...] [1.0, 0.9644981507636873, 0.9554469702621406, 0.9550674725047994, 0.9500765764997154, ...]
    brinox roma ['chaleira em inox com indução e apito brinox roma 2,7 litros vanilla', 'chaleira roma inox 2,7l com apito de indução vanilla 4875/100 brinox', 'chaleira brinox inox com apito preto marble 2,7l roma'] [1.0, 0.29334574198355634, 0.0]
  • Loss: RankNetLoss with these parameters:
    {
        "k": null,
        "sigma": 1.0,
        "eps": 1e-10,
        "reduction_log": "binary",
        "activation_fn": "torch.nn.modules.linear.Identity",
        "mini_batch_size": null
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • eval_strategy: epoch
  • per_device_train_batch_size: 64
  • per_device_eval_batch_size: 16
  • num_train_epochs: 10
  • warmup_ratio: 0.1
  • log_level: info
  • tf32: True
  • load_best_model_at_end: True
  • push_to_hub: True
  • hub_strategy: end
  • hub_private_repo: True
  • hub_always_push: True
  • eval_on_start: True

All Hyperparameters

Click to expand
  • overwrite_output_dir: False
  • do_predict: False
  • eval_strategy: epoch
  • prediction_loss_only: True
  • per_device_train_batch_size: 64
  • per_device_eval_batch_size: 16
  • per_gpu_train_batch_size: None
  • per_gpu_eval_batch_size: None
  • gradient_accumulation_steps: 1
  • eval_accumulation_steps: None
  • torch_empty_cache_steps: None
  • learning_rate: 5e-05
  • weight_decay: 0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • max_grad_norm: 1.0
  • num_train_epochs: 10
  • max_steps: -1
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_ratio: 0.1
  • warmup_steps: 0
  • log_level: info
  • log_level_replica: warning
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • save_safetensors: True
  • save_on_each_node: False
  • save_only_model: False
  • restore_callback_states_from_checkpoint: False
  • no_cuda: False
  • use_cpu: False
  • use_mps_device: False
  • seed: 42
  • data_seed: None
  • jit_mode_eval: False
  • bf16: False
  • fp16: False
  • fp16_opt_level: O1
  • half_precision_backend: auto
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: True
  • local_rank: 0
  • ddp_backend: None
  • tpu_num_cores: None
  • tpu_metrics_debug: False
  • debug: []
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_prefetch_factor: None
  • past_index: -1
  • disable_tqdm: False
  • remove_unused_columns: True
  • label_names: None
  • load_best_model_at_end: True
  • ignore_data_skip: False
  • fsdp: []
  • fsdp_min_num_params: 0
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • fsdp_transformer_layer_cls_to_wrap: None
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • deepspeed: None
  • label_smoothing_factor: 0.0
  • optim: adamw_torch
  • optim_args: None
  • adafactor: False
  • group_by_length: False
  • length_column_name: length
  • project: huggingface
  • trackio_space_id: trackio
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • skip_memory_metrics: True
  • use_legacy_prediction_loop: False
  • push_to_hub: True
  • resume_from_checkpoint: None
  • hub_model_id: None
  • hub_strategy: end
  • hub_private_repo: True
  • hub_always_push: True
  • hub_revision: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • include_inputs_for_metrics: False
  • include_for_metrics: []
  • eval_do_concat_batches: True
  • fp16_backend: auto
  • push_to_hub_model_id: None
  • push_to_hub_organization: None
  • mp_parameters:
  • auto_find_batch_size: False
  • full_determinism: False
  • torchdynamo: None
  • ray_scope: last
  • ddp_timeout: 1800
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • include_tokens_per_second: False
  • include_num_input_tokens_seen: no
  • neftune_noise_alpha: None
  • optim_target_modules: None
  • batch_eval_metrics: False
  • eval_on_start: True
  • use_liger_kernel: False
  • liger_kernel_config: None
  • eval_use_gather_object: False
  • average_tokens_across_devices: True
  • prompts: None
  • batch_sampler: batch_sampler
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss Validation Loss
0 0 - 1.0000
1.0 2576 0.9894 0.9858
2.0 5152 0.9678 0.9775
3.0 7728 0.9525 0.9700
4.0 10304 0.9406 0.9628
5.0 12880 0.9309 0.9607
6.0 15456 0.9244 0.9584
7.0 18032 0.9188 0.9573
8.0 20608 0.9138 0.9571
9.0 23184 0.9114 0.9544
10.0 25760 0.9084 0.9543
  • The bold row denotes the saved checkpoint.

Framework Versions

  • Python: 3.12.11
  • Sentence Transformers: 5.2.0
  • Transformers: 4.57.6
  • PyTorch: 2.7.1+cu128
  • Accelerate: 1.9.0
  • Datasets: 4.1.1
  • Tokenizers: 0.22.1

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

RankNetLoss

@inproceedings{burges2005learning,
  title={Learning to Rank using Gradient Descent},
  author={Burges, Chris and Shaked, Tal and Renshaw, Erin and Lazier, Ari and Deeds, Matt and Hamilton, Nicole and Hullender, Greg},
  booktitle={Proceedings of the 22nd international conference on Machine learning},
  pages={89--96},
  year={2005}
}
Downloads last month
14
Safetensors
Model size
17.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lucasflins/CE-MiniLM-L3-H384-uncased-RankNet-20260207-052807

Finetuned
(5)
this model

Paper for lucasflins/CE-MiniLM-L3-H384-uncased-RankNet-20260207-052807