--- tags: - sentence-transformers - cross-encoder - reranker - generated_from_trainer - dataset_size:164849 - loss:RankNetLoss base_model: nreimers/MiniLM-L3-H384-uncased pipeline_tag: text-ranking library_name: sentence-transformers --- # CrossEncoder based on nreimers/MiniLM-L3-H384-uncased This is a [Cross Encoder](https://www.sbert.net/docs/cross_encoder/usage/usage.html) model finetuned from [nreimers/MiniLM-L3-H384-uncased](https://huggingface.co/nreimers/MiniLM-L3-H384-uncased) using the [sentence-transformers](https://www.SBERT.net) library. It computes scores for pairs of texts, which can be used for text reranking and semantic search. ## Model Details ### Model Description - **Model Type:** Cross Encoder - **Base model:** [nreimers/MiniLM-L3-H384-uncased](https://huggingface.co/nreimers/MiniLM-L3-H384-uncased) - **Maximum Sequence Length:** 64 tokens - **Number of Output Labels:** 1 label ### Model Sources - **Documentation:** [Sentence Transformers Documentation](https://sbert.net) - **Documentation:** [Cross Encoder Documentation](https://www.sbert.net/docs/cross_encoder/usage/usage.html) - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers) - **Hugging Face:** [Cross Encoders on Hugging Face](https://huggingface.co/models?library=sentence-transformers&other=cross-encoder) ## Usage ### Direct Usage (Sentence Transformers) First install the Sentence Transformers library: ```bash pip install -U sentence-transformers ``` Then you can load this model and run inference. ```python from sentence_transformers import CrossEncoder # Download from the 🤗 Hub model = CrossEncoder("lucasflins/CE-MiniLM-L3-H384-uncased-RankNet-20260207-052807") # Get scores for pairs of texts pairs = [ ['escada 5', 'escada mor alumínio 5 degraus uso doméstico'], ['escada 5', 'escada alumínio 5 degraus mor'], ['escada 5', 'escada doméstica 05 degraus alumínio - real'], ['escada 5', 'escada alumínio 5 degraus - mor'], ['escada 5', 'escada de alumínio 5 degraus com fita de segurança mor'], ] scores = model.predict(pairs) print(scores.shape) # (5,) # Or rank different texts based on similarity to a single text ranks = model.rank( 'escada 5', [ 'escada mor alumínio 5 degraus uso doméstico', 'escada alumínio 5 degraus mor', 'escada doméstica 05 degraus alumínio - real', 'escada alumínio 5 degraus - mor', 'escada de alumínio 5 degraus com fita de segurança mor', ] ) # [{'corpus_id': ..., 'score': ...}, {'corpus_id': ..., 'score': ...}, ...] ``` ## Training Details ### Training Dataset #### Unnamed Dataset * Size: 164,849 training samples * Columns: query, text, and label * Approximate statistics based on the first 1000 samples: | | query | text | label | |:--------|:---------------------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------------| | type | string | list | list | | details | | | | * Samples: | query | text | label | |:---------------------------------------|:-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:----------------------------------------------------------------------------------------------------------------------| | blusas femininas de verao | ['blusa de malha feminina com decote v - manga curta - moda verão 2024', 'kit 2 regata feminina verão soltinhas confortáveis plus', 'blusa feminina manga princesa crepe verão básico em tendência 2025.', 'regata feminina malha fria senhoras plus verão', 'blusa cropped feminino tricot modal decote cruzado com alça verão', ...] | [1.0, 0.9726344517752807, 0.8374873988889217, 0.8008303826878321, 0.7915794516443119, ...] | | boneca grande bebe reborn | ['bebê reborn boneca princesa grande com vários acessórios', 'boneca reborn bebê realista grande com vestido florido', 'boneca bebe reborn grande 60cm corpo de pano com girafinha de pelúcia', 'bebe reborn loira realista grande 100% silicone 48cm banho', 'bebe reborn realista princesa grande 100% silicone 20 itens', ...] | [1.0, 0.9116494092065375, 0.8928073970303235, 0.8879649368952169, 0.8731986703319667, ...] | | porta tempero de inox | ['kit porta temperos e condimentos vidro tampa inox 6un 200ml', 'kit potes de vidro hermético 150ml porta tempero tampa inox 3 un', 'porta tempero condimento inox magnético imã cozinha 3 potes', 'kit 6 porta tempero condimento vidro suporte giratório inox cozinha pote organizador', 'porta tempero condimento giratório redondo 20 peças inox', ...] | [0.9655845721090657, 0.8396654217445704, 0.7738607871762022, 0.6343641152653602, 0.566489008709851, ...] | * Loss: [RankNetLoss](https://sbert.net/docs/package_reference/cross_encoder/losses.html#ranknetloss) with these parameters: ```json { "k": null, "sigma": 1.0, "eps": 1e-10, "reduction_log": "binary", "activation_fn": "torch.nn.modules.linear.Identity", "mini_batch_size": null } ``` ### Evaluation Dataset #### Unnamed Dataset * Size: 54,809 evaluation samples * Columns: query, text, and label * Approximate statistics based on the first 1000 samples: | | query | text | label | |:--------|:----------------------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------------| | type | string | list | list | | details | | | | * Samples: | query | text | label | |:--------------------------|:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------------------| | escada 5 | ['escada mor alumínio 5 degraus uso doméstico', 'escada alumínio 5 degraus mor', 'escada doméstica 05 degraus alumínio - real', 'escada alumínio 5 degraus - mor', 'escada de alumínio 5 degraus com fita de segurança mor', ...] | [1.0, 0.7096172820312041, 0.6436050498616379, 0.6233409190446045, 0.5943286560299376, ...] | | mascara lola | ['lola cosmetics danos vorazes máscara de reparação intensiva', 'máscara disciplinante xapadinha 450g lola', 'máscara nutritiva lola bossa crespos e cachos 450g', 'máscara lola morte súbita 450gr', 'kit argan lola cosmetics - shampoo + máscara', ...] | [1.0, 0.9644981507636873, 0.9554469702621406, 0.9550674725047994, 0.9500765764997154, ...] | | brinox roma | ['chaleira em inox com indução e apito brinox roma 2,7 litros vanilla', 'chaleira roma inox 2,7l com apito de indução vanilla 4875/100 brinox', 'chaleira brinox inox com apito preto marble 2,7l roma'] | [1.0, 0.29334574198355634, 0.0] | * Loss: [RankNetLoss](https://sbert.net/docs/package_reference/cross_encoder/losses.html#ranknetloss) with these parameters: ```json { "k": null, "sigma": 1.0, "eps": 1e-10, "reduction_log": "binary", "activation_fn": "torch.nn.modules.linear.Identity", "mini_batch_size": null } ``` ### Training Hyperparameters #### Non-Default Hyperparameters - `eval_strategy`: epoch - `per_device_train_batch_size`: 64 - `per_device_eval_batch_size`: 16 - `num_train_epochs`: 10 - `warmup_ratio`: 0.1 - `log_level`: info - `tf32`: True - `load_best_model_at_end`: True - `push_to_hub`: True - `hub_strategy`: end - `hub_private_repo`: True - `hub_always_push`: True - `eval_on_start`: True #### All Hyperparameters
Click to expand - `overwrite_output_dir`: False - `do_predict`: False - `eval_strategy`: epoch - `prediction_loss_only`: True - `per_device_train_batch_size`: 64 - `per_device_eval_batch_size`: 16 - `per_gpu_train_batch_size`: None - `per_gpu_eval_batch_size`: None - `gradient_accumulation_steps`: 1 - `eval_accumulation_steps`: None - `torch_empty_cache_steps`: None - `learning_rate`: 5e-05 - `weight_decay`: 0 - `adam_beta1`: 0.9 - `adam_beta2`: 0.999 - `adam_epsilon`: 1e-08 - `max_grad_norm`: 1.0 - `num_train_epochs`: 10 - `max_steps`: -1 - `lr_scheduler_type`: linear - `lr_scheduler_kwargs`: None - `warmup_ratio`: 0.1 - `warmup_steps`: 0 - `log_level`: info - `log_level_replica`: warning - `log_on_each_node`: True - `logging_nan_inf_filter`: True - `save_safetensors`: True - `save_on_each_node`: False - `save_only_model`: False - `restore_callback_states_from_checkpoint`: False - `no_cuda`: False - `use_cpu`: False - `use_mps_device`: False - `seed`: 42 - `data_seed`: None - `jit_mode_eval`: False - `bf16`: False - `fp16`: False - `fp16_opt_level`: O1 - `half_precision_backend`: auto - `bf16_full_eval`: False - `fp16_full_eval`: False - `tf32`: True - `local_rank`: 0 - `ddp_backend`: None - `tpu_num_cores`: None - `tpu_metrics_debug`: False - `debug`: [] - `dataloader_drop_last`: False - `dataloader_num_workers`: 0 - `dataloader_prefetch_factor`: None - `past_index`: -1 - `disable_tqdm`: False - `remove_unused_columns`: True - `label_names`: None - `load_best_model_at_end`: True - `ignore_data_skip`: False - `fsdp`: [] - `fsdp_min_num_params`: 0 - `fsdp_config`: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False} - `fsdp_transformer_layer_cls_to_wrap`: None - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None} - `parallelism_config`: None - `deepspeed`: None - `label_smoothing_factor`: 0.0 - `optim`: adamw_torch - `optim_args`: None - `adafactor`: False - `group_by_length`: False - `length_column_name`: length - `project`: huggingface - `trackio_space_id`: trackio - `ddp_find_unused_parameters`: None - `ddp_bucket_cap_mb`: None - `ddp_broadcast_buffers`: False - `dataloader_pin_memory`: True - `dataloader_persistent_workers`: False - `skip_memory_metrics`: True - `use_legacy_prediction_loop`: False - `push_to_hub`: True - `resume_from_checkpoint`: None - `hub_model_id`: None - `hub_strategy`: end - `hub_private_repo`: True - `hub_always_push`: True - `hub_revision`: None - `gradient_checkpointing`: False - `gradient_checkpointing_kwargs`: None - `include_inputs_for_metrics`: False - `include_for_metrics`: [] - `eval_do_concat_batches`: True - `fp16_backend`: auto - `push_to_hub_model_id`: None - `push_to_hub_organization`: None - `mp_parameters`: - `auto_find_batch_size`: False - `full_determinism`: False - `torchdynamo`: None - `ray_scope`: last - `ddp_timeout`: 1800 - `torch_compile`: False - `torch_compile_backend`: None - `torch_compile_mode`: None - `include_tokens_per_second`: False - `include_num_input_tokens_seen`: no - `neftune_noise_alpha`: None - `optim_target_modules`: None - `batch_eval_metrics`: False - `eval_on_start`: True - `use_liger_kernel`: False - `liger_kernel_config`: None - `eval_use_gather_object`: False - `average_tokens_across_devices`: True - `prompts`: None - `batch_sampler`: batch_sampler - `multi_dataset_batch_sampler`: proportional - `router_mapping`: {} - `learning_rate_mapping`: {}
### Training Logs | Epoch | Step | Training Loss | Validation Loss | |:--------:|:---------:|:-------------:|:---------------:| | 0 | 0 | - | 1.0000 | | 1.0 | 2576 | 0.9894 | 0.9858 | | 2.0 | 5152 | 0.9678 | 0.9775 | | 3.0 | 7728 | 0.9525 | 0.9700 | | 4.0 | 10304 | 0.9406 | 0.9628 | | 5.0 | 12880 | 0.9309 | 0.9607 | | 6.0 | 15456 | 0.9244 | 0.9584 | | 7.0 | 18032 | 0.9188 | 0.9573 | | 8.0 | 20608 | 0.9138 | 0.9571 | | 9.0 | 23184 | 0.9114 | 0.9544 | | **10.0** | **25760** | **0.9084** | **0.9543** | * The bold row denotes the saved checkpoint. ### Framework Versions - Python: 3.12.11 - Sentence Transformers: 5.2.0 - Transformers: 4.57.6 - PyTorch: 2.7.1+cu128 - Accelerate: 1.9.0 - Datasets: 4.1.1 - Tokenizers: 0.22.1 ## Citation ### BibTeX #### Sentence Transformers ```bibtex @inproceedings{reimers-2019-sentence-bert, title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks", author = "Reimers, Nils and Gurevych, Iryna", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing", month = "11", year = "2019", publisher = "Association for Computational Linguistics", url = "https://arxiv.org/abs/1908.10084", } ``` #### RankNetLoss ```bibtex @inproceedings{burges2005learning, title={Learning to Rank using Gradient Descent}, author={Burges, Chris and Shaked, Tal and Renshaw, Erin and Lazier, Ari and Deeds, Matt and Hamilton, Nicole and Hullender, Greg}, booktitle={Proceedings of the 22nd international conference on Machine learning}, pages={89--96}, year={2005} } ```