--- language: - en license: apache-2.0 tags: - sentence-transformers - sentence-similarity - feature-extraction - dense - generated_from_trainer - dataset_size:5175 - loss:CachedMultipleNegativesRankingLoss base_model: nomic-ai/modernbert-embed-base widget: - source_sentence: What court case is referenced with citation 546 F.2d at 1015 n.14? sentences: - "3023980, at *4 (D.D.C. May 25, 2016). The Court will not exercise its discretion\ \ to do so here, \ngiven the potentially important national security interests\ \ at stake. Still, the Court advises the \n \n3 The Court also questions whether\ \ the Department has sufficiently shown that it conducted \nan adequate segregability\ \ analysis. FOIA requires that “[a]ny reasonably segregable portion of a" - "records, using methods which can be reasonably expected to produce the information\ \ requested.” \nOglesby v. U.S. Dep’t of Army, 920 F.2d 57, 68 (D.C. Cir. 1990).\ \ “Summary judgment may be \nbased on affidavit, if the declaration sets forth\ \ sufficiently detailed information ‘for a court to \ndetermine if the search\ \ was adequate.’” Students Against Genocide v. Dep’t of State, 257 F.3d" - "546 F.2d at 1015 n.14, in contravention of the D.C. Circuit’s limitation on that\ \ provision’s scope. \nCf. Milner, 131 S. Ct. at 1269 (observing that “[u]nder\ \ this interpretation, an agency’s ‘internal \npersonnel rules and practices’\ \ appears to mean all its internal rules and practices,” and thus “[t]he \nmodifier\ \ ‘personnel’ . . . does no modifying work”)." - source_sentence: What does the Court need to infer regarding the ODNI's submission? sentences: - "Mr. Zimmerman identified Mr. Mooney as the person in the video wearing the white\ \ \nshirt and identified himself as the person in the pink shirt. The circuit\ \ court also admitted \ninto evidence, without objection, two other videos, identified\ \ as State’s Exhibits 1B and 2, \nand the videos were played for the jury. Mr.\ \ Zimmerman testified that State’s Exhibit 1B" - "withholding information. The ODNI thus apparently leaves it to the Court to\ \ infer which \nportions of which of the thirty-four heavily redacted (and sometimes\ \ illegibly blurry) pages were \nwithheld under Exemption 5. This sort of submission\ \ is utterly unhelpful to the Court in \ndetermining whether a FOIA exemption\ \ applies to particular portions of particular records," - "withholding information. The ODNI thus apparently leaves it to the Court to\ \ infer which \nportions of which of the thirty-four heavily redacted (and sometimes\ \ illegibly blurry) pages were \nwithheld under Exemption 5. This sort of submission\ \ is utterly unhelpful to the Court in \ndetermining whether a FOIA exemption\ \ applies to particular portions of particular records," - source_sentence: What type of company is Senetas Corporation Limited? sentences: - "books and records that are necessary, sufficient and essential for the plaintiff’s\ \ \nproper purpose, prior to the Court making further determinations. This is\ \ a final \nreport. \nI. \nBackground \nA. Factual Background1 \n \nPlaintiff\ \ Senetas Corporation Limited (“Senetas”) is a publicly listed \nAustralian network\ \ encryption company.2 Senetas’ non-executive chairperson is" - "nevertheless maintains that “the Court has no way of knowing which of the withheld\ \ information \nactually does fall within these criteria.” Id. \nThe Court disagrees.\ \ The plaintiff is reading the CIA’s declaration and Vaughn index in \nisolation,\ \ rather than reading both documents together. Although the CIA’s Vaughn index\ \ only" - ". . . I seen him walk, and I didn’t know which way he went and so I’m looking.\ \ \nAnd then, I’m thinking he’s going to come run up to my front side of my --\ \ \nmy passen -- or the drive’s seat. . . . And, um, I look out my, I cracked\ \ my \ndoor and I’m looking out and I didn’t see him. As soon as I sat back that’s\ \ \nwhen the gunshots happened." - source_sentence: What case is referenced with the citation 352 U.S. 29, 33 (1956)? sentences: - "Id. at 1–2. The CIA eventually agreed to process the fourth option, but after\ \ the plaintiff refused \n \n\ 26 Additionally, requiring agencies to respond to FOIA requests like this would\ \ open them up to a whole new \ncategory of legal challenges regarding the adequacy\ \ of their search efforts. For example, a requester could later" - "14 In Weeks Marine, the Federal Circuit ascribed the “non-trivial competitive\ \ injury” standard for \nprejudice within the context of standing to file a pre-award\ \ bid protest. See Weeks Marine, 575 \nF.3d 1359–63 (“[W]e conclude that in a\ \ pre-award protest such as the one before us, [a] prospective" - "undoubtedly “touch[es] the rights and duties of the United States,” see Bank\ \ of Am. Nat’l Trust & Sav. Ass’n v. \nParnell, 352 U.S. 29, 33 (1956), it therefore\ \ likely qualifies as one of the “few areas . . . involving ‘uniquely federal\ \ \ninterests’” that requires the development of federal common law principles,\ \ see Boyle v. United Techs. Corp., 487" - source_sentence: What sections of the document are referenced in the location Supplement 2, AR? sentences: - "statement from Congress in the FOIA regarding assignments, the common-law principles\ \ \nregarding the recognition of assignments presumably apply, and, as discussed\ \ above, under \n \n16 Since the\ \ CIA’s Assignment of Rights Policy is categorical, the Court need not decide\ \ in what circumstances an" - "“Based on this misunderstanding, the CIA attorney incorrectly cited some of the\ \ justifications for \nredacting the material to the DOJ attorney, who in turn\ \ shared that information with plaintiff.” \nId. ¶ 9. \nE. \nProcedural History\ \ \nThe plaintiff filed the Complaints in each of these three actions on February\ \ 28, 2011," - "the Polaris Solicitations as currently drafted do not comply with Section 3306(c)(3).\ \ In its request \nto apply Section 3306(c)(3) to the Polaris Solicitations,\ \ GSA stated that \n \n \n \nSupplement 2, AR at 2907–08. Because GSA adopted\ \ an overly broad understanding of Section \n3306(c)(3)’s scope, GSA stated the\ \ Solicitations will include a “full range of order types,”" datasets: - AdamLucek/legal-rag-positives-synthetic pipeline_tag: sentence-similarity library_name: sentence-transformers metrics: - cosine_accuracy@1 - cosine_accuracy@3 - cosine_accuracy@5 - cosine_accuracy@10 - cosine_precision@1 - cosine_precision@3 - cosine_precision@5 - cosine_precision@10 - cosine_recall@1 - cosine_recall@3 - cosine_recall@5 - cosine_recall@10 - cosine_ndcg@10 - cosine_mrr@10 - cosine_map@100 model-index: - name: ModernBERT Embed Base Legal Fine-tuned results: - task: type: information-retrieval name: Information Retrieval dataset: name: ir eval test type: ir_eval_test metrics: - type: cosine_accuracy@1 value: 0.34312210200927357 name: Cosine Accuracy@1 - type: cosine_accuracy@3 value: 0.5131375579598145 name: Cosine Accuracy@3 - type: cosine_accuracy@5 value: 0.5749613601236476 name: Cosine Accuracy@5 - type: cosine_accuracy@10 value: 0.678516228748068 name: Cosine Accuracy@10 - type: cosine_precision@1 value: 0.34312210200927357 name: Cosine Precision@1 - type: cosine_precision@3 value: 0.1710458526532715 name: Cosine Precision@3 - type: cosine_precision@5 value: 0.1149922720247295 name: Cosine Precision@5 - type: cosine_precision@10 value: 0.06785162287480678 name: Cosine Precision@10 - type: cosine_recall@1 value: 0.34312210200927357 name: Cosine Recall@1 - type: cosine_recall@3 value: 0.5131375579598145 name: Cosine Recall@3 - type: cosine_recall@5 value: 0.5749613601236476 name: Cosine Recall@5 - type: cosine_recall@10 value: 0.678516228748068 name: Cosine Recall@10 - type: cosine_ndcg@10 value: 0.5027845721427848 name: Cosine Ndcg@10 - type: cosine_mrr@10 value: 0.4476730207796664 name: Cosine Mrr@10 - type: cosine_map@100 value: 0.45652768049118125 name: Cosine Map@100 - task: type: information-retrieval name: Information Retrieval dataset: name: ir eval eval type: ir_eval_eval metrics: - type: cosine_accuracy@1 value: 0.5795981452859351 name: Cosine Accuracy@1 - type: cosine_accuracy@3 value: 0.7527047913446677 name: Cosine Accuracy@3 - type: cosine_accuracy@5 value: 0.8176197836166924 name: Cosine Accuracy@5 - type: cosine_accuracy@10 value: 0.8639876352395672 name: Cosine Accuracy@10 - type: cosine_precision@1 value: 0.5795981452859351 name: Cosine Precision@1 - type: cosine_precision@3 value: 0.2509015971148892 name: Cosine Precision@3 - type: cosine_precision@5 value: 0.16352395672333847 name: Cosine Precision@5 - type: cosine_precision@10 value: 0.08639876352395671 name: Cosine Precision@10 - type: cosine_recall@1 value: 0.5795981452859351 name: Cosine Recall@1 - type: cosine_recall@3 value: 0.7527047913446677 name: Cosine Recall@3 - type: cosine_recall@5 value: 0.8176197836166924 name: Cosine Recall@5 - type: cosine_recall@10 value: 0.8639876352395672 name: Cosine Recall@10 - type: cosine_ndcg@10 value: 0.7246312019713473 name: Cosine Ndcg@10 - type: cosine_mrr@10 value: 0.6795276121783075 name: Cosine Mrr@10 - type: cosine_map@100 value: 0.6842337248584138 name: Cosine Map@100 --- # ModernBERT Embed Base Legal Fine-tuned This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [nomic-ai/modernbert-embed-base](https://huggingface.co/nomic-ai/modernbert-embed-base) on the [legal-rag-positives-synthetic](https://huggingface.co/datasets/AdamLucek/legal-rag-positives-synthetic) dataset. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more. ## Model Details ### Model Description - **Model Type:** Sentence Transformer - **Base model:** [nomic-ai/modernbert-embed-base](https://huggingface.co/nomic-ai/modernbert-embed-base) - **Maximum Sequence Length:** 8192 tokens - **Output Dimensionality:** 768 dimensions - **Similarity Function:** Cosine Similarity - **Training Dataset:** - [legal-rag-positives-synthetic](https://huggingface.co/datasets/AdamLucek/legal-rag-positives-synthetic) - **Language:** en - **License:** apache-2.0 ### Model Sources - **Documentation:** [Sentence Transformers Documentation](https://sbert.net) - **Repository:** [Sentence Transformers on GitHub](https://github.com/huggingface/sentence-transformers) - **Hugging Face:** [Sentence Transformers on Hugging Face](https://huggingface.co/models?library=sentence-transformers) ### Full Model Architecture ``` SentenceTransformer( (0): Transformer({'max_seq_length': 8192, 'do_lower_case': False, 'architecture': 'ModernBertModel'}) (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True}) (2): Normalize() ) ``` ## Usage ### Direct Usage (Sentence Transformers) First install the Sentence Transformers library: ```bash pip install -U sentence-transformers ``` Then you can load this model and run inference. ```python from sentence_transformers import SentenceTransformer # Download from the 🤗 Hub model = SentenceTransformer("aaa961/modernbert-embed-base-legal-no_MRL_symmetricMNRL_3sets") # Run inference sentences = [ 'What sections of the document are referenced in the location Supplement 2, AR?', 'the Polaris Solicitations as currently drafted do not comply with Section 3306(c)(3). In its request \nto apply Section 3306(c)(3) to the Polaris Solicitations, GSA stated that \n \n \n \nSupplement 2, AR at 2907–08. Because GSA adopted an overly broad understanding of Section \n3306(c)(3)’s scope, GSA stated the Solicitations will include a “full range of order types,”', '“Based on this misunderstanding, the CIA attorney incorrectly cited some of the justifications for \nredacting the material to the DOJ attorney, who in turn shared that information with plaintiff.” \nId. ¶ 9. \nE. \nProcedural History \nThe plaintiff filed the Complaints in each of these three actions on February 28, 2011,', ] embeddings = model.encode(sentences) print(embeddings.shape) # [3, 768] # Get the similarity scores for the embeddings similarities = model.similarity(embeddings, embeddings) print(similarities) # tensor([[1.0000, 0.3908, 0.0520], # [0.3908, 1.0000, 0.0703], # [0.0520, 0.0703, 1.0000]]) ``` ## Evaluation ### Metrics #### Information Retrieval * Datasets: `ir_eval_test` and `ir_eval_eval` * Evaluated with [InformationRetrievalEvaluator](https://sbert.net/docs/package_reference/sentence_transformer/evaluation.html#sentence_transformers.evaluation.InformationRetrievalEvaluator) | Metric | ir_eval_test | ir_eval_eval | |:--------------------|:-------------|:-------------| | cosine_accuracy@1 | 0.3431 | 0.5796 | | cosine_accuracy@3 | 0.5131 | 0.7527 | | cosine_accuracy@5 | 0.575 | 0.8176 | | cosine_accuracy@10 | 0.6785 | 0.864 | | cosine_precision@1 | 0.3431 | 0.5796 | | cosine_precision@3 | 0.171 | 0.2509 | | cosine_precision@5 | 0.115 | 0.1635 | | cosine_precision@10 | 0.0679 | 0.0864 | | cosine_recall@1 | 0.3431 | 0.5796 | | cosine_recall@3 | 0.5131 | 0.7527 | | cosine_recall@5 | 0.575 | 0.8176 | | cosine_recall@10 | 0.6785 | 0.864 | | **cosine_ndcg@10** | **0.5028** | **0.7246** | | cosine_mrr@10 | 0.4477 | 0.6795 | | cosine_map@100 | 0.4565 | 0.6842 | ## Training Details ### Training Dataset #### legal-rag-positives-synthetic * Dataset: [legal-rag-positives-synthetic](https://huggingface.co/datasets/AdamLucek/legal-rag-positives-synthetic) at [f11534a](https://huggingface.co/datasets/AdamLucek/legal-rag-positives-synthetic/tree/f11534aeed060a3245f55f8f9d944cf8132c780d) * Size: 5,175 training samples * Columns: anchor and positive * Approximate statistics based on the first 1000 samples: | | anchor | positive | |:--------|:----------------------------------------------------------------------------------|:------------------------------------------------------------------------------------| | type | string | string | | details | | | * Samples: | anchor | positive | |:--------------------------------------------------------------------------------------------------------|:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | Where is the similar statement to the one about business judgment and scoring merit found? | is with each bidder itself, and its own business judgment in forming a team and what score it thinks
is enough to merit an award.”) (emphasis in original); VCH MJAR at 21–22 (same); Oral Ar. Tr.
at 10:5–7 (“[C]ompetition involves . . . some sort of tradeoff between offerors, some sort of
evaluation of how offerors are against one another, and that’s not the case here. The case here is
| | Who do lawyers generally employ as assistants in their practice? | abide by the Rules of Professional Conduct. See rule 4-5.2(a).

RULE 4-5.3.
RESPONSIBILITIES REGARDING NONLAWYER
ASSISTANTS
(a) – (c) [No Change]
Comment
Lawyers generally employ assistants in their practice,
including secretaries, investigators, law student interns, and
paraprofessionals such as paralegals and legal assistants. Such
| | Which court case is cited with a page number of 1327? | 30; VCH MJAR at 28–30 (same).
As noted, this Court applies the same interpretive rules to analyze both statutes and federal
regulations. See Boeing, 983 F.3d at 1327 (citing Mass. Mut. Life Ins. Co., 782 F.3d at 1365); see
also supra Discussion Section I. It is a “fundamental canon of statutory construction that the words
| * Loss: [CachedMultipleNegativesRankingLoss](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#cachedmultiplenegativesrankingloss) with these parameters: ```json { "scale": 20.0, "similarity_fct": "cos_sim", "mini_batch_size": 32, "gather_across_devices": false, "directions": [ "query_to_doc", "doc_to_query" ], "partition_mode": "per_direction", "hardness_mode": null, "hardness_strength": 0.0 } ``` ### Training Hyperparameters #### Non-Default Hyperparameters - `per_device_train_batch_size`: 32 - `num_train_epochs`: 4 - `learning_rate`: 2e-05 - `lr_scheduler_type`: cosine - `warmup_steps`: 0.1 - `optim`: adamw_torch_fused - `gradient_accumulation_steps`: 16 - `bf16`: True - `tf32`: True - `eval_strategy`: epoch - `per_device_eval_batch_size`: 16 - `load_best_model_at_end`: True #### All Hyperparameters
Click to expand - `per_device_train_batch_size`: 32 - `num_train_epochs`: 4 - `max_steps`: -1 - `learning_rate`: 2e-05 - `lr_scheduler_type`: cosine - `lr_scheduler_kwargs`: None - `warmup_steps`: 0.1 - `optim`: adamw_torch_fused - `optim_args`: None - `weight_decay`: 0.0 - `adam_beta1`: 0.9 - `adam_beta2`: 0.999 - `adam_epsilon`: 1e-08 - `optim_target_modules`: None - `gradient_accumulation_steps`: 16 - `average_tokens_across_devices`: True - `max_grad_norm`: 1.0 - `label_smoothing_factor`: 0.0 - `bf16`: True - `fp16`: False - `bf16_full_eval`: False - `fp16_full_eval`: False - `tf32`: True - `gradient_checkpointing`: False - `gradient_checkpointing_kwargs`: None - `torch_compile`: False - `torch_compile_backend`: None - `torch_compile_mode`: None - `use_liger_kernel`: False - `liger_kernel_config`: None - `use_cache`: False - `neftune_noise_alpha`: None - `torch_empty_cache_steps`: None - `auto_find_batch_size`: False - `log_on_each_node`: True - `logging_nan_inf_filter`: True - `include_num_input_tokens_seen`: no - `log_level`: passive - `log_level_replica`: warning - `disable_tqdm`: False - `project`: huggingface - `trackio_space_id`: trackio - `eval_strategy`: epoch - `per_device_eval_batch_size`: 16 - `prediction_loss_only`: True - `eval_on_start`: False - `eval_do_concat_batches`: True - `eval_use_gather_object`: False - `eval_accumulation_steps`: None - `include_for_metrics`: [] - `batch_eval_metrics`: False - `save_only_model`: False - `save_on_each_node`: False - `enable_jit_checkpoint`: False - `push_to_hub`: False - `hub_private_repo`: None - `hub_model_id`: None - `hub_strategy`: every_save - `hub_always_push`: False - `hub_revision`: None - `load_best_model_at_end`: True - `ignore_data_skip`: False - `restore_callback_states_from_checkpoint`: False - `full_determinism`: False - `seed`: 42 - `data_seed`: None - `use_cpu`: False - `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None} - `parallelism_config`: None - `dataloader_drop_last`: False - `dataloader_num_workers`: 0 - `dataloader_pin_memory`: True - `dataloader_persistent_workers`: False - `dataloader_prefetch_factor`: None - `remove_unused_columns`: True - `label_names`: None - `train_sampling_strategy`: random - `length_column_name`: length - `ddp_find_unused_parameters`: None - `ddp_bucket_cap_mb`: None - `ddp_broadcast_buffers`: False - `ddp_backend`: None - `ddp_timeout`: 1800 - `fsdp`: [] - `fsdp_config`: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False} - `deepspeed`: None - `debug`: [] - `skip_memory_metrics`: True - `do_predict`: False - `resume_from_checkpoint`: None - `warmup_ratio`: None - `local_rank`: -1 - `prompts`: None - `batch_sampler`: batch_sampler - `multi_dataset_batch_sampler`: proportional - `router_mapping`: {} - `learning_rate_mapping`: {}
### Training Logs | Epoch | Step | Training Loss | ir_eval_test_cosine_ndcg@10 | ir_eval_eval_cosine_ndcg@10 | |:-------:|:------:|:-------------:|:---------------------------:|:---------------------------:| | -1 | -1 | - | 0.5028 | - | | 0.9877 | 10 | 0.9180 | - | - | | 1.0 | 11 | - | - | 0.6723 | | 1.8889 | 20 | 0.4042 | - | - | | 2.0 | 22 | - | - | 0.7082 | | 2.7901 | 30 | 0.2940 | - | - | | 3.0 | 33 | - | - | 0.7236 | | 3.6914 | 40 | 0.2646 | - | - | | **4.0** | **44** | **-** | **-** | **0.7246** | * The bold row denotes the saved checkpoint. ### Framework Versions - Python: 3.12.11 - Sentence Transformers: 5.3.0 - Transformers: 5.3.0 - PyTorch: 2.5.1+cu121 - Accelerate: 1.13.0 - Datasets: 4.8.2 - Tokenizers: 0.22.2 ## Citation ### BibTeX #### Sentence Transformers ```bibtex @inproceedings{reimers-2019-sentence-bert, title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks", author = "Reimers, Nils and Gurevych, Iryna", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing", month = "11", year = "2019", publisher = "Association for Computational Linguistics", url = "https://arxiv.org/abs/1908.10084", } ``` #### CachedMultipleNegativesRankingLoss ```bibtex @misc{gao2021scaling, title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup}, author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan}, year={2021}, eprint={2101.06983}, archivePrefix={arXiv}, primaryClass={cs.LG} } ```