GiacomoSignorile commited on
Commit
8e441ee
·
verified ·
1 Parent(s): b74754c

Initial upload of fine-tuned PatentSBERTa Specialist (eHealth)

Browse files
Files changed (4) hide show
  1. README.md +50 -47
  2. model.safetensors +1 -1
  3. sentence_bert_config.json +1 -1
  4. tokenizer.json +1 -1
README.md CHANGED
@@ -5,38 +5,36 @@ tags:
5
  - feature-extraction
6
  - dense
7
  - generated_from_trainer
8
- - dataset_size:3907
9
  - loss:MultipleNegativesRankingLoss
 
10
  base_model: AI-Growth-Lab/PatentSBERTa
11
  widget:
12
- - source_sentence: 'TMEK2: DEVICE FOR THE QUANTIFICATION OF BIOLOGICAL COMPONENTS
13
- IN A FLUID'
14
  sentences:
15
- - 'Technical Classification: C09D_11'
16
- - 'Technical Classification: G01M'
17
- - 'Technical Classification: G01N_33'
18
- - source_sentence: Portable optical analyzer for agri-food sector
19
  sentences:
20
- - 'Technical Classification: A61L_27'
21
- - 'Technical Classification: H01L_21'
22
- - 'Technical Classification: G01N_21'
23
- - source_sentence: Determining the antioxidant power of biological fluids via palladium
24
- nanoparticles
25
  sentences:
26
- - 'Technical Classification: E01C_23'
27
- - 'Technical Classification: A61K_31'
28
- - 'Technical Classification: G01N_33'
29
- - source_sentence: Helical scanning system for curved tubes
30
  sentences:
31
- - 'Technical Classification: G01N_29'
32
- - 'Technical Classification: B01D_53'
33
- - 'Technical Classification: A61P_25'
34
- - source_sentence: MICROSCOPY METHOD AND APPARATUS FOR OPTICAL TRACKING OF EMITTER
35
- OBJECTS
36
  sentences:
37
- - 'Technical Classification: H04N_13'
38
- - 'Technical Classification: E04G_21'
39
- - 'Technical Classification: B64D_27'
40
  pipeline_tag: sentence-similarity
41
  library_name: sentence-transformers
42
  ---
@@ -50,7 +48,7 @@ This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [A
50
  ### Model Description
51
  - **Model Type:** Sentence Transformer
52
  - **Base model:** [AI-Growth-Lab/PatentSBERTa](https://huggingface.co/AI-Growth-Lab/PatentSBERTa) <!-- at revision 3ff1d553c861d8f5bfd902333d97fc95eb6b4c8f -->
53
- - **Maximum Sequence Length:** 512 tokens
54
  - **Output Dimensionality:** 768 dimensions
55
  - **Similarity Function:** Cosine Similarity
56
  <!-- - **Training Dataset:** Unknown -->
@@ -67,7 +65,7 @@ This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [A
67
 
68
  ```
69
  SentenceTransformer(
70
- (0): Transformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'MPNetModel'})
71
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
72
  )
73
  ```
@@ -90,9 +88,9 @@ from sentence_transformers import SentenceTransformer
90
  model = SentenceTransformer("sentence_transformers_model_id")
91
  # Run inference
92
  sentences = [
93
- 'MICROSCOPY METHOD AND APPARATUS FOR OPTICAL TRACKING OF EMITTER OBJECTS',
94
- 'Technical Classification: H04N_13',
95
- 'Technical Classification: B64D_27',
96
  ]
97
  embeddings = model.encode(sentences)
98
  print(embeddings.shape)
@@ -101,9 +99,9 @@ print(embeddings.shape)
101
  # Get the similarity scores for the embeddings
102
  similarities = model.similarity(embeddings, embeddings)
103
  print(similarities)
104
- # tensor([[1.0000, 0.4365, 0.3320],
105
- # [0.4365, 1.0000, 0.4859],
106
- # [0.3320, 0.4859, 1.0000]])
107
  ```
108
 
109
  <!--
@@ -148,19 +146,19 @@ You can finetune this model on your own dataset.
148
 
149
  #### Unnamed Dataset
150
 
151
- * Size: 3,907 training samples
152
  * Columns: <code>sentence_0</code> and <code>sentence_1</code>
153
  * Approximate statistics based on the first 1000 samples:
154
- | | sentence_0 | sentence_1 |
155
- |:--------|:----------------------------------------------------------------------------------|:---------------------------------------------------------------------------------|
156
- | type | string | string |
157
- | details | <ul><li>min: 4 tokens</li><li>mean: 12.62 tokens</li><li>max: 38 tokens</li></ul> | <ul><li>min: 7 tokens</li><li>mean: 9.89 tokens</li><li>max: 15 tokens</li></ul> |
158
  * Samples:
159
- | sentence_0 | sentence_1 |
160
- |:--------------------------------------------------------------------|:----------------------------------------------------|
161
- | <code>INNOVATIVE PLATFORM FOR THE STUDY OF AGING PATHOLOGIES</code> | <code>Technical Classification: G16B_25</code> |
162
- | <code>AGENT FOR THE TREATMENT OF LEIOMYOMAS</code> | <code>Technical Classification: A61K_31_7105</code> |
163
- | <code>POLY DIVINYLBENZENE FOR POLYPEPTIDE SYNTHESIS</code> | <code>Technical Classification: C08F_12</code> |
164
  * Loss: [<code>MultipleNegativesRankingLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#multiplenegativesrankingloss) with these parameters:
165
  ```json
166
  {
@@ -173,14 +171,14 @@ You can finetune this model on your own dataset.
173
  ### Training Hyperparameters
174
  #### Non-Default Hyperparameters
175
 
176
- - `per_device_train_batch_size`: 16
177
- - `per_device_eval_batch_size`: 16
178
  - `multi_dataset_batch_sampler`: round_robin
179
 
180
  #### All Hyperparameters
181
  <details><summary>Click to expand</summary>
182
 
183
- - `per_device_train_batch_size`: 16
184
  - `num_train_epochs`: 3
185
  - `max_steps`: -1
186
  - `learning_rate`: 5e-05
@@ -223,7 +221,7 @@ You can finetune this model on your own dataset.
223
  - `project`: huggingface
224
  - `trackio_space_id`: trackio
225
  - `eval_strategy`: no
226
- - `per_device_eval_batch_size`: 16
227
  - `prediction_loss_only`: True
228
  - `eval_on_start`: False
229
  - `eval_do_concat_batches`: True
@@ -283,7 +281,12 @@ You can finetune this model on your own dataset.
283
  ### Training Logs
284
  | Epoch | Step | Training Loss |
285
  |:------:|:----:|:-------------:|
286
- | 2.0408 | 500 | 2.0066 |
 
 
 
 
 
287
 
288
 
289
  ### Framework Versions
 
5
  - feature-extraction
6
  - dense
7
  - generated_from_trainer
8
+ - dataset_size:2633
9
  - loss:MultipleNegativesRankingLoss
10
+ - dataset_size:7861
11
  base_model: AI-Growth-Lab/PatentSBERTa
12
  widget:
13
+ - source_sentence: CONSORTIUM OF PROBIOTICS.
 
14
  sentences:
15
+ - NANOSTRUCTURED FILTERING MEMBRANE FOR ANALYTICAL APPLICATIONS.
16
+ - CONSORTIUM OF PROBIOTICS.
17
+ - Real-time communications over Bluetooth Low Energy.
18
+ - source_sentence: Device for the collection and analysis of a biological fluid.
19
  sentences:
20
+ - Device for the collection and analysis of a biological fluid.
21
+ - STED MICROSCOPY BASED ON SYNCHRONOUS DETECTION OF FLUORESCENCE EMISSION.
22
+ - Robotic apparatus for minimally invasive surgery.
23
+ - source_sentence: Method for obtaining a porous semiconductor.
 
24
  sentences:
25
+ - 'Racetrack Logic: smart in-memory computing.'
26
+ - Method for obtaining a porous semiconductor.
27
+ - Salts of benzimidazole compounds, their use and synthetic preparation.
28
+ - source_sentence: Enthalpy exchangers with polymeric membrane.
29
  sentences:
30
+ - Closure system for engine connecting rods.
31
+ - Enthalpy exchangers with polymeric membrane.
32
+ - Soluble protein with high angiogenic activity.
33
+ - source_sentence: FLEXIBLE CAPACITIVE KEYBOARD FOR TEXTILE USE.
 
34
  sentences:
35
+ - 'GreenValve III: System for spherical segment valves.'
36
+ - FLEXIBLE CAPACITIVE KEYBOARD FOR TEXTILE USE.
37
+ - Mandibular advancement device.
38
  pipeline_tag: sentence-similarity
39
  library_name: sentence-transformers
40
  ---
 
48
  ### Model Description
49
  - **Model Type:** Sentence Transformer
50
  - **Base model:** [AI-Growth-Lab/PatentSBERTa](https://huggingface.co/AI-Growth-Lab/PatentSBERTa) <!-- at revision 3ff1d553c861d8f5bfd902333d97fc95eb6b4c8f -->
51
+ - **Maximum Sequence Length:** 384 tokens
52
  - **Output Dimensionality:** 768 dimensions
53
  - **Similarity Function:** Cosine Similarity
54
  <!-- - **Training Dataset:** Unknown -->
 
65
 
66
  ```
67
  SentenceTransformer(
68
+ (0): Transformer({'max_seq_length': 384, 'do_lower_case': False, 'architecture': 'MPNetModel'})
69
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
70
  )
71
  ```
 
88
  model = SentenceTransformer("sentence_transformers_model_id")
89
  # Run inference
90
  sentences = [
91
+ 'FLEXIBLE CAPACITIVE KEYBOARD FOR TEXTILE USE.',
92
+ 'FLEXIBLE CAPACITIVE KEYBOARD FOR TEXTILE USE.',
93
+ 'GreenValve III: System for spherical segment valves.',
94
  ]
95
  embeddings = model.encode(sentences)
96
  print(embeddings.shape)
 
99
  # Get the similarity scores for the embeddings
100
  similarities = model.similarity(embeddings, embeddings)
101
  print(similarities)
102
+ # tensor([[1.0000, 1.0000, 0.3686],
103
+ # [1.0000, 1.0000, 0.3686],
104
+ # [0.3686, 0.3686, 1.0000]])
105
  ```
106
 
107
  <!--
 
146
 
147
  #### Unnamed Dataset
148
 
149
+ * Size: 7,861 training samples
150
  * Columns: <code>sentence_0</code> and <code>sentence_1</code>
151
  * Approximate statistics based on the first 1000 samples:
152
+ | | sentence_0 | sentence_1 |
153
+ |:--------|:----------------------------------------------------------------------------------|:----------------------------------------------------------------------------------|
154
+ | type | string | string |
155
+ | details | <ul><li>min: 4 tokens</li><li>mean: 12.32 tokens</li><li>max: 32 tokens</li></ul> | <ul><li>min: 8 tokens</li><li>mean: 12.01 tokens</li><li>max: 36 tokens</li></ul> |
156
  * Samples:
157
+ | sentence_0 | sentence_1 |
158
+ |:--------------------------------------------------------------------------------------------------------------|:----------------------------------------------------------------|
159
+ | <code>Genetic diagnostic technique for SCA1-3,6,7</code> | <code>Applicant/Organization: UNIV DEGLI STUDI DI TORINO</code> |
160
+ | <code>Method for detecting Macrophomina phaseolina</code> | <code>Technical Classification: C12Q</code> |
161
+ | <code>System for controlled administration of a substance from a human -body-implanted infusion device</code> | <code>Applicant/Organization: Stefanini Cesare</code> |
162
  * Loss: [<code>MultipleNegativesRankingLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#multiplenegativesrankingloss) with these parameters:
163
  ```json
164
  {
 
171
  ### Training Hyperparameters
172
  #### Non-Default Hyperparameters
173
 
174
+ - `per_device_train_batch_size`: 12
175
+ - `per_device_eval_batch_size`: 12
176
  - `multi_dataset_batch_sampler`: round_robin
177
 
178
  #### All Hyperparameters
179
  <details><summary>Click to expand</summary>
180
 
181
+ - `per_device_train_batch_size`: 12
182
  - `num_train_epochs`: 3
183
  - `max_steps`: -1
184
  - `learning_rate`: 5e-05
 
221
  - `project`: huggingface
222
  - `trackio_space_id`: trackio
223
  - `eval_strategy`: no
224
+ - `per_device_eval_batch_size`: 12
225
  - `prediction_loss_only`: True
226
  - `eval_on_start`: False
227
  - `eval_do_concat_batches`: True
 
281
  ### Training Logs
282
  | Epoch | Step | Training Loss |
283
  |:------:|:----:|:-------------:|
284
+ | 1.5152 | 500 | 0.0002 |
285
+ | 3.0303 | 1000 | 0.0000 |
286
+ | 4.5455 | 1500 | 0.0000 |
287
+ | 0.7622 | 500 | 2.3552 |
288
+ | 1.5244 | 1000 | 1.9437 |
289
+ | 2.2866 | 1500 | 1.8059 |
290
 
291
 
292
  ### Framework Versions
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:3c205fd35ee2f54179664085aee58168bd124d632f7a65423e0ffbf59b3342cb
3
  size 437967648
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7124e7a341b5a27bb2703ada4533fd5cbeaab50189f7313fa499b4f2c132263f
3
  size 437967648
sentence_bert_config.json CHANGED
@@ -1,4 +1,4 @@
1
  {
2
- "max_seq_length": 512,
3
  "do_lower_case": false
4
  }
 
1
  {
2
+ "max_seq_length": 384,
3
  "do_lower_case": false
4
  }
tokenizer.json CHANGED
@@ -2,7 +2,7 @@
2
  "version": "1.0",
3
  "truncation": {
4
  "direction": "Right",
5
- "max_length": 512,
6
  "strategy": "LongestFirst",
7
  "stride": 0
8
  },
 
2
  "version": "1.0",
3
  "truncation": {
4
  "direction": "Right",
5
+ "max_length": 384,
6
  "strategy": "LongestFirst",
7
  "stride": 0
8
  },