mawe2 commited on
Commit
697719c
·
verified ·
1 Parent(s): 656aa5a

Add full model card

Browse files
Files changed (1) hide show
  1. README.md +217 -142
README.md CHANGED
@@ -1,199 +1,274 @@
1
  ---
2
- library_name: transformers
3
- tags: []
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  ---
5
 
6
- # Model Card for Model ID
7
 
8
- <!-- Provide a quick summary of what the model is/does. -->
 
 
9
 
 
 
10
 
 
11
 
12
- ## Model Details
13
-
14
- ### Model Description
15
-
16
- <!-- Provide a longer summary of what this model is. -->
17
-
18
- This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
19
-
20
- - **Developed by:** [More Information Needed]
21
- - **Funded by [optional]:** [More Information Needed]
22
- - **Shared by [optional]:** [More Information Needed]
23
- - **Model type:** [More Information Needed]
24
- - **Language(s) (NLP):** [More Information Needed]
25
- - **License:** [More Information Needed]
26
- - **Finetuned from model [optional]:** [More Information Needed]
27
-
28
- ### Model Sources [optional]
29
-
30
- <!-- Provide the basic links for the model. -->
31
-
32
- - **Repository:** [More Information Needed]
33
- - **Paper [optional]:** [More Information Needed]
34
- - **Demo [optional]:** [More Information Needed]
35
-
36
- ## Uses
37
-
38
- <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
39
-
40
- ### Direct Use
41
-
42
- <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
43
-
44
- [More Information Needed]
45
-
46
- ### Downstream Use [optional]
47
-
48
- <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
49
-
50
- [More Information Needed]
51
-
52
- ### Out-of-Scope Use
53
-
54
- <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
55
-
56
- [More Information Needed]
57
-
58
- ## Bias, Risks, and Limitations
59
-
60
- <!-- This section is meant to convey both technical and sociotechnical limitations. -->
61
-
62
- [More Information Needed]
63
-
64
- ### Recommendations
65
-
66
- <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
67
-
68
- Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
69
-
70
- ## How to Get Started with the Model
71
-
72
- Use the code below to get started with the model.
73
-
74
- [More Information Needed]
75
-
76
- ## Training Details
77
-
78
- ### Training Data
79
-
80
- <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
81
-
82
- [More Information Needed]
83
-
84
- ### Training Procedure
85
-
86
- <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
87
-
88
- #### Preprocessing [optional]
89
-
90
- [More Information Needed]
91
-
92
-
93
- #### Training Hyperparameters
94
-
95
- - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
96
-
97
- #### Speeds, Sizes, Times [optional]
98
 
99
- <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
 
 
 
 
100
 
101
- [More Information Needed]
102
 
103
- ## Evaluation
104
 
105
- <!-- This section describes the evaluation protocols and provides the results. -->
 
 
 
 
 
 
106
 
107
- ### Testing Data, Factors & Metrics
108
 
109
- #### Testing Data
110
 
111
- <!-- This should link to a Dataset Card if possible. -->
 
 
 
 
 
 
 
 
 
 
112
 
113
- [More Information Needed]
114
 
115
- #### Factors
116
 
117
- <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
 
 
118
 
119
- [More Information Needed]
 
 
 
 
 
 
 
120
 
121
- #### Metrics
122
 
123
- <!-- These are the evaluation metrics being used, ideally with a description of why. -->
124
 
125
- [More Information Needed]
 
 
 
 
126
 
127
- ### Results
128
 
129
- [More Information Needed]
 
 
 
130
 
131
- #### Summary
132
 
 
133
 
 
 
 
 
134
 
135
- ## Model Examination [optional]
 
136
 
137
- <!-- Relevant interpretability work for the model goes here -->
 
138
 
139
- [More Information Needed]
 
 
140
 
141
- ## Environmental Impact
 
 
 
142
 
143
- <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
 
 
 
144
 
145
- Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
 
146
 
147
- - **Hardware Type:** [More Information Needed]
148
- - **Hours used:** [More Information Needed]
149
- - **Cloud Provider:** [More Information Needed]
150
- - **Compute Region:** [More Information Needed]
151
- - **Carbon Emitted:** [More Information Needed]
152
 
153
- ## Technical Specifications [optional]
 
 
154
 
155
- ### Model Architecture and Objective
156
 
157
- [More Information Needed]
158
 
159
- ### Compute Infrastructure
 
 
 
 
 
 
160
 
161
- [More Information Needed]
 
 
 
 
162
 
163
- #### Hardware
 
 
164
 
165
- [More Information Needed]
166
 
167
- #### Software
168
 
169
- [More Information Needed]
 
 
 
170
 
171
- ## Citation [optional]
 
 
 
 
172
 
173
- <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
 
 
 
174
 
175
- **BibTeX:**
 
176
 
177
- [More Information Needed]
178
 
179
- **APA:**
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
180
 
181
- [More Information Needed]
182
 
183
- ## Glossary [optional]
184
 
185
- <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
186
 
187
- [More Information Needed]
 
 
 
 
 
 
 
188
 
189
- ## More Information [optional]
190
 
191
- [More Information Needed]
192
 
193
- ## Model Card Authors [optional]
 
 
 
 
 
 
 
194
 
195
- [More Information Needed]
196
 
197
- ## Model Card Contact
198
 
199
- [More Information Needed]
 
 
1
  ---
2
+ language:
3
+ - en
4
+ license: apache-2.0
5
+ tags:
6
+ - biology
7
+ - protein
8
+ - longevity
9
+ - aging
10
+ - ESM-2
11
+ - LoRA
12
+ - sequence-classification
13
+ datasets:
14
+ - GenAge
15
+ - SwissProt
16
+ metrics:
17
+ - auprc
18
+ - roc_auc
19
+ base_model: facebook/esm2_t30_150M_UR50D
20
  ---
21
 
22
+ # Longevity Protein Classifier v6
23
 
24
+ Fine-tuned ESM-2 150M for binary classification of protein sequences
25
+ as longevity-associated or not, trained on multi-species GenAge data
26
+ with LoRA adapters.
27
 
28
+ Built as part of a personal ML learning arc — Week 3 of 8 —
29
+ connecting protein language models to longevity biology.
30
 
31
+ ---
32
 
33
+ ## Model Description
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
34
 
35
+ - **Model type:** ESM-2 150M + LoRA (r=16) sequence classifier
36
+ - **Base model:** facebook/esm2_t30_150M_UR50D
37
+ - **Task:** Binary classification — longevity-associated vs non-longevity
38
+ - **Developed by:** Mo Elzek
39
+ - **License:** Apache 2.0
40
 
41
+ ---
42
 
43
+ ## Performance
44
 
45
+ | Metric | Value |
46
+ |--------|-------|
47
+ | Test AUPRC | 0.335 |
48
+ | Test AUC-ROC | 0.696 |
49
+ | Random AUPRC baseline | 0.061 |
50
+ | Improvement over random | 5.5x |
51
+ | Training epochs | 10 (early stopping) |
52
 
53
+ ---
54
 
55
+ ## Benchmark Results
56
 
57
+ | Protein | Score | Expected | Notes |
58
+ |---------|-------|----------|-------|
59
+ | SIRT1 | 0.996 | HIGH | NAD+ deacetylase, caloric restriction mediator |
60
+ | SIRT3 | 0.998 | HIGH | Mitochondrial sirtuin |
61
+ | TP53 | 0.974 | HIGH | Tumour suppressor, aging roles |
62
+ | MYH9 | 0.000 | LOW | Structural myosin — negative control |
63
+ | ACTB | 0.000 | LOW | Beta actin — negative control |
64
+ | ALB | 0.000 | LOW | Serum albumin — negative control |
65
+ | FOXO3 | 0.000 | HIGH | **Fails** — see limitations |
66
+ | MTOR | 0.000 | HIGH | **Fails** — see limitations |
67
+ | TERT | 0.000 | HIGH | **Fails** — see limitations |
68
 
69
+ ---
70
 
71
+ ## Novel Predictions Not in GenAge
72
 
73
+ Proteins scoring above 0.50 that are not present in GenAge human
74
+ database. These are the model's predictions of longevity-relevant
75
+ proteins not yet catalogued — not validated findings.
76
 
77
+ | Protein | Score | Biological relevance |
78
+ |---------|-------|----------------------|
79
+ | TFEB | 0.502 | Master regulator of autophagy and lysosomal biogenesis. Overexpression extends lifespan in C. elegans. Regulated by mTOR. Strongest novel prediction. |
80
+ | NEIL1 | 0.951 | DNA glycosylase, base excision repair of oxidative damage. DNA repair capacity correlates with species lifespan. |
81
+ | GSTA1 | 0.871 | Glutathione S-transferase. Antioxidant defence. GST family implicated in longevity across multiple species. |
82
+ | GRHL1 | 0.880 | Grainyhead-like transcription factor. Epithelial barrier maintenance — tissue integrity declines with age. |
83
+ | EXO1 | 0.550 | Exonuclease involved in DNA mismatch repair and double-strand break repair. |
84
+ | MSH4 | 0.546 | DNA mismatch repair. Related family members (MSH2, MSH6) are established longevity-associated genes. |
85
 
86
+ ---
87
 
88
+ ## Recommended Thresholds
89
 
90
+ | Use case | Threshold | Precision | Recall |
91
+ |----------|-----------|-----------|--------|
92
+ | Screening — cast wide net | 0.05 | ~0.20 | ~29% |
93
+ | Balanced | 0.06 | ~0.41 | ~29% |
94
+ | High confidence hits only | 0.50 | ~0.61 | ~24% |
95
 
96
+ Optimised threshold from val set: **0.06** (F1: 0.358)
97
 
98
+ The model produces a bimodal distribution — proteins it recognises
99
+ score very high (above 0.50), proteins it does not score near zero.
100
+ The flat recall curve from 0.05 to 0.70 reflects this — most
101
+ longevity proteins are either clearly found or clearly missed.
102
 
103
+ ---
104
 
105
+ ## Known Limitations — Read Before Use
106
 
107
+ ### 1. Protein length truncation
108
+ Sequences longer than 512 amino acids are truncated from the
109
+ C-terminus. This causes systematic failures on long proteins where
110
+ the functional domain sits in the C-terminal half:
111
 
112
+ - **MTOR** (2,549 aa): kinase domain at residues 2181-2431 — truncated away
113
+ - **TERT** (1,132 aa): reverse transcriptase domain at 600-900 — truncated away
114
 
115
+ Do not use this model to score proteins above 800 amino acids
116
+ without validating on known examples from that protein family first.
117
 
118
+ ### 2. Family-specific blind spots
119
+ The model learned sirtuin and tumour suppressor sequence features
120
+ well but has insufficient training examples to generalise to:
121
 
122
+ - **Forkhead transcription factors** (FOXO3 scores 0.000 despite
123
+ being a canonical longevity gene and fitting within the 512 aa window)
124
+ - **Large kinases** (truncation compounds this)
125
+ - **Telomerase complex** proteins
126
 
127
+ ### 3. Direction of effect not captured
128
+ The model cannot distinguish between:
129
+ - Pro-longevity proteins (overexpression extends lifespan)
130
+ - Anti-aging-disease proteins (loss of function accelerates aging)
131
 
132
+ Both may score high. A high score means "associated with longevity
133
+ biology" not "activating this protein extends lifespan."
134
 
135
+ ### 4. Not validated experimentally
136
+ Novel predictions are model outputs only. No wet lab validation has
137
+ been performed. TFEB is the strongest prediction based on prior
138
+ literature but this model did not discover TFEB — it independently
139
+ ranked it highly, consistent with existing biology.
140
 
141
+ ### 5. Not for clinical use
142
+ This is a research screening tool. Do not use for any clinical,
143
+ diagnostic, or therapeutic decision-making.
144
 
145
+ ---
146
 
147
+ ## Training Data
148
 
149
+ **Positive set:** GenAge database (genomics.senescence.info)
150
+ - Human GenAge: 306 human longevity-associated genes
151
+ - Model organism GenAge: Pro-Longevity genes only from 4 species
152
+ - C. elegans: 283 genes
153
+ - D. melanogaster: 125 genes
154
+ - M. musculus: 85 genes
155
+ - Total positives: ~574
156
 
157
+ **Negative set:** Swiss-Prot reviewed proteins from same species
158
+ - Sampled proportionally per species (NEG_RATIO=10)
159
+ - Species weights applied: human 2.0x, mouse 1.5x, worm/fly 1.0x
160
+ - "Necessary for fitness" genes excluded from universe entirely
161
+ - Anti-Longevity genes excluded from positives
162
 
163
+ **Filtering:**
164
+ - Sequence length: 50-1500 amino acids
165
+ - Swiss-Prot reviewed only (manually curated)
166
 
167
+ ---
168
 
169
+ ## Training Procedure
170
 
171
+ **Architecture:** ESM-2 150M + LoRA adapters
172
+ - LoRA rank: r=16, alpha=32, dropout=0.15
173
+ - Target modules: query, value attention projections
174
+ - Trainable parameters: ~4.7M of 150M total (3.1%)
175
 
176
+ **Loss function:** Focal loss with contrastive margin penalty
177
+ - gamma=1.0 (softer than standard gamma=2.0)
178
+ - Label smoothing=0.1
179
+ - Contrastive margin=0.30 (explicit separation penalty)
180
+ - Class weights: balanced
181
 
182
+ **Optimiser:** AdamW, lr=2e-4, weight_decay=0.01
183
+ **Schedule:** Cosine with warmup (10% warmup steps)
184
+ **Early stopping:** Patience=4 on val AUPRC
185
+ **Best epoch:** 10 of 20
186
 
187
+ **Hardware:** NVIDIA T4 16GB (Kaggle)
188
+ **Training time:** ~2 hours
189
 
190
+ ---
191
 
192
+ ## How to Use
193
+ ```
194
+ from transformers import AutoTokenizer, EsmForSequenceClassification
195
+ from peft import PeftModel
196
+ import torch
197
+
198
+ # Load model
199
+ base = EsmForSequenceClassification.from_pretrained(
200
+ "facebook/esm2_t30_150M_UR50D",
201
+ num_labels=2,
202
+ ignore_mismatched_sizes=True
203
+ )
204
+ model = PeftModel.from_pretrained(base, "YOUR_USERNAME/longevity-esm2-v6")
205
+ tokenizer = AutoTokenizer.from_pretrained("YOUR_USERNAME/longevity-esm2-v6")
206
+
207
+ device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
208
+ model = model.to(device)
209
+ model.eval()
210
+
211
+ def score_sequence(sequence, threshold=0.06):
212
+ inputs = tokenizer(
213
+ sequence,
214
+ max_length=512,
215
+ padding="max_length",
216
+ truncation=True,
217
+ return_tensors="pt"
218
+ )
219
+ with torch.no_grad():
220
+ outputs = model(
221
+ input_ids=inputs["input_ids"].to(device),
222
+ attention_mask=inputs["attention_mask"].to(device)
223
+ )
224
+ prob = torch.softmax(outputs.logits, dim=1)[:, 1].item()
225
+ return {
226
+ "probability": round(prob, 4),
227
+ "prediction": "Longevity" if prob >= threshold else "Non-longevity",
228
+ "threshold": threshold,
229
+ "warning": "Truncated to 512 aa" if len(sequence) > 512 else None
230
+ }
231
+
232
+ # Example
233
+ result = score_sequence("MKTAYIAKQRQISFVK...")
234
+ print(result)
235
+ ```
236
+
237
+ **Recommended thresholds:**
238
+ - 0.05-0.06 for screening (maximise recall)
239
+ - 0.50 for high-confidence hits only
240
 
241
+ ---
242
 
243
+ ## Experiment History
244
 
245
+ This model is v6 in a series of iterative experiments:
246
 
247
+ | Version | Key change | Test AUPRC |
248
+ |---------|-----------|------------|
249
+ | v1 | Frozen encoder, 186 positives | Collapsed |
250
+ | v2 | LoRA r=8, 277 positives | 0.027 |
251
+ | v3 | ESM-2 150M, multi-species, ~2000 positives | 0.302 |
252
+ | v4 | Pro-Longevity filter, focal loss gamma=2 | 0.250 |
253
+ | v5 | Cleaned species, gamma=1, label smoothing | 0.323 |
254
+ | v6 (this) | Pathway-stratified split, contrastive margin | **0.335** |
255
 
256
+ ---
257
 
258
+ ## Citation
259
 
260
+ If you use this model in research, please cite:
261
+ @misc{elzek2026longevity,
262
+ author = {Elzek, Mo},
263
+ title = {Longevity Protein Classifier: Multi-species ESM-2 Fine-tuning},
264
+ year = {2026},
265
+ publisher = {HuggingFace},
266
+ url = {https://huggingface.co/YOUR_USERNAME/longevity-esm2-v6}
267
+ }
268
 
269
+ ---
270
 
271
+ ## Contact
272
 
273
+ Built by Mo Elzek as part of the London Longevity Network ML project arc.
274
+ Feedback and collaboration welcome.