Text Generation
Transformers
Safetensors
PEFT
llama
protein
ptm
methylation
phosphorylation
ubiquitination
lora
text-generation-inference
Instructions to use jbenbudd/ptm-llama with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jbenbudd/ptm-llama with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jbenbudd/ptm-llama")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("jbenbudd/ptm-llama") model = AutoModelForCausalLM.from_pretrained("jbenbudd/ptm-llama", device_map="auto") - PEFT
How to use jbenbudd/ptm-llama with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jbenbudd/ptm-llama with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jbenbudd/ptm-llama" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jbenbudd/ptm-llama", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/jbenbudd/ptm-llama
- SGLang
How to use jbenbudd/ptm-llama with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jbenbudd/ptm-llama" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jbenbudd/ptm-llama", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jbenbudd/ptm-llama" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jbenbudd/ptm-llama", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use jbenbudd/ptm-llama with Docker Model Runner:
docker model run hf.co/jbenbudd/ptm-llama
Push evaluation artifacts and model card
Browse files- .gitattributes +2 -0
- README.md +163 -2
- calibration_threshold_sweep.png +3 -0
- cross_instruction_ablation.png +0 -0
- cross_instruction_auc.csv +4 -0
- inference_config.json +17 -0
- metrics_summary.json +192 -0
- per_residue_breakdown.csv +7 -0
- test_metrics.png +3 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
calibration_threshold_sweep.png filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
test_metrics.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -9,10 +9,171 @@ tags:
|
|
| 9 |
- lora
|
| 10 |
- peft
|
| 11 |
library_name: transformers
|
|
|
|
| 12 |
---
|
| 13 |
|
| 14 |
# PTM-LLaMA
|
| 15 |
|
| 16 |
-
LoRA
|
| 17 |
|
| 18 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
- lora
|
| 10 |
- peft
|
| 11 |
library_name: transformers
|
| 12 |
+
pipeline_tag: text-generation
|
| 13 |
---
|
| 14 |
|
| 15 |
# PTM-LLaMA
|
| 16 |
|
| 17 |
+
LoRA fine-tune of [`GreatCaptainNemo/ProLLaMA_Stage_1`](https://huggingface.co/GreatCaptainNemo/ProLLaMA_Stage_1) instruction-tuned to predict post-translational modification (PTM) sites for three PTM types in a single adapter: **methylation**, **phosphorylation**, and **ubiquitination**.
|
| 18 |
|
| 19 |
+
## Task
|
| 20 |
+
|
| 21 |
+
Given a 21-residue peptide window and a PTM-type instruction, the model generates the list of modified positions in the format `Sites=<R5,D12,...>` (residue letter + 1-indexed position within the window). The PTM type to predict is selected by the instruction prompt; the output format is shared across PTM types.
|
| 22 |
+
|
| 23 |
+
## Inference architecture
|
| 24 |
+
|
| 25 |
+
Full-protein inference for a single PTM type proceeds in three steps:
|
| 26 |
+
|
| 27 |
+
1. **Sliding windows.** A 21-residue window is slid across the input sequence with stride 5; a tail window is appended so the final residues are covered.
|
| 28 |
+
2. **Per-window generation.** The model generates `Sites=<...>` for each window, prompted with the PTM-type instruction. Each predicted residue letter is validated against the window sequence at the indicated position; mismatches are discarded. Validated predictions are mapped to full-protein coordinates and accumulated into a per-residue **consensus score**, defined as `(# windows predicting the residue as a site) / (# windows covering the residue)`.
|
| 29 |
+
3. **Thresholding.** A per-PTM-type F1-optimal threshold (derived from the held-out calibration set) is applied to the consensus scores.
|
| 30 |
+
|
| 31 |
+
The per-PTM-type thresholds, windowing parameters, and prompt template are persisted in [`inference_config.json`](./inference_config.json).
|
| 32 |
+
|
| 33 |
+
### Reference implementation
|
| 34 |
+
|
| 35 |
+
```python
|
| 36 |
+
import re, json, torch
|
| 37 |
+
from huggingface_hub import hf_hub_download
|
| 38 |
+
from transformers import AutoTokenizer, AutoModelForCausalLM
|
| 39 |
+
|
| 40 |
+
REPO = "jbenbudd/ptm-llama"
|
| 41 |
+
|
| 42 |
+
cfg = json.load(open(hf_hub_download(REPO, "inference_config.json")))
|
| 43 |
+
tok = AutoTokenizer.from_pretrained(REPO)
|
| 44 |
+
if tok.pad_token is None:
|
| 45 |
+
tok.pad_token = tok.unk_token
|
| 46 |
+
mdl = AutoModelForCausalLM.from_pretrained(
|
| 47 |
+
REPO, torch_dtype=torch.float16, device_map="auto"
|
| 48 |
+
).eval()
|
| 49 |
+
|
| 50 |
+
SITE_RE = re.compile(r"^([A-Z])(\d+)$")
|
| 51 |
+
SITES_RE = re.compile(r"Sites=<([^>]*)>")
|
| 52 |
+
|
| 53 |
+
|
| 54 |
+
@torch.no_grad()
|
| 55 |
+
def predict_sites(seq: str, ptm_type: str):
|
| 56 |
+
"""Predict PTM-site positions in a full protein for a given PTM type.
|
| 57 |
+
|
| 58 |
+
Returns sites like ['K161', 'S203'] in 1-indexed full-protein coords.
|
| 59 |
+
"""
|
| 60 |
+
if ptm_type not in cfg["instructions"]:
|
| 61 |
+
raise ValueError(f"Unknown PTM type {ptm_type}; supported: {list(cfg['instructions'])}")
|
| 62 |
+
w, s = cfg["window_size"], cfg["stride"]
|
| 63 |
+
t = cfg["consensus_thresholds"][ptm_type]
|
| 64 |
+
instruction = cfg["instructions"][ptm_type]
|
| 65 |
+
L = len(seq)
|
| 66 |
+
|
| 67 |
+
if L <= w:
|
| 68 |
+
starts = [0]
|
| 69 |
+
else:
|
| 70 |
+
starts = list(range(0, L - w + 1, s))
|
| 71 |
+
if starts[-1] + w < L:
|
| 72 |
+
starts.append(L - w)
|
| 73 |
+
|
| 74 |
+
covered, predicted = [0] * L, [0] * L
|
| 75 |
+
for st in starts:
|
| 76 |
+
win = seq[st:st + w]
|
| 77 |
+
prompt = cfg["prompt_template"].format(instruction=instruction, input=f"Seq=<{win}>")
|
| 78 |
+
enc = tok(prompt, return_tensors="pt").to(mdl.device)
|
| 79 |
+
out = mdl.generate(
|
| 80 |
+
**enc, max_new_tokens=cfg["max_new_tokens"], do_sample=False,
|
| 81 |
+
pad_token_id=tok.pad_token_id,
|
| 82 |
+
)
|
| 83 |
+
text = tok.decode(out[0][enc.input_ids.shape[1]:], skip_special_tokens=True)
|
| 84 |
+
|
| 85 |
+
for i in range(st, min(st + w, L)):
|
| 86 |
+
covered[i] += 1
|
| 87 |
+
m = SITES_RE.search(text)
|
| 88 |
+
if not m:
|
| 89 |
+
continue
|
| 90 |
+
for part in m.group(1).split(","):
|
| 91 |
+
mm = SITE_RE.match(part.strip())
|
| 92 |
+
if not mm:
|
| 93 |
+
continue
|
| 94 |
+
letter, pos_local = mm.group(1), int(mm.group(2))
|
| 95 |
+
pos_full = st + pos_local
|
| 96 |
+
if 1 <= pos_full <= L and seq[pos_full - 1] == letter:
|
| 97 |
+
predicted[pos_full - 1] += 1
|
| 98 |
+
|
| 99 |
+
return [f"{seq[i]}{i + 1}" for i in range(L)
|
| 100 |
+
if covered[i] > 0 and predicted[i] / covered[i] >= t]
|
| 101 |
+
|
| 102 |
+
|
| 103 |
+
# Example usage.
|
| 104 |
+
sites = predict_sites(
|
| 105 |
+
"MASDEGKLFVGGLSFDTNEQALEQVFSKYGQISEVVVVKDRETQRSRGFGFVTFENIDDAKDAMMAMNGK",
|
| 106 |
+
ptm_type="Phosphorylation",
|
| 107 |
+
)
|
| 108 |
+
print(sites)
|
| 109 |
+
```
|
| 110 |
+
|
| 111 |
+
## Training
|
| 112 |
+
|
| 113 |
+
- **Base model:** `GreatCaptainNemo/ProLLaMA_Stage_1`
|
| 114 |
+
- **Method:** LoRA SFT via `trl.SFTTrainer` + `peft.LoraConfig`
|
| 115 |
+
- **LoRA config:** r=64, alpha=128, dropout=0.05, target modules = q,k,v,o,gate,down,up_proj
|
| 116 |
+
- **Optimizer:** AdamW, lr=3e-4, cosine schedule, warmup=40 steps, max_grad_norm=1.0
|
| 117 |
+
- **Batching:** per-device batch 16 × grad_accum 8 (effective 128) at bf16, max_seq_length 2048
|
| 118 |
+
- **Epochs:** up to 8, with `EarlyStoppingCallback(patience=3)` on `eval_loss` and `load_best_model_at_end=True`
|
| 119 |
+
- **Source data:** `datasets/all_ptm_sites_site_level.csv` (long format: one row per annotated PTM site across methylation, phosphorylation, and ubiquitination)
|
| 120 |
+
- **Split:** protein-level partition with seed `42` and ratios 0.80 / 0.10 / 0.10 (train / calibration / test). All annotations of a given `uniprot_id` are assigned to a single split.
|
| 121 |
+
- **Training distribution:** natural — the relative training-set abundance across PTM types reflects the natural distribution of annotated sites in the source data (phosphorylation ≫ ubiquitination ≫ methylation). Each `(protein, PTM_type)` record contributes all sliding windows over its sequence, including windows with no in-window site of that PTM type as in-context negatives.
|
| 122 |
+
|
| 123 |
+

|
| 124 |
+
|
| 125 |
+
## Evaluation methodology
|
| 126 |
+
|
| 127 |
+
Two disjoint protein-level partitions are used to separate per-PTM-type threshold selection from final metric reporting:
|
| 128 |
+
|
| 129 |
+
- **Calibration set** (5,943 proteins): the F1-optimal threshold over the per-residue consensus ROC is selected per PTM type.
|
| 130 |
+
- **Test set** (5,944 proteins): the calibration-derived per-PTM-type thresholds are applied without modification; the metrics below are unbiased point estimates of generalization.
|
| 131 |
+
|
| 132 |
+
### Per-PTM-type results
|
| 133 |
+
|
| 134 |
+
| PTM type | Test residues | Positive prevalence | Calibration AUC | Locked t | Test AUC | Accuracy | Precision | Recall | Specificity | F1 | TN / FP / FN / TP |
|
| 135 |
+
|---|---|---|---|---|---|---|---|---|---|---|---|
|
| 136 |
+
| Methylation | 399,838 | 0.37% | 0.619 | 0.333 | 0.613 | 99.44% | 20.18% | 16.70% | 99.75% | 18.28% | 397,362 / 985 / 1,242 / 249 |
|
| 137 |
+
| Phosphorylation | 1,732,712 | 2.52% | 0.652 | 0.200 | 0.653 | 96.33% | 29.45% | 32.61% | 97.98% | 30.95% | 1,654,945 / 34,107 / 29,424 / 14,236 |
|
| 138 |
+
| Ubiquitination | 2,321,360 | 0.77% | 0.762 | 0.200 | 0.765 | 97.85% | 18.91% | 54.75% | 98.19% | 28.11% | 2,261,794 / 41,772 / 8,051 / 9,743 |
|
| 139 |
+
|
| 140 |
+

|
| 141 |
+
|
| 142 |
+
### Per-residue-type breakdown on test
|
| 143 |
+
|
| 144 |
+
| PTM type | Residue | Total | Positive | AUC | Accuracy | Precision | Recall | Specificity | F1 |
|
| 145 |
+
|---|---|---|---|---|---|---|---|---|---|
|
| 146 |
+
| Methylation | K | 24,891 | 699 | 0.516 | 96.99% | 14.29% | 1.43% | 99.75% | 2.60% |
|
| 147 |
+
| Methylation | R | 24,176 | 746 | 0.674 | 94.08% | 20.53% | 32.04% | 96.05% | 25.03% |
|
| 148 |
+
| Phosphorylation | S | 149,662 | 26,557 | 0.611 | 73.21% | 30.83% | 40.99% | 80.16% | 35.19% |
|
| 149 |
+
| Phosphorylation | T | 95,276 | 11,688 | 0.574 | 82.52% | 26.37% | 23.71% | 90.75% | 24.97% |
|
| 150 |
+
| Phosphorylation | Y | 44,179 | 5,211 | 0.531 | 85.11% | 22.90% | 11.09% | 95.01% | 14.95% |
|
| 151 |
+
| Ubiquitination | K | 146,021 | 17,792 | 0.623 | 65.88% | 18.91% | 54.76% | 67.42% | 28.12% |
|
| 152 |
+
|
| 153 |
+
### Cross-instruction ablation
|
| 154 |
+
|
| 155 |
+
For a sub-sample of test proteins, sliding-window inference was run three times — once per PTM-type instruction — on the same protein sequences. AUC is reported against each PTM type's ground-truth labels.
|
| 156 |
+
|
| 157 |
+
A working instruction-tuned model should have higher AUC on the diagonal (matched instruction) than off-diagonal (mismatched instruction).
|
| 158 |
+
|
| 159 |
+
| | Methylation | Phosphorylation | Ubiquitination |
|
| 160 |
+
|---|---|---|---|
|
| 161 |
+
| **Methylation** | 0.657 | 0.485 | 0.580 |
|
| 162 |
+
| **Phosphorylation** | 0.498 | 0.656 | 0.491 |
|
| 163 |
+
| **Ubiquitination** | 0.514 | 0.490 | 0.743 |
|
| 164 |
+
|
| 165 |
+
Instruction-following rate (fraction of windows whose output differs across the three prompts on a 50-protein sub-sample): **15.58%**.
|
| 166 |
+
|
| 167 |
+

|
| 168 |
+
|
| 169 |
+

|
| 170 |
+
|
| 171 |
+
## Limitations
|
| 172 |
+
|
| 173 |
+
- **Window-local outputs.** The model emits positions inside a 21-residue window. Full-protein predictions are produced by the sliding-window aggregation described in *Inference architecture*; very short proteins (`length < 21`) are scored as a single window.
|
| 174 |
+
- **No structural context.** The model sees only primary sequence; structurally-disfavored false positives cannot be filtered without external 3D information.
|
| 175 |
+
- **Three PTM types only.** Predictions are well-defined for methylation, phosphorylation, and ubiquitination. Generalization to other PTM types is not evaluated.
|
| 176 |
+
|
| 177 |
+
## Reproduction
|
| 178 |
+
|
| 179 |
+
Execute `training/train_ptm_llama.ipynb` followed by `evaluation/evaluate_ptm_llama.ipynb` from the source repository. Both notebooks are self-contained and intended for execution in Google Colab. The protein-level split is deterministic given `SPLIT_SEED = 42`.
|
calibration_threshold_sweep.png
ADDED
|
Git LFS Details
|
cross_instruction_ablation.png
ADDED
|
cross_instruction_auc.csv
ADDED
|
@@ -0,0 +1,4 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
,Methylation,Phosphorylation,Ubiquitination
|
| 2 |
+
Methylation,0.6565676387753249,0.4845552830017262,0.5799763786681202
|
| 3 |
+
Phosphorylation,0.49783048886317616,0.6563748251514029,0.49054773775289695
|
| 4 |
+
Ubiquitination,0.5141878873945271,0.4897625870439934,0.7429060321209946
|
inference_config.json
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"window_size": 21,
|
| 3 |
+
"stride": 5,
|
| 4 |
+
"max_new_tokens": 64,
|
| 5 |
+
"prompt_template": "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\n{instruction}\n{input}\n\n### Response:\n",
|
| 6 |
+
"instructions": {
|
| 7 |
+
"Methylation": "[Predict the methylation sites given the peptide sequence]",
|
| 8 |
+
"Phosphorylation": "[Predict the phosphorylation sites given the peptide sequence]",
|
| 9 |
+
"Ubiquitination": "[Predict the ubiquitination sites given the peptide sequence]"
|
| 10 |
+
},
|
| 11 |
+
"consensus_thresholds": {
|
| 12 |
+
"Methylation": 0.3333333432674408,
|
| 13 |
+
"Phosphorylation": 0.20000000298023224,
|
| 14 |
+
"Ubiquitination": 0.20000000298023224
|
| 15 |
+
},
|
| 16 |
+
"output_format": "Sites=<X1,X2,...> where Xi is one-letter residue + 1-indexed position within the window"
|
| 17 |
+
}
|
metrics_summary.json
ADDED
|
@@ -0,0 +1,192 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model_repo": "jbenbudd/ptm-llama",
|
| 3 |
+
"data_csv": "datasets/all_ptm_sites_site_level.csv",
|
| 4 |
+
"split_seed": 42,
|
| 5 |
+
"split_ratios": {
|
| 6 |
+
"train": 0.8,
|
| 7 |
+
"calibration": 0.1,
|
| 8 |
+
"test": 0.1
|
| 9 |
+
},
|
| 10 |
+
"window_size": 21,
|
| 11 |
+
"stride": 5,
|
| 12 |
+
"calibration": {
|
| 13 |
+
"Methylation": {
|
| 14 |
+
"n_proteins": 639,
|
| 15 |
+
"n_residues_evaluated": 423176,
|
| 16 |
+
"positive_prevalence": 0.0036675047734276048,
|
| 17 |
+
"roc_auc": 0.6188406072489271,
|
| 18 |
+
"locked_threshold": 0.3333333432674408,
|
| 19 |
+
"f1_at_threshold": 0.2036967182195398,
|
| 20 |
+
"precision_at_threshold": 0.2456778889899909,
|
| 21 |
+
"recall_at_threshold": 0.17396907216494845
|
| 22 |
+
},
|
| 23 |
+
"Phosphorylation": {
|
| 24 |
+
"n_proteins": 2970,
|
| 25 |
+
"n_residues_evaluated": 1681943,
|
| 26 |
+
"positive_prevalence": 0.027348132487248378,
|
| 27 |
+
"roc_auc": 0.6524042839293995,
|
| 28 |
+
"locked_threshold": 0.20000000298023224,
|
| 29 |
+
"f1_at_threshold": 0.32371403661726245,
|
| 30 |
+
"precision_at_threshold": 0.32454875223984964,
|
| 31 |
+
"recall_at_threshold": 0.3228836036349407
|
| 32 |
+
},
|
| 33 |
+
"Ubiquitination": {
|
| 34 |
+
"n_proteins": 3914,
|
| 35 |
+
"n_residues_evaluated": 2340476,
|
| 36 |
+
"positive_prevalence": 0.007890702575031746,
|
| 37 |
+
"roc_auc": 0.7622045108091328,
|
| 38 |
+
"locked_threshold": 0.20000000298023224,
|
| 39 |
+
"f1_at_threshold": 0.2864640844363501,
|
| 40 |
+
"precision_at_threshold": 0.19483078123476055,
|
| 41 |
+
"recall_at_threshold": 0.540827377084687
|
| 42 |
+
}
|
| 43 |
+
},
|
| 44 |
+
"test": {
|
| 45 |
+
"Methylation": {
|
| 46 |
+
"n_proteins": 649,
|
| 47 |
+
"n_residues_evaluated": 399838,
|
| 48 |
+
"positive_prevalence": 0.003729010249150906,
|
| 49 |
+
"threshold_applied": 0.3333333432674408,
|
| 50 |
+
"accuracy": 0.9944302442489208,
|
| 51 |
+
"precision": 0.20178282009724474,
|
| 52 |
+
"recall": 0.16700201207243462,
|
| 53 |
+
"specificity": 0.9975272814907605,
|
| 54 |
+
"f1": 0.18275229357798167,
|
| 55 |
+
"roc_auc": 0.6129889531736042,
|
| 56 |
+
"confusion_matrix": {
|
| 57 |
+
"TN": 397362,
|
| 58 |
+
"FP": 985,
|
| 59 |
+
"FN": 1242,
|
| 60 |
+
"TP": 249
|
| 61 |
+
}
|
| 62 |
+
},
|
| 63 |
+
"Phosphorylation": {
|
| 64 |
+
"n_proteins": 3040,
|
| 65 |
+
"n_residues_evaluated": 1732712,
|
| 66 |
+
"positive_prevalence": 0.02519749387087987,
|
| 67 |
+
"threshold_applied": 0.20000000298023224,
|
| 68 |
+
"accuracy": 0.9633343567771216,
|
| 69 |
+
"precision": 0.2944790352274373,
|
| 70 |
+
"recall": 0.32606504809894643,
|
| 71 |
+
"specificity": 0.979807016006612,
|
| 72 |
+
"f1": 0.3094681695162114,
|
| 73 |
+
"roc_auc": 0.6533860232975732,
|
| 74 |
+
"confusion_matrix": {
|
| 75 |
+
"TN": 1654945,
|
| 76 |
+
"FP": 34107,
|
| 77 |
+
"FN": 29424,
|
| 78 |
+
"TP": 14236
|
| 79 |
+
}
|
| 80 |
+
},
|
| 81 |
+
"Ubiquitination": {
|
| 82 |
+
"n_proteins": 3850,
|
| 83 |
+
"n_residues_evaluated": 2321360,
|
| 84 |
+
"positive_prevalence": 0.007665334114484612,
|
| 85 |
+
"threshold_applied": 0.20000000298023224,
|
| 86 |
+
"accuracy": 0.9785371506358341,
|
| 87 |
+
"precision": 0.1891293797922935,
|
| 88 |
+
"recall": 0.5475441159941553,
|
| 89 |
+
"specificity": 0.9818663758711493,
|
| 90 |
+
"f1": 0.2811467486185055,
|
| 91 |
+
"roc_auc": 0.7653999295817027,
|
| 92 |
+
"confusion_matrix": {
|
| 93 |
+
"TN": 2261794,
|
| 94 |
+
"FP": 41772,
|
| 95 |
+
"FN": 8051,
|
| 96 |
+
"TP": 9743
|
| 97 |
+
}
|
| 98 |
+
}
|
| 99 |
+
},
|
| 100 |
+
"cross_instruction_auc": {
|
| 101 |
+
"Methylation": {
|
| 102 |
+
"Methylation": 0.6565676387753249,
|
| 103 |
+
"Phosphorylation": 0.49783048886317616,
|
| 104 |
+
"Ubiquitination": 0.5141878873945271
|
| 105 |
+
},
|
| 106 |
+
"Phosphorylation": {
|
| 107 |
+
"Methylation": 0.4845552830017262,
|
| 108 |
+
"Phosphorylation": 0.6563748251514029,
|
| 109 |
+
"Ubiquitination": 0.4897625870439934
|
| 110 |
+
},
|
| 111 |
+
"Ubiquitination": {
|
| 112 |
+
"Methylation": 0.5799763786681202,
|
| 113 |
+
"Phosphorylation": 0.49054773775289695,
|
| 114 |
+
"Ubiquitination": 0.7429060321209946
|
| 115 |
+
}
|
| 116 |
+
},
|
| 117 |
+
"instruction_follow_rate": 0.1557890855457227,
|
| 118 |
+
"per_residue_breakdown": [
|
| 119 |
+
{
|
| 120 |
+
"ptm_type": "Methylation",
|
| 121 |
+
"residue": "K",
|
| 122 |
+
"n_residues": 24891,
|
| 123 |
+
"n_positive": 699,
|
| 124 |
+
"auc": 0.5158489475706036,
|
| 125 |
+
"accuracy": 0.9699088023783697,
|
| 126 |
+
"precision": 0.14285714285714285,
|
| 127 |
+
"recall": 0.01430615164520744,
|
| 128 |
+
"specificity": 0.9975198412698413,
|
| 129 |
+
"f1": 0.02600780234070221
|
| 130 |
+
},
|
| 131 |
+
{
|
| 132 |
+
"ptm_type": "Methylation",
|
| 133 |
+
"residue": "R",
|
| 134 |
+
"n_residues": 24176,
|
| 135 |
+
"n_positive": 746,
|
| 136 |
+
"auc": 0.6735408592590558,
|
| 137 |
+
"accuracy": 0.9407677035076109,
|
| 138 |
+
"precision": 0.20532646048109965,
|
| 139 |
+
"recall": 0.3203753351206434,
|
| 140 |
+
"specificity": 0.9605206999573197,
|
| 141 |
+
"f1": 0.250261780104712
|
| 142 |
+
},
|
| 143 |
+
{
|
| 144 |
+
"ptm_type": "Phosphorylation",
|
| 145 |
+
"residue": "S",
|
| 146 |
+
"n_residues": 149662,
|
| 147 |
+
"n_positive": 26557,
|
| 148 |
+
"auc": 0.6111038426325143,
|
| 149 |
+
"accuracy": 0.732096323716107,
|
| 150 |
+
"precision": 0.30830879021295876,
|
| 151 |
+
"recall": 0.40994841284783673,
|
| 152 |
+
"specificity": 0.8015921367937939,
|
| 153 |
+
"f1": 0.3519371575425496
|
| 154 |
+
},
|
| 155 |
+
{
|
| 156 |
+
"ptm_type": "Phosphorylation",
|
| 157 |
+
"residue": "T",
|
| 158 |
+
"n_residues": 95276,
|
| 159 |
+
"n_positive": 11688,
|
| 160 |
+
"auc": 0.573697049783009,
|
| 161 |
+
"accuracy": 0.8252130652000503,
|
| 162 |
+
"precision": 0.26372894260968877,
|
| 163 |
+
"recall": 0.2370807665982204,
|
| 164 |
+
"specificity": 0.9074508302627171,
|
| 165 |
+
"f1": 0.2496958774498761
|
| 166 |
+
},
|
| 167 |
+
{
|
| 168 |
+
"ptm_type": "Phosphorylation",
|
| 169 |
+
"residue": "Y",
|
| 170 |
+
"n_residues": 44179,
|
| 171 |
+
"n_positive": 5211,
|
| 172 |
+
"auc": 0.530815164126421,
|
| 173 |
+
"accuracy": 0.8510830937775866,
|
| 174 |
+
"precision": 0.22900158478605387,
|
| 175 |
+
"recall": 0.11091920936480522,
|
| 176 |
+
"specificity": 0.9500615889960994,
|
| 177 |
+
"f1": 0.14945054945054945
|
| 178 |
+
},
|
| 179 |
+
{
|
| 180 |
+
"ptm_type": "Ubiquitination",
|
| 181 |
+
"residue": "K",
|
| 182 |
+
"n_residues": 146021,
|
| 183 |
+
"n_positive": 17792,
|
| 184 |
+
"auc": 0.6234033873578442,
|
| 185 |
+
"accuracy": 0.6588093493401634,
|
| 186 |
+
"precision": 0.1891293797922935,
|
| 187 |
+
"recall": 0.5476056654676259,
|
| 188 |
+
"specificity": 0.6742390566876447,
|
| 189 |
+
"f1": 0.28115486170228116
|
| 190 |
+
}
|
| 191 |
+
]
|
| 192 |
+
}
|
per_residue_breakdown.csv
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
ptm_type,residue,n_residues,n_positive,auc,accuracy,precision,recall,specificity,f1
|
| 2 |
+
Methylation,K,24891,699,0.5158489475706036,0.9699088023783697,0.14285714285714285,0.01430615164520744,0.9975198412698413,0.02600780234070221
|
| 3 |
+
Methylation,R,24176,746,0.6735408592590558,0.9407677035076109,0.20532646048109965,0.3203753351206434,0.9605206999573197,0.250261780104712
|
| 4 |
+
Phosphorylation,S,149662,26557,0.6111038426325143,0.732096323716107,0.30830879021295876,0.40994841284783673,0.8015921367937939,0.3519371575425496
|
| 5 |
+
Phosphorylation,T,95276,11688,0.573697049783009,0.8252130652000503,0.26372894260968877,0.2370807665982204,0.9074508302627171,0.2496958774498761
|
| 6 |
+
Phosphorylation,Y,44179,5211,0.530815164126421,0.8510830937775866,0.22900158478605387,0.11091920936480522,0.9500615889960994,0.14945054945054945
|
| 7 |
+
Ubiquitination,K,146021,17792,0.6234033873578442,0.6588093493401634,0.1891293797922935,0.5476056654676259,0.6742390566876447,0.28115486170228116
|
test_metrics.png
ADDED
|
Git LFS Details
|