f-galkin commited on
Commit
366add9
·
verified ·
1 Parent(s): 837d0b8

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +40 -61
README.md CHANGED
@@ -16,78 +16,57 @@ tags:
16
  - insilicomedicine
17
  language:
18
  - en
19
- model-index:
20
- - name: longevity-llm
21
- results:
22
- - task:
23
- type: text-generation
24
- dataset:
25
- name: LongeBench (aging clocks, methylation, proteomics, NHANES)
26
- type: in-house
27
- metrics:
28
- - type: accuracy
29
- value: 0.565
30
- name: Classification avg accuracy (29 tasks)
31
- - type: mae
32
- value: 21.71
33
- name: Regression avg MAE (8 tasks)
34
- - type: jaccard
35
- value: 0.510
36
- name: Generation avg Jaccard (7 tasks)
37
  ---
38
 
39
  # Longevity-LLM (L-LLM)
40
 
41
- A domain-adapted **Qwen3.5-9B** for aging and longevity biology. L-LLM is the
42
- phase-1 continued-pretraining + phase-2 supervised fine-tuning (with a third
43
- reasoning-augmented continuation pass) of Qwen3.5-9B on ≈1.88M examples
44
- (≈7.0B tokens) drawn from clinical aging, epigenomics, transcriptomics,
45
- proteomics, and genetics tasks. The two trained LoRA adapters were
46
- concatenated into a single rank-64 LoRA and merged into the base weights to
47
- produce this standalone bf16 checkpoint.
48
 
49
  ## Methods
50
 
51
  L-LLM was built by LoRA fine-tuning Qwen3.5-9B, a 9B-parameter hybrid
52
  transformer that interleaves Gated DeltaNet linear-attention layers with
53
- standard self-attention in a 3:1 ratio.
54
-
55
- Training data: ≈1.88M examples / ≈7.0B tokens over 50 tasks in seven datasets.
56
- Two were domain-knowledge priors — **QPT_Gene** (≈224K examples: UniProt
57
- annotations, GO terms, PPIs, pathway membership) and **QPT_Clock** (≈926K
58
- examples: published aging-clock formulas, CpG-site coefficients, gene-to-clock
59
- mappings, feature importance from Biolearn). The remaining ≈727K examples
60
- covered prediction tasks across **Clinical** (NHANES age/mortality),
61
- **Epigenomics** (GEO DNAm age, CpG methylation, clock proxies),
62
- **Transcriptomics** (GTEx age, TCGA survival, expression generation),
63
- **Proteomics** (Olink age, protein generation, clock proxies), and
64
- **Genetics** (OpenGenes expression directionality, SynergyAge lifespan,
65
- CellAge senescence, anti-aging target classification). ≈90K prediction
66
- examples carried frontier-model thinking-trace reasoning. A continuation
67
- corpus of 286K reasoning-augmented examples (≈1.13B tokens) was assembled by
68
- re-mixing prediction tasks with additional reasoning traces.
69
-
70
- Three training stages on Qwen3.5-9B:
71
-
72
- 1. **Continued pretraining** (phase 1) — rank 32, α=32, rsLoRA scaling,
73
- dropout 0.0, LR 2 × 10⁻⁵, 3 epochs, raw text packed into 4,096-token
74
- blocks, effective BS 32.
75
- 2. **Supervised fine-tuning** (phase 2) — rank 32, α=32, standard α/r
76
- scaling, dropout 0.05, LR 1 × 10⁻⁴, 3 epochs, conversation format,
77
- effective BS 16. Initialized from the phase-1 adapter via LLaMA-Factory's
78
- `adapter_name_or_path`.
79
- 3. **Reasoning continuation** (phase 2-continue) — continued the SFT adapter
80
- from its 9,500-step checkpoint for ≈2,504 additional steps (≈1 epoch) on
81
- the 286K reasoning-augmented corpus, LR 3 × 10⁻⁵, 3% warmup, context
82
- 32,768 with example packing, effective BS 16.
83
 
84
  All adapters targeted the 12 linear projections including the GatedDeltaNet
85
- modules (`in_proj_qkv`, `in_proj_z`, `in_proj_a`, `in_proj_b`, `out_proj`).
86
- After training, the CPT adapter and the continued-SFT adapter were
87
- **concatenated into a single rank-64 LoRA and merged into the base weights**
88
- to produce this checkpoint. DeepSpeed ZeRO-2 on 2× NVIDIA H100 NVL 94GB GPUs,
89
- bf16, tf32 matmul, gradient checkpointing, Liger fused kernels (continuation
90
- stage), flash-attention 2 for self-attention layers.
91
 
92
  ## Example usage
93
 
 
16
  - insilicomedicine
17
  language:
18
  - en
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
19
  ---
20
 
21
  # Longevity-LLM (L-LLM)
22
 
23
+ A domain-adapted **Qwen3.5-9B** for aging and longevity biology. L-LLM is
24
+ the result of continued pretraining + supervised fine-tuning + a
25
+ reasoning-augmented continuation pass on a multi-domain corpus spanning
26
+ clinical aging, epigenomics, transcriptomics, proteomics, and genetics.
27
+ The two trained LoRA adapters were concatenated into a single rank-64 LoRA
28
+ and merged into the base weights to produce this standalone bf16
29
+ checkpoint.
30
 
31
  ## Methods
32
 
33
  L-LLM was built by LoRA fine-tuning Qwen3.5-9B, a 9B-parameter hybrid
34
  transformer that interleaves Gated DeltaNet linear-attention layers with
35
+ standard self-attention in a 3:1 ratio. Training data was assembled across
36
+ three main domains:
37
+
38
+ | Domain | Sources |
39
+ |---|---|
40
+ | Knowledge priors | UniProt protein/gene annotations, Gene Ontology, protein–protein interactions, pathway membership; published aging-clock formulas and CpG-site coefficients (Biolearn) |
41
+ | Clinical & epidemiology | NHANES (age, mortality) |
42
+ | Epigenomics | GEO DNA-methylation cohorts, CpG methylation profiles, aging-clock proxy tasks |
43
+ | Transcriptomics | GTEx (tissue age), TCGA (cancer survival), expression-profile generation |
44
+ | Proteomics | Olink plasma-proteomics panels, proteomic clock proxy tasks |
45
+ | Genetics & longevity | OpenGenes (expression directionality), SynergyAge (lifespan), CellAge (senescence), anti-aging target classification |
46
+ | Reasoning corpus | Prediction tasks augmented with frontier-model chain-of-thought traces |
47
+
48
+ Approximate scale across all domains: ≈10⁶ training prompts at the order of
49
+ several billion tokens total. Exact composition, prompt counts, and token
50
+ counts will be reported in the forthcoming preprint.
51
+
52
+ Training proceeded in three stages on Qwen3.5-9B:
53
+
54
+ 1. **Continued pretraining** — knowledge priors only, raw text packed into
55
+ 4,096-token blocks. Rank-32 LoRA with rsLoRA scaling, LR 2 × 10⁻⁵,
56
+ 3 epochs.
57
+ 2. **Supervised fine-tuning** — aging prediction tasks in conversation
58
+ format. Rank-32 LoRA initialized from the phase-1 adapter, standard α/r
59
+ scaling, LR 1 × 10⁻⁴, 3 epochs.
60
+ 3. **Reasoning continuation** — continued the SFT adapter on the reasoning
61
+ corpus, LR 3 × 10⁻⁵, ≈1 epoch, context length 32,768 with example
62
+ packing.
 
 
63
 
64
  All adapters targeted the 12 linear projections including the GatedDeltaNet
65
+ modules. After training, the CPT and continued-SFT adapters were
66
+ concatenated into a single rank-64 LoRA and merged into the base weights to
67
+ produce this checkpoint. DeepSpeed ZeRO-2 on 2× NVIDIA H100 NVL 94 GB,
68
+ bf16, flash-attention 2 on self-attention layers. Full details in the
69
+ forthcoming preprint.
 
70
 
71
  ## Example usage
72