pyaging-data / PAD000022_README.md
lucascamillomd's picture
Add real PAD000022 proteomics tutorial dataset with provenance
ab0b228 verified
|
Raw History Blame Contribute Delete
3.62 kB

PAD000022 proteomics example

This example contains 32 real human plasma samples and 134 measured proteins from the Olink Explore 3072 portion of Human Plasma Discovery Proteomics, deposited by Sara Ahadi in PRIDE, PAD000022. The PRIDE record releases these source data under CC0 1.0. The clock weights have their own research-use terms.

Please cite Kirsher, D. Y., Chand, S., Phong, A., Nguyen, B., Szoke, B. G., and Ahadi, S. Current landscape of plasma proteomics from technical innovations to biological insights and biomarker discovery. Communications Chemistry 8, 279, 2025. Publication, dataset.

Files

  • PAD000022_subset.pkl is the pandas DataFrame used by the tutorial.
  • PAD000022_subset.csv contains the same table in an open text format.
  • PAD000022_assays.csv identifies each assay by its source name, Olink ID, UniProt annotation and panel.
  • PAD000022_provenance.json records source URLs, SHA-256 checksums, sample identifiers, preparation steps and output checksums.

Rows retain the published Olink sample IDs. subject_id contains the source de-identified subject identifier and belongs in metadata_cols. All other columns are source protein names with float64 NPX values.

Selection and preparation

The subset excludes controls, technical replicates and samples marked "Sample not analyzed." in the authors' sample spreadsheet. Among the remaining subjects, it selects the first 32 sorted Olink sample IDs with finite NPX and PASS in both QC_Warning and Assay_Warning for every selected assay. Spreadsheet sample IDs have leading zero padding, which is removed solely to join them to the source NPX CSV. Each retained row is a distinct subject.

The 134 assays are the union of features for the brain, heart and kidney OrganAge models. Every selected protein has exactly one Olink assay in this deposit, so no duplicate-assay aggregation or gene-name substitution is used. The matrix contains unchanged measured NPX, including values below LOD. There is no synthetic data, imputation, exponentiation, centering or scaling.

Use and limits

This subset covers all features of organagechronologicalbrain, organagechronologicalheart, organagechronologicalkidney, and organagemortalitybrain. It does not contain the full Explore 3072 panel. Exact chronological ages are unavailable in the public subject annotations, so this example cannot run PAC, HPS or PAOPAC. The tutorial does not substitute age-group midpoints for measured ages.

The source reports plate-control-normalized NPX, a log2 relative-abundance measure. These values have not been harmonized to the OrganAge UK Biobank training data. The example demonstrates scoring and feature coverage, not clock accuracy or calibration in this cohort. Chronological models return years; the mortality model returns relative natural-log mortality hazard. Changing assays or sample type requires separate preparation.

Reproduce

From the pyaging repository root:

uv run --with openpyxl python tutorials/data/prepare_pad000022.py \
    --source-dir /tmp/pad000022 --output-dir hf_static_data/repo

The script verifies both original source files against fixed SHA-256 hashes before selecting any data. The assay manifest is committed beside the script. The pandas pickle uses protocol 4; the CSV provides a version-independent copy.