Donor-Liver Macrosteatosis Models

Frozen weights for predicting donor-liver macroscopic steatosis β€” biopsy-defined macrovesicular fat β‰₯ 30% (MACRO_FAT_LI_DON) β€” from donor CT and transplant-registry variables. Three models at increasing modality coverage:

Model Weights Inputs
No-vision ensemble no_vision.pkl tabular features (clinical / body-composition / organ-HU / HU-histogram / topology)
SuPreM vision vision_suprem.pt 96Β³ liver CT volume
Full ensemble full.pkl + full_vision.pt tabular features and 96Β³ CT

The tabular models stack LightGBM + XGBoost + TabPFN with a logistic-regression meta-learner; the full ensemble adds a SuPreM SwinUNETR image probability as a fourth base learner. All are trained on the 5-class fat bin and report the binary probability of fat β‰₯ 30%.

⚠️ Research use only. These models are for retrospective research and are not a medical device and not for clinical decision-making. Donor allocation decisions must not be based on this output.

Code

The training/inference code lives in the companion GitHub repository: https://github.com/sayoni-c98/liver-transplant-project (no_vision_model.py, vision_model.py, full_model.py, common.py). Each script exposes train and predict subcommands. These weights are the released artifacts that let you run predict without the training data.

Usage

pip install -r requirements.txt   # from the code repo; TabPFN MUST be 2.2.1 (see below)

# download the weights into ./weights
python -c "from huggingface_hub import snapshot_download; \
snapshot_download('philmorekoung/liver-steatosis', local_dir='weights')"

# tabular, no vision  β€” one row per donor, columns matching the feature list
python no_vision_model.py predict --csv your_donors.csv --out predictions.csv

# vision  β€” npz with uuids[str] and images[N,96,96,96] scaled to [0,1]
python vision_model.py predict --weights weights/vision_suprem.pt \
    --images your_volumes.npz --out predictions.csv

# full  β€” features + volumes joined on uuid
python full_model.py predict --weights weights/full.pkl \
    --csv your_donors.csv --images your_volumes.npz --out predictions.csv

Output columns: uuid, p_macrosteatosis (probability of fat β‰₯ 30%), and high_risk (1 if p β‰₯ threshold). The operating threshold is chosen on a held-out split at training time and stored inside each weight file.

Required tabular features

The tabular models expect 218 numeric features per donor: registry clinical/donor-history variables, body-composition & organ-HU statistics, the per-donor liver HU histogram, and 150 topological persistence features f0..f149 (f0–49 = H0, f50–99 = H1, f100–149 = H2). Missing values are allowed. The exact column order is stored in the pickle (feature_cols); predict reorders your columns by name and errors out naming any that are missing. Producing the topology and HU-histogram features for a new CT requires the upstream liver-segmentation + persistence pipeline used to build the cohort.

Training data

A national, multi-site cohort of 2,709 biopsy-verified donor liver CTs (prevalence of fat β‰₯ 30% = 11.2%, 303/2,709). Labels are biopsy macrosteatosis percentage, binned to 5 classes (0 / 1–9 / 10–29 / 30–49 / β‰₯50 %) for training and collapsed to the β‰₯30% binary at prediction. The CT, segmentation, and registry data are access-restricted and are not distributed with these weights.

Performance

Benchmark ROC AUC for the β‰₯30% task (10 seeds, held-out test split, from the paper):

Model ROC AUC PR AUC
Clinical only 0.644 0.19
SuPreM image only 0.778 0.43
Tabular (clin + HU + topology) β€” no-vision 0.816 0.48
Full (tabular + SuPreM image) 0.815 0.48

Discrimination plateaus once tabular modalities are combined; adding the image branch does not significantly improve over the tabular ensemble at this cohort size (prevalence 11.2%, so PR AUC is the more informative operating-point metric).

These released checkpoints are fit for deployment rather than the multi-seed benchmark, so their self-reported metrics differ slightly:

  • no_vision.pkl β€” 5-fold out-of-fold ROC AUC 0.797, PR AUC 0.435; base learners then refit on all 2,709 donors.
  • vision_suprem.pt β€” validation ROC AUC 0.821.

Requirements & caveats

  • TabPFN must be 2.2.1. The .pkl files embed a fitted TabPFN object (a tabular foundation model that carries its training table β€” this is why the pickles are large). Newer TabPFN releases change the model and add an interactive license gate; loading the pickle with a different version may fail. Pin tabpfn==2.2.1.
  • Keep full.pkl and full_vision.pt together. full.pkl references its image branch by relative filename (full_vision.pt) and full_model.py predict looks for it in the same folder.
  • Library versions for unpickling: scikit-learn==1.6.1, lightgbm==4.6.0, xgboost==3.2.0 (see the code repo's requirements.txt).
  • The vision weights are a fine-tune of SuPreM's SwinUNETR; prediction needs only this checkpoint, but your use is subject to SuPreM's license (https://github.com/MrGiovanni/SuPreM) in addition to this card's license.

License

Released under CC BY-NC 4.0 (non-commercial).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support