--- license: openrail++ language: - en tags: - histopathology - he-staining - breast-cancer - diffusion - counterfactuals - explainability - film - conditional-image-generation base_model: Manojb/stable-diffusion-2-1-base base_model_relation: finetune --- # CPathOGen **Spatially and Morphologically Controlled H&E Counterfactuals for Probing Pathology Models** - **Code and complete instructions:** [a12dongithub/PathOGen](https://github.com/a12dongithub/PathOGen) - **Model repository:** [a12donhf/CPathOGen](https://huggingface.co/a12donhf/CPathOGen) - **Authors:** Samarth Singhal and Varang Rai ## What the model does CPathOGen synthesizes a 512 x 512 H&E tile from a cellular spatial map and a morphology/appearance vector. A learned spatial encoder creates latent-space features that are concatenated with the noisy image latent. Blockwise FiLM modules apply the morphology controls during denoising. Researchers can keep diffusion noise fixed, change a requested control, and compare outputs of a frozen pathology model on the resulting matched images. This release contains the generator; CellViT++ candidate ranking and independent nucleus analysis require their own software and checkpoints. The checkpoint uses latent concatenation with a learned spatial encoder. It is not a standard Diffusers `ControlNetModel` or standalone `DiffusionPipeline` checkpoint. ## Figures from the paper ![Controllable image generation and recorded black-box predictions for post-hoc interpretation](assets/principle.png) **Counterfactual probing:** change a control, generate a matched image, and measure the downstream model's prediction response. ![CPathOGen pipeline from CellViT++ weak labels to spatial and morphology-conditioned latent diffusion](assets/pipeline.png) **Generation pipeline:** CellViT++ supplies cellular maps and morphology/appearance summaries; the spatial encoder and FiLM condition latent diffusion synthesis. ![Three examples showing the cell map, real H&E tile, and generated H&E tile](assets/spatial_fidelity.png) **Spatial examples:** input maps, associated real tiles, and condition-matched generated tiles from the paper. Map colors are tumor (white), immune (cyan), stroma (green), dead (yellow), and non-neoplastic epithelium (orange). ![Five-level sweeps of nuclear size, eccentricity, solidity, gradient, and RGB appearance](assets/morphology_controls.png) **Morphology and appearance examples:** the paper's five-level control sweeps. Most columns use relative standardized offsets; the historical eccentricity illustration instead uses absolute standardized coordinates with 20 steps and spatial strength 1. Nuclear size changes area and perimeter together. These are paper illustrations, not new generations from the quickstart. ## Run one real example in Colab Select a **GPU runtime** in Colab, then run the cell below. It downloads one prepared paper-tile condition, so you do not need the full dataset, Google Drive, or CellViT++ to try the generator. The public weights require no Hugging Face token. ```python from pathlib import Path import torch if not torch.cuda.is_available(): raise RuntimeError("Select Runtime > Change runtime type > GPU in Colab.") %cd /content if not Path("/content/CPathOGen-example/.git").is_dir(): !git clone -q --depth 1 https://github.com/a12dongithub/PathOGen.git /content/CPathOGen-example %cd /content/CPathOGen-example %pip install -q -r inference/requirements.txt from huggingface_hub import hf_hub_download from IPython.display import display from PIL import Image MODEL_ID = "a12donhf/CPathOGen" map_path = hf_hub_download(MODEL_ID, "examples/paper_tile/map.npz") morphology_path = hf_hub_download(MODEL_ID, "examples/paper_tile/morphology.json") preview_path = hf_hub_download(MODEL_ID, "examples/paper_tile/input_map.png") !python inference/generate.py --spatial-map "{map_path}" --morphology-json "{morphology_path}" --tile TCGA-E2-A15D_x46080_y24576_TR --seed 1872879198 --steps 30 --spatial-strength 2 --output outputs/paper_example.png print("Input spatial map") display(Image.open(preview_path)) print("Generated H&E") display(Image.open("outputs/paper_example.png")) print("Controls and run metadata: outputs/paper_example.json") ``` The first run downloads approximately **4.16 GB of generator weights**, plus the base-model text encoder, tokenizer, and scheduler. Downloads are cached. This produces one 512 x 512 PNG and its condition/provenance JSON; it does not run eight-seed selection or downstream probing. Keep the seed fixed when comparing edited controls. Exact pixels can differ across hardware and numerical precision. This model card documents GPU inference through the project code; it is not a hosted Hugging Face inference widget. ## Local inference and custom inputs Use Python 3.10/3.11, a compatible CUDA-enabled PyTorch build, and an NVIDIA GPU. ```bash git clone https://github.com/a12dongithub/PathOGen.git cd PathOGen python -m pip install -r inference/requirements.txt python inference/generate.py --synthetic-example --seed 42 --steps 30 --spatial-strength 2 --output outputs/demo.png ``` This downloads the model weights from this repository and the text encoder, tokenizer, and scheduler from the named Stable Diffusion 2.1 base model. Downloads are cached. The artificial demo layout is for software testing and does not reproduce the paper's evaluation. For a prepared dataset tile: ```bash python inference/generate.py --data-root /path/to/512_final_dataset --tile YOUR_TILE_ID --seed 42 --output outputs/tile.png ``` For custom conditions: ```bash python inference/generate.py --spatial-map conditions/map.npz --morphology-json conditions/morphology.json --seed 42 --output outputs/tile.png ``` See the [GitHub README](https://github.com/a12dongithub/PathOGen#custom-spatial-and-morphology-controls) for the complete schema, Colab cell, local-checkpoint usage, and paired interventions. ## Inputs The spatial NPZ must contain a `map` array with shape `(512, 512, 5)` or `(5, 512, 512)`. Channel order is neoplastic, inflammatory, connective, dead, and epithelial. Zero values in all channels represent background. Input intensities are uint8 in `[0, 255]` or floats in `[0, 1]`; the input is a map of smoothed cell centroids rather than a color-rendered illustration. The 16 morphology entries are standardized values in this order: ```text area_mean, area_var, eccentricity_mean, eccentricity_var, solidity_mean, solidity_var, perimeter_mean, perimeter_var, grad_mean, grad_var, r_mean, r_var, g_mean, g_var, b_mean, b_var ``` The historical preprocessing script did not save the fitted `StandardScaler`. Reuse already standardized dataset values; no verified training-scaler artifact is included in this release. A newly fitted scaler on another cohort can alter stain and geometry controls. ## Released files ```text checkpoint-30000/ unet/config.json unet/diffusion_pytorch_model.safetensors vae/config.json vae/diffusion_pytorch_model.safetensors film_mlps.pt spatial_encoder.pt assets/ principle.png pipeline.png spatial_fidelity.png morphology_controls.png examples/paper_tile/ map.npz morphology.json metadata.json input_map.png reference_generated.png ``` The FiLM and spatial-encoder files contain PyTorch state dictionaries and are loaded with `weights_only=True`. Optimizer, scheduler training state, and random-state pickle files are excluded. Original weights are preserved; the checkpoint is not quantized or converted. `release_manifest.json` records per-file SHA-256 hashes and sizes. ## Training and evaluation The paper describes training on approximately 1.4 million 512 x 512 tiles from TCGA-BRCA, retaining tiles with more than two detected tumor cells. CellViT++ supplies weak nucleus contours, locations, and labels. Training used eight Tesla V100 GPUs. The project reports the following distributional results: | Generation protocol | FID | KID mean | | --- | ---: | ---: | | No seed filtering | 42.9228 | 0.0246604 | | CellViT++ selection from eight seeds | 36.7168 | 0.0196440 | These values describe the project's evaluation protocol and its reference set. The software demo is not an FID/KID reproduction. Selection requires candidate generation and CellViT++ ranking; a single model call does not include selection. ## Intended use and limitations Research uses include controlled histopathology synthesis and probing model sensitivity to nuclear morphology, cellular organization, and appearance. The generator is not validated for clinical diagnosis or patient management. Shared diffusion noise and fixed conditions do not guarantee that every unmeasured property remains unchanged. Global morphology summaries compress cell-level heterogeneity, and out-of-distribution controls can produce artifacts. Synthetic intervention responses measure model behavior rather than biological or treatment causality. ## License Model weights retain the CreativeML Open RAIL++-M terms inherited from Stable Diffusion 2.1; see `LICENSE-MODEL`. Third-party software, analyzers, and datasets retain their respective licenses and access requirements. Paper figures are shared under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) with attribution to Samarth Singhal and Varang Rai; see [assets/README.md](assets/README.md). This does not change the model-weight license.