Public-pretrained CGM model with MOMENT temporal cross-attention

Selected seed-43, final-3,500-update research checkpoint. Development classification PR-AUC 60.709051%, AUROC 68.609122%, macro-F1 60.049084%. The three training seeds score 60.709051%, 58.829941%, and 59.975307%; mean 59.838100%, sample SD 0.947039 percentage points. This is not seed-stable >=60%, independent test confirmation, or evidence that cross-attention is necessary.

Model and training boundary

Raw 24-hour CGM and its observation mask feed a trainable CGM encoder. A frozen MOMENT prior supplies hourly tokens through a public-fitted, frozen PCA8 input transform and trainable cross-attention. The fused representation has 128 dimensions. A head jointly learned during public pretraining produces eight shape coordinates. The primary downstream input concatenates these with the actual observed mean of that individual window: 9 dimensions. The mean is an input statistic, not a cohort statistic fitted using test participants. The secondary input is 128D+mean.

The public CGM adaptation uses 1,094 windows, seed43, batch32 and exactly3,500 updates. Its loss is shape MSE + mean MSE +0.1 temporal JEPA; context JEPA is zero, standard-deviation supervision is zero, and EMA is retained. Teacher PCA targets and input token PCA are fitted on public pretraining rows before model training. The encoder, attention, head and temporal predictors genuinely train. Official MOMENT weights remain frozen. External TSFM pretraining is separate from the public-only CGM adaptation claim. Wear-CGM is not used.

For downstream classification, freeze all feature-path parameters and buffers. Fit only the logistic classifier and training-fold feature standardization. Never fit a new downstream PCA, ridge readout, encoder, attention or target head. Labels and downstream CGM must not enter representation pretraining or its PCA fit.

Contents

  • Original unchanged model.pt (SHA256 ebc40a0221e1e09bfa15d166f29abd3d5e89bbe1a8a05f81e732190e43e1b9ff).
  • Public aggregate token/teacher bases and teacher target normalizers.
  • Exact model dependency source, checksummed loader, numerical runtime and metadata.
  • No CGM records, participant embeddings/predictions, downstream classifier, or optimizer.

The loader downloads official AutonLab/MOMENT-1-large at the revision in bundle.json. For historical constructor parity, it also loads pinned Chronos-2 temporarily; the original constructor removes Chronos before any model computation. Chronos is not a prior, predictor or teacher in this model. Neither upstream backbone's weights are redistributed here. Their upstream terms continue to apply.

Load frozen features

hf download Yotto3108/cgm-moment-temporal-3500 --local-dir cgm-moment-temporal-3500
cd cgm-moment-temporal-3500
# Use an isolated environment and install the appropriate torch build first.
python -m pip install -r requirements.txt
python load_moment_cgm.py --device cpu

Record the returned Hub commit and use --revision <commit> for repeatable downloads. This is a PyTorch research bundle, not a Transformers AutoModel repository.

import torch
from load_moment_cgm import load_moment_cgm, extract_features
model, metadata = load_moment_cgm(device='cpu')
x = torch.full((2, 288), 110., dtype=torch.float32)  # synthetic loading example
observed = torch.ones_like(x, dtype=torch.bool)
starts = torch.zeros(2, dtype=torch.long)  # five-minute local-day slot, 0..287
features = extract_features(model, x, observed, starts)  # [2,9], frozen

Preserve raw mg/dL, missingness and physical time; do not compact missing samples or pre-normalize glucose. The synthetic example is not a performance test. Exact classification replay uses CPU native-feature batches of16, applies the head to the complete extracted feature matrix, and standardizes LP inputs in float64. See the executable evaluator and current Udit instructions in gluco-fm-bench.

Evaluation uses 14 dataset/task cells, five subject folds, ten repeats, splitseed42, L2 logistic regression C1. The result is the mean of cell-wise fold means. Repeated development selection and unverified global participant identity, including session-level Shanghai grouping, limit interpretation. A competitive aligned-fusion control scored60.632668%; conditional paired AP intervals did not establish learned-attention superiority. Do not infer clinical validation.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support