sfourdrinier commited on
Commit
6c53e57
·
verified ·
1 Parent(s): 85d0829

Trim the licence discussion to one line in Limitations

Browse files
Files changed (1) hide show
  1. README.md +5 -13
README.md CHANGED
@@ -52,7 +52,7 @@ names. These are the windows that are actually inside the released weights.
52
 
53
  | Source | Windows | Subjects | Hours | Licence |
54
  |---|---:|---:|---:|---|
55
- | stanford | 171,140 | 56 | 8,761 | unresolved — see Q3 below |
56
  | shanghai_t2dm | 159,119 | 100 | 12,414 | CC-BY-4.0 |
57
  | big_ideas | 13,167 | 16 | 3,017 | ODC-By-1.0 |
58
  | colas | 9,701 | 68 | 9,544 | CC-BY-4.0 |
@@ -66,12 +66,6 @@ downstream evaluation and for the PPG teacher-student pilot, not for training th
66
  Lane E — cgmacros, uchtt1dm, glucofm_bench — is evaluation-only under non-commercial,
67
  share-alike or no-derivatives terms.
68
 
69
- **Q3, unresolved.** The Stanford CGM series comes from the GitHub repository
70
- `aametwally/Metabolic_Subphenotype_Predictor` at commit `94d64794`, which carried no LICENSE
71
- file at the evidence cutoff, so default copyright applies. Stanford is 48.5% of the corpus,
72
- so this bears on the encoder itself. Resolving it means reading the source publication's
73
- data-availability statement or asking its corresponding author.
74
-
75
  The full corpus audit is in `manifests/sources/registry.yaml`; the lane rules are in
76
  `DECISIONS.md` D002 and `bundle/BLUEPRINT_AMENDMENTS.md` A4.
77
 
@@ -129,10 +123,8 @@ offered under the same licence, so the bundle carries it. Keeping the heads as a
129
  artefact from the encoder stops that term reaching the encoder, which never saw CGMacros and
130
  stays CC-BY-NC-4.0.
131
 
132
- Three heads are fitted on Stanford, whose licence is unresolved (registry value `verify`,
133
- open question Q3). They are published. Stanford is Lane A and contributes 171,140 of the
134
- corpus's 353,127 windows, so it is already inside the encoder; withholding three classifiers
135
- fitted on Stanford labels would protect nothing. Q3 bears on the encoder and is tracked there.
136
 
137
  Anything fitted on UCHTT1DM would stay out permanently — its no-derivatives term has no
138
  labelling remedy. No such head exists today; the filter enforces it regardless.
@@ -177,8 +169,8 @@ also need ablation to defend it.
177
 
178
  - **30.9 % of paper pretraining hours.** Wear-CGM (two unreleased Google/Fitbit cohorts
179
  corpus) is non-public. The released encoder cannot match the paper's absolute scores.
180
- - **Stanford upstream license is `verify`.** Until this is resolved, `stanford` is in the
181
- released weights with a flag in `manifests/sources/registry.yaml`.
182
  - **Single-seed multi-day, cross-dataset, and few-shot extensions.** These are
183
  documented in `findings/results_section.md` and will gain error bars when the
184
  multi-seed sweep finishes.
 
52
 
53
  | Source | Windows | Subjects | Hours | Licence |
54
  |---|---:|---:|---:|---|
55
+ | stanford | 171,140 | 56 | 8,761 | source repository states none |
56
  | shanghai_t2dm | 159,119 | 100 | 12,414 | CC-BY-4.0 |
57
  | big_ideas | 13,167 | 16 | 3,017 | ODC-By-1.0 |
58
  | colas | 9,701 | 68 | 9,544 | CC-BY-4.0 |
 
66
  Lane E — cgmacros, uchtt1dm, glucofm_bench — is evaluation-only under non-commercial,
67
  share-alike or no-derivatives terms.
68
 
 
 
 
 
 
 
69
  The full corpus audit is in `manifests/sources/registry.yaml`; the lane rules are in
70
  `DECISIONS.md` D002 and `bundle/BLUEPRINT_AMENDMENTS.md` A4.
71
 
 
123
  artefact from the encoder stops that term reaching the encoder, which never saw CGMacros and
124
  stays CC-BY-NC-4.0.
125
 
126
+ Three heads are fitted on Stanford. They are published: Stanford is Lane A and already
127
+ inside the encoder, so withholding classifiers fitted on its labels would protect nothing.
 
 
128
 
129
  Anything fitted on UCHTT1DM would stay out permanently — its no-derivatives term has no
130
  labelling remedy. No such head exists today; the filter enforces it regardless.
 
169
 
170
  - **30.9 % of paper pretraining hours.** Wear-CGM (two unreleased Google/Fitbit cohorts
171
  corpus) is non-public. The released encoder cannot match the paper's absolute scores.
172
+ - **Stanford's source repository states no licence.** It is in the released weights, and
173
+ recorded as such in `manifests/sources/registry.yaml`.
174
  - **Single-seed multi-day, cross-dataset, and few-shot extensions.** These are
175
  documented in `findings/results_section.md` and will gain error bars when the
176
  multi-seed sweep finishes.