andyshu commited on
Commit
8f7cfb6
·
verified ·
1 Parent(s): 1a0a7fb

Publish verified OpenSysOne training snapshot pointer and model card

Browse files
Files changed (2) hide show
  1. CURRENT_SNAPSHOT.json +8 -7
  2. README.md +27 -8
CURRENT_SNAPSHOT.json CHANGED
@@ -1,14 +1,15 @@
1
  {
2
  "repo_id": "andyshu/opensysone",
3
- "snapshot_id": "20260917T022835Z-hf-backup",
4
- "path": "snapshots/20260917T022835Z-hf-backup",
5
- "payload_commit": "e60eab96fdd568c5ce7c9e880640874bee63e88b",
6
- "source_commit": "b156513bf6ac1475eaa24fd79e7beb94d159364f",
7
  "training_source_commits": [
 
8
  "4a60423c39d70f8d50472ce4f4f7fa4a4bd9fce1"
9
  ],
10
- "manifest_path": "publications/20260917T022835Z-hf-backup/b156513bf6ac1475eaa24fd79e7beb94d159364f/manifest.json",
11
- "manifest_sha256": "5fc9dab067477b1d2e38c362bbdc3f45aa435cdcb800d3710d148f85cc217ce6",
12
- "published_utc": "2026-09-17T03:59:39.663515+00:00",
13
  "calibrated": false
14
  }
 
1
  {
2
  "repo_id": "andyshu/opensysone",
3
+ "snapshot_id": "20260917-expanded-pilot-hf-backup",
4
+ "path": "snapshots/20260917-expanded-pilot-hf-backup",
5
+ "payload_commit": "1a0a7fb679b88043c84325eadbbfc37c83345a31",
6
+ "source_commit": "fe7e48f806a7e947d689b9e966d38bc0dbedcc2e",
7
  "training_source_commits": [
8
+ "24b8ccf60d388f9cbb184e03a6ae260a1f5a8b86",
9
  "4a60423c39d70f8d50472ce4f4f7fa4a4bd9fce1"
10
  ],
11
+ "manifest_path": "publications/20260917-expanded-pilot-hf-backup/fe7e48f806a7e947d689b9e966d38bc0dbedcc2e/manifest.json",
12
+ "manifest_sha256": "9d6d8875cbd6c908db84a5cba2b7ac348df1c9d559c1f4836dedf198ec895fbc",
13
+ "published_utc": "2026-09-17T07:24:35.048421+00:00",
14
  "calibrated": false
15
  }
README.md CHANGED
@@ -32,21 +32,25 @@ calibrated deployment have not completed. The original campaign ends on
32
  `FINAL_MODEL.json` will identify the separately verified calibrated release when
33
  finalization and upload succeed; its absence means no final release is recorded.
34
 
35
- Validation audit on 17 September, 02:10–02:15 UTC:
36
 
37
  | Selected candidate | Step | Validation accuracy | Crossfit selection NLL |
38
  | --- | ---: | ---: | ---: |
39
  | Qwen3 4B, main run | 2,500 | 93.55% | 0.188640 |
40
- | Qwen3 4B, lower learning rate | 2,500 | 92.58% | 0.218012 |
41
  | Qwen3.5 2B | 2,000 | 89.84% | 0.255294 |
 
42
 
43
  All use the same 512 validation decisions. Selection uses four source-group-
44
  disjoint temperature-crossfit folds with fixed seed 431; smaller macro-family NLL
45
  is better. This is **validation-selection evidence**, not a held-out quality claim.
46
  The 2B run stopped cleanly at step 6,000 after eight evaluations without a new
47
- best. Both 4B runs continued, and another lower-rate 4B refinement was started
48
- from the main run's selected weights. See `source/NEXT_STEPS.md` for findings and
49
- why independent trials were retained over unmeasured distributed training.
 
 
 
50
 
51
  ## Contents and reconstruction
52
 
@@ -64,6 +68,8 @@ hosted inference application.
64
  and reconstruction metadata. These initial backups have no fitted temperature.
65
  - `snapshots/<id>/artifacts/<candidate>/checkpoint.pt`: resumable training state,
66
  including Adam and random states. Its step may be later than the selected best.
 
 
67
  - `final/<campaign>/`: final calibrated model and evaluation evidence, created only
68
  after successful finalization and publication.
69
 
@@ -91,7 +97,7 @@ Pinned bases, recorded as Apache-2.0 in their provenance:
91
 
92
  ## Data, evaluation and limitations
93
 
94
- Public training sources are SNLI, BoolQ, ARC and four-choice Banking77 routing.
95
  Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences,
96
  raw/split hashes and group/deduplication audits are included in
97
  `source/results/public-decisions-v1-manifest.json`. That manifest credits
@@ -100,6 +106,17 @@ and records their original dataset licences. The original repository's
100
  `license: unknown` metadata is retained; no new licence for the project artifacts
101
  is assigned by this backup.
102
 
 
 
 
 
 
 
 
 
 
 
 
103
  Validation selects checkpoints. Separate calibration fits one global temperature
104
  only after selection is frozen. Final test/holdout comparisons use the unchanged
105
  pretrained scorer, with a separately fitted base temperature and source-group
@@ -112,5 +129,7 @@ The harness implements local scoring, a loopback API, hosted Jev requests and
112
  response/timing comparisons. The local model identifies itself as OpenSysOne;
113
  compatibility with the request shape does not make it Jev. Local confidence is
114
  normalized entropy, not calibrated probability of correctness. Authenticated
115
- hosted Jev calls require `TYPESAFE_API_KEY` and have not been tested. Credentials,
116
- pretrained base weights and raw training datasets are not part of this backup.
 
 
 
32
  `FINAL_MODEL.json` will identify the separately verified calibrated release when
33
  finalization and upload succeed; its absence means no final release is recorded.
34
 
35
+ Validation audit on 17 September, 07:00 UTC:
36
 
37
  | Selected candidate | Step | Validation accuracy | Crossfit selection NLL |
38
  | --- | ---: | ---: | ---: |
39
  | Qwen3 4B, main run | 2,500 | 93.55% | 0.188640 |
40
+ | Qwen3 4B, lower learning rate | 4,000 | 93.75% | 0.190020 |
41
  | Qwen3.5 2B | 2,000 | 89.84% | 0.255294 |
42
+ | Qwen3 4B, refinement | 1,500 | 94.73% | 0.170150 |
43
 
44
  All use the same 512 validation decisions. Selection uses four source-group-
45
  disjoint temperature-crossfit folds with fixed seed 431; smaller macro-family NLL
46
  is better. This is **validation-selection evidence**, not a held-out quality claim.
47
  The 2B run stopped cleanly at step 6,000 after eight evaluations without a new
48
+ best. At 07:07 UTC the original GX10 4B run stopped cleanly at step 4,380,
49
+ preserving its selected step 2,500 and full resumable state. GX10 now verifies an
50
+ expanded-data candidate from the refinement's frozen selected step 1,500, while
51
+ both Spark 4B runs continue. See `source/EXPANDED_DATA.md` for its latest run
52
+ state, verification evidence and the broader training mix. Independent trials
53
+ retain the original absolute deadlines.
54
 
55
  ## Contents and reconstruction
56
 
 
68
  and reconstruction metadata. These initial backups have no fitted temperature.
69
  - `snapshots/<id>/artifacts/<candidate>/checkpoint.pt`: resumable training state,
70
  including Adam and random states. Its step may be later than the selected best.
71
+ - Expanded-data snapshots include transformed decision files and their source
72
+ notices, split hashes, diagnostic exclusions and original dataset lineage.
73
  - `final/<campaign>/`: final calibrated model and evaluation evidence, created only
74
  after successful finalization and publication.
75
 
 
97
 
98
  ## Data, evaluation and limitations
99
 
100
+ The original public training sources are SNLI, BoolQ, ARC and four-choice Banking77 routing.
101
  Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences,
102
  raw/split hashes and group/deduplication audits are included in
103
  `source/results/public-decisions-v1-manifest.json`. That manifest credits
 
106
  `license: unknown` metadata is retained; no new licence for the project artifacts
107
  is assigned by this backup.
108
 
109
+ The expanded candidate adds official training rows from HellaSwag (MIT), PIQA
110
+ (AFL-3.0 according to its creator's pinned README) and CommonsenseQA (MIT).
111
+ After the 512-token filter there are 80,765 training decisions, including all
112
+ 40,915 original retained examples. All four original reserved split files and
113
+ their tokenized rows are unchanged. Another 383 retained new-source diagnostic
114
+ decisions are excluded from both training and checkpoint selection. Source pins,
115
+ credits, licences and transformations are in
116
+ `source/results/20260917-expanded-data/dataset-manifest.json`. The unchanged
117
+ selection set measures the original tasks; expansion alone is not evidence of
118
+ better performance on the three added tasks.
119
+
120
  Validation selects checkpoints. Separate calibration fits one global temperature
121
  only after selection is frozen. Final test/holdout comparisons use the unchanged
122
  pretrained scorer, with a separately fitted base temperature and source-group
 
129
  response/timing comparisons. The local model identifies itself as OpenSysOne;
130
  compatibility with the request shape does not make it Jev. Local confidence is
131
  normalized entropy, not calibrated probability of correctness. Authenticated
132
+ hosted Jev calls require `TYPESAFE_API_KEY` and have not been tested. Credentials
133
+ and pretrained base weights are not part of this backup. Expanded-data snapshots
134
+ may include the transformed training data, with upstream notices retained;
135
+ raw upstream archives are referenced by pinned revision and checksum.