Publish verified OpenSysOne training snapshot pointer and model card
Browse files- CURRENT_SNAPSHOT.json +8 -7
- README.md +27 -8
CURRENT_SNAPSHOT.json
CHANGED
|
@@ -1,14 +1,15 @@
|
|
| 1 |
{
|
| 2 |
"repo_id": "andyshu/opensysone",
|
| 3 |
-
"snapshot_id": "
|
| 4 |
-
"path": "snapshots/
|
| 5 |
-
"payload_commit": "
|
| 6 |
-
"source_commit": "
|
| 7 |
"training_source_commits": [
|
|
|
|
| 8 |
"4a60423c39d70f8d50472ce4f4f7fa4a4bd9fce1"
|
| 9 |
],
|
| 10 |
-
"manifest_path": "publications/
|
| 11 |
-
"manifest_sha256": "
|
| 12 |
-
"published_utc": "2026-09-
|
| 13 |
"calibrated": false
|
| 14 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"repo_id": "andyshu/opensysone",
|
| 3 |
+
"snapshot_id": "20260917-expanded-pilot-hf-backup",
|
| 4 |
+
"path": "snapshots/20260917-expanded-pilot-hf-backup",
|
| 5 |
+
"payload_commit": "1a0a7fb679b88043c84325eadbbfc37c83345a31",
|
| 6 |
+
"source_commit": "fe7e48f806a7e947d689b9e966d38bc0dbedcc2e",
|
| 7 |
"training_source_commits": [
|
| 8 |
+
"24b8ccf60d388f9cbb184e03a6ae260a1f5a8b86",
|
| 9 |
"4a60423c39d70f8d50472ce4f4f7fa4a4bd9fce1"
|
| 10 |
],
|
| 11 |
+
"manifest_path": "publications/20260917-expanded-pilot-hf-backup/fe7e48f806a7e947d689b9e966d38bc0dbedcc2e/manifest.json",
|
| 12 |
+
"manifest_sha256": "9d6d8875cbd6c908db84a5cba2b7ac348df1c9d559c1f4836dedf198ec895fbc",
|
| 13 |
+
"published_utc": "2026-09-17T07:24:35.048421+00:00",
|
| 14 |
"calibrated": false
|
| 15 |
}
|
README.md
CHANGED
|
@@ -32,21 +32,25 @@ calibrated deployment have not completed. The original campaign ends on
|
|
| 32 |
`FINAL_MODEL.json` will identify the separately verified calibrated release when
|
| 33 |
finalization and upload succeed; its absence means no final release is recorded.
|
| 34 |
|
| 35 |
-
Validation audit on 17 September,
|
| 36 |
|
| 37 |
| Selected candidate | Step | Validation accuracy | Crossfit selection NLL |
|
| 38 |
| --- | ---: | ---: | ---: |
|
| 39 |
| Qwen3 4B, main run | 2,500 | 93.55% | 0.188640 |
|
| 40 |
-
| Qwen3 4B, lower learning rate |
|
| 41 |
| Qwen3.5 2B | 2,000 | 89.84% | 0.255294 |
|
|
|
|
| 42 |
|
| 43 |
All use the same 512 validation decisions. Selection uses four source-group-
|
| 44 |
disjoint temperature-crossfit folds with fixed seed 431; smaller macro-family NLL
|
| 45 |
is better. This is **validation-selection evidence**, not a held-out quality claim.
|
| 46 |
The 2B run stopped cleanly at step 6,000 after eight evaluations without a new
|
| 47 |
-
best.
|
| 48 |
-
|
| 49 |
-
|
|
|
|
|
|
|
|
|
|
| 50 |
|
| 51 |
## Contents and reconstruction
|
| 52 |
|
|
@@ -64,6 +68,8 @@ hosted inference application.
|
|
| 64 |
and reconstruction metadata. These initial backups have no fitted temperature.
|
| 65 |
- `snapshots/<id>/artifacts/<candidate>/checkpoint.pt`: resumable training state,
|
| 66 |
including Adam and random states. Its step may be later than the selected best.
|
|
|
|
|
|
|
| 67 |
- `final/<campaign>/`: final calibrated model and evaluation evidence, created only
|
| 68 |
after successful finalization and publication.
|
| 69 |
|
|
@@ -91,7 +97,7 @@ Pinned bases, recorded as Apache-2.0 in their provenance:
|
|
| 91 |
|
| 92 |
## Data, evaluation and limitations
|
| 93 |
|
| 94 |
-
|
| 95 |
Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences,
|
| 96 |
raw/split hashes and group/deduplication audits are included in
|
| 97 |
`source/results/public-decisions-v1-manifest.json`. That manifest credits
|
|
@@ -100,6 +106,17 @@ and records their original dataset licences. The original repository's
|
|
| 100 |
`license: unknown` metadata is retained; no new licence for the project artifacts
|
| 101 |
is assigned by this backup.
|
| 102 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 103 |
Validation selects checkpoints. Separate calibration fits one global temperature
|
| 104 |
only after selection is frozen. Final test/holdout comparisons use the unchanged
|
| 105 |
pretrained scorer, with a separately fitted base temperature and source-group
|
|
@@ -112,5 +129,7 @@ The harness implements local scoring, a loopback API, hosted Jev requests and
|
|
| 112 |
response/timing comparisons. The local model identifies itself as OpenSysOne;
|
| 113 |
compatibility with the request shape does not make it Jev. Local confidence is
|
| 114 |
normalized entropy, not calibrated probability of correctness. Authenticated
|
| 115 |
-
hosted Jev calls require `TYPESAFE_API_KEY` and have not been tested. Credentials
|
| 116 |
-
pretrained base weights
|
|
|
|
|
|
|
|
|
| 32 |
`FINAL_MODEL.json` will identify the separately verified calibrated release when
|
| 33 |
finalization and upload succeed; its absence means no final release is recorded.
|
| 34 |
|
| 35 |
+
Validation audit on 17 September, 07:00 UTC:
|
| 36 |
|
| 37 |
| Selected candidate | Step | Validation accuracy | Crossfit selection NLL |
|
| 38 |
| --- | ---: | ---: | ---: |
|
| 39 |
| Qwen3 4B, main run | 2,500 | 93.55% | 0.188640 |
|
| 40 |
+
| Qwen3 4B, lower learning rate | 4,000 | 93.75% | 0.190020 |
|
| 41 |
| Qwen3.5 2B | 2,000 | 89.84% | 0.255294 |
|
| 42 |
+
| Qwen3 4B, refinement | 1,500 | 94.73% | 0.170150 |
|
| 43 |
|
| 44 |
All use the same 512 validation decisions. Selection uses four source-group-
|
| 45 |
disjoint temperature-crossfit folds with fixed seed 431; smaller macro-family NLL
|
| 46 |
is better. This is **validation-selection evidence**, not a held-out quality claim.
|
| 47 |
The 2B run stopped cleanly at step 6,000 after eight evaluations without a new
|
| 48 |
+
best. At 07:07 UTC the original GX10 4B run stopped cleanly at step 4,380,
|
| 49 |
+
preserving its selected step 2,500 and full resumable state. GX10 now verifies an
|
| 50 |
+
expanded-data candidate from the refinement's frozen selected step 1,500, while
|
| 51 |
+
both Spark 4B runs continue. See `source/EXPANDED_DATA.md` for its latest run
|
| 52 |
+
state, verification evidence and the broader training mix. Independent trials
|
| 53 |
+
retain the original absolute deadlines.
|
| 54 |
|
| 55 |
## Contents and reconstruction
|
| 56 |
|
|
|
|
| 68 |
and reconstruction metadata. These initial backups have no fitted temperature.
|
| 69 |
- `snapshots/<id>/artifacts/<candidate>/checkpoint.pt`: resumable training state,
|
| 70 |
including Adam and random states. Its step may be later than the selected best.
|
| 71 |
+
- Expanded-data snapshots include transformed decision files and their source
|
| 72 |
+
notices, split hashes, diagnostic exclusions and original dataset lineage.
|
| 73 |
- `final/<campaign>/`: final calibrated model and evaluation evidence, created only
|
| 74 |
after successful finalization and publication.
|
| 75 |
|
|
|
|
| 97 |
|
| 98 |
## Data, evaluation and limitations
|
| 99 |
|
| 100 |
+
The original public training sources are SNLI, BoolQ, ARC and four-choice Banking77 routing.
|
| 101 |
Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences,
|
| 102 |
raw/split hashes and group/deduplication audits are included in
|
| 103 |
`source/results/public-decisions-v1-manifest.json`. That manifest credits
|
|
|
|
| 106 |
`license: unknown` metadata is retained; no new licence for the project artifacts
|
| 107 |
is assigned by this backup.
|
| 108 |
|
| 109 |
+
The expanded candidate adds official training rows from HellaSwag (MIT), PIQA
|
| 110 |
+
(AFL-3.0 according to its creator's pinned README) and CommonsenseQA (MIT).
|
| 111 |
+
After the 512-token filter there are 80,765 training decisions, including all
|
| 112 |
+
40,915 original retained examples. All four original reserved split files and
|
| 113 |
+
their tokenized rows are unchanged. Another 383 retained new-source diagnostic
|
| 114 |
+
decisions are excluded from both training and checkpoint selection. Source pins,
|
| 115 |
+
credits, licences and transformations are in
|
| 116 |
+
`source/results/20260917-expanded-data/dataset-manifest.json`. The unchanged
|
| 117 |
+
selection set measures the original tasks; expansion alone is not evidence of
|
| 118 |
+
better performance on the three added tasks.
|
| 119 |
+
|
| 120 |
Validation selects checkpoints. Separate calibration fits one global temperature
|
| 121 |
only after selection is frozen. Final test/holdout comparisons use the unchanged
|
| 122 |
pretrained scorer, with a separately fitted base temperature and source-group
|
|
|
|
| 129 |
response/timing comparisons. The local model identifies itself as OpenSysOne;
|
| 130 |
compatibility with the request shape does not make it Jev. Local confidence is
|
| 131 |
normalized entropy, not calibrated probability of correctness. Authenticated
|
| 132 |
+
hosted Jev calls require `TYPESAFE_API_KEY` and have not been tested. Credentials
|
| 133 |
+
and pretrained base weights are not part of this backup. Expanded-data snapshots
|
| 134 |
+
may include the transformed training data, with upstream notices retained;
|
| 135 |
+
raw upstream archives are referenced by pinned revision and checksum.
|