andyshu commited on
Commit
75a4b6a
·
verified ·
1 Parent(s): 75c3d90

Publish verified OpenSysOne directory index and manifest pointer

Browse files
Files changed (2) hide show
  1. PUBLICATION.json +10 -0
  2. README.md +69 -139
PUBLICATION.json ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "accepted_lfs_attributes": null,
3
+ "format": "opensysone-publication-pointer-v1",
4
+ "manifest_path": "publication-manifest.json",
5
+ "manifest_sha256": "a92002e3605de0e9cdafe8691fca44e482f0465da26c189958838c4df01250b6",
6
+ "payload_commit": "75c3d90272aa383c2633df564a78b56370855b54",
7
+ "repo_id": "andyshu/opensysone",
8
+ "source_commit": "363ea4987da8e4a80ae8595e52f4f9db0866d4e7",
9
+ "verified_utc": "2026-09-17T10:04:57.045039+00:00"
10
+ }
README.md CHANGED
@@ -13,142 +13,72 @@ tags:
13
 
14
  # OpenSysOne
15
 
16
- **Credit:** OpenSysOne is inspired by [Jev](https://typesafe.ai/), TypeSafe.ai's
17
- System One model for structured decisions with probabilities. Credit goes to the
18
- TypeSafe team for motivating this project. OpenSysOne is an independent
19
- experimental implementation.
20
-
21
- OpenSysOne is an experimental natural-language decision scorer. It scores an
22
- arbitrary state, question and candidate answer, then normalizes mutually exclusive
23
- choices into probabilities. It does not generate chat responses.
24
-
25
- This repository backs up the ongoing three-machine training experiment. It
26
- contains custom adapter/head checkpoints, reproducible source, run provenance
27
- and small result files. The full pretrained base weights are referenced by pinned
28
- revision and are not duplicated here. **These are custom OpenSysOne artifacts,
29
- not a drop-in Transformers or PEFT adapter repository.**
30
-
31
- ## Current status
32
-
33
- The snapshot is **training-stage, uncalibrated**. Final held-out evaluation and
34
- calibrated deployment have not completed. The original campaign ends on
35
- **17 September 2026 at 18:16:10 UTC**, with training stopping by 16:00 UTC.
36
- `CURRENT_SNAPSHOT.json` identifies the backup and its source revision.
37
- `FINAL_MODEL.json` will identify the separately verified calibrated release when
38
- finalization and upload succeed; its absence means no final release is recorded.
39
-
40
- Validation audit on 17 September, 07:00 UTC:
41
-
42
- | Selected candidate | Step | Validation accuracy | Crossfit selection NLL |
43
- | --- | ---: | ---: | ---: |
44
- | Qwen3 4B, main run | 2,500 | 93.55% | 0.188640 |
45
- | Qwen3 4B, lower learning rate | 4,000 | 93.75% | 0.190020 |
46
- | Qwen3.5 2B | 2,000 | 89.84% | 0.255294 |
47
- | Qwen3 4B, refinement | 1,500 | 94.73% | 0.170150 |
48
-
49
- All use the same 512 validation decisions. Selection uses four source-group-
50
- disjoint temperature-crossfit folds with fixed seed 431; smaller macro-family NLL
51
- is better. This is **validation-selection evidence**, not a held-out quality claim.
52
- The 2B run stopped cleanly at step 6,000 after eight evaluations without a new
53
- best. At 07:07 UTC the original GX10 4B run stopped cleanly at step 4,380,
54
- preserving its selected step 2,500 and full resumable state. GX10 now verifies an
55
- expanded-data candidate from the refinement's frozen selected step 1,500, while
56
- both Spark 4B runs continue. See `source/EXPANDED_DATA.md` for its latest run
57
- state, verification evidence and the broader training mix. Independent trials
58
- retain the original absolute deadlines.
59
-
60
- ## Contents and reconstruction
61
-
62
- The [browser playground](source/PLAYGROUND.md) provides a context box, question,
63
- options list and probability bars, with a model selector. Its lightweight local
64
- server uses these custom checkpoints and the pinned base weights. It runs on
65
- GX10 over a loopback connection/SSH tunnel; this repository is a backup, not a
66
- hosted inference application.
67
-
68
- - `source/`: a committed source/documentation snapshot, including the training
69
- scripts, tests, Jev-compatible harness, findings and operational handover.
70
- - `snapshots/<id>/backup-manifest.json`: artifact roles, originating hosts/paths,
71
- checkpoint steps, model/data/source provenance and SHA-256 checksums.
72
- - `snapshots/<id>/artifacts/<candidate>/best.pt`: selected adapter/head weights
73
- and reconstruction metadata. These initial backups have no fitted temperature.
74
- - `snapshots/<id>/artifacts/<candidate>/checkpoint.pt`: resumable training state,
75
- including Adam and random states. Its step may be later than the selected best.
76
- - Expanded-data snapshots include transformed decision files and their source
77
- notices, split hashes, diagnostic exclusions and original dataset lineage.
78
- - `final/<campaign>/`: final calibrated model and evaluation evidence, created only
79
- after successful finalization and publication.
80
-
81
- Keep `checkpoint.pt`, its sibling `best.pt` and matching validation evidence
82
- together when resuming. Checkpoint metadata records the original local paths;
83
- reconstruction needs the pinned base/data and matching project source. The current
84
- harness works on the recorded GX10/Spark layout. This backup does not establish
85
- portability to a new operating system or a generic hosted inference service.
86
- Model and optimizer files use PyTorch serialization; use the project's known,
87
- hash-verified artifacts and its supplied reconstruction code.
88
-
89
- The main 4B model uses FP32, rank-8 additive adapters (alpha 16) and a learned
90
- scalar head initialized from pretrained yes-minus-no logits: 16,517,633 trainable
91
- parameters. The 2B text decoder uses rank 16 / alpha 32, with 16,821,249 trainable
92
- parameters. Training limits are 512 and 768 complete-chat tokens respectively;
93
- 1,024-token inference was separately verified. Training retains a 16 GiB per-job
94
- CUDA allocation cap. BF16 did not pass the project's numerical invariance gate.
95
-
96
- Pinned bases, recorded as Apache-2.0 in their provenance:
97
-
98
- - `Qwen/Qwen3-4B-Instruct-2507` at
99
- `cdbee75f17c01a7cc42f958dc650907174af0554`.
100
- - `Qwen/Qwen3.5-2B` at
101
- `15852e8c16360a2fea060d615a32b45270f8a8fc`.
102
-
103
- ## Data, evaluation and limitations
104
-
105
- The original public training sources are SNLI, BoolQ, ARC and four-choice Banking77 routing.
106
- Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences,
107
- raw/split hashes and group/deduplication audits are included in
108
- `source/results/public-decisions-v1-manifest.json`. That manifest credits
109
- Stanford NLP (SNLI), Google (BoolQ), AllenAI (ARC/Social IQA) and PolyAI (Banking77),
110
- and records their original dataset licences. The original repository's
111
- `license: unknown` metadata is retained; no new licence for the project artifacts
112
- is assigned by this backup.
113
-
114
- The expanded candidate adds official training rows from HellaSwag (MIT), PIQA
115
- (AFL-3.0 according to its creator's pinned README) and CommonsenseQA (MIT).
116
- After the 512-token filter there are 80,765 training decisions, including all
117
- 40,915 original retained examples. All four original reserved split files and
118
- their tokenized rows are unchanged. Another 383 retained new-source diagnostic
119
- decisions are excluded from both training and checkpoint selection. Source pins,
120
- credits, licences and transformations are in
121
- `source/results/20260917-expanded-data/dataset-manifest.json`. The unchanged
122
- selection set measures the original tasks; expansion alone is not evidence of
123
- better performance on the three added tasks.
124
-
125
- Validation selects checkpoints. Separate calibration fits one global temperature
126
- only after selection is frozen. Final test/holdout comparisons use the unchanged
127
- pretrained scorer, with a separately fitted base temperature and source-group
128
- uncertainty. Frozen grouping and exact deduplication do not rule out pretraining
129
- contamination or semantic duplicates. Four-choice Banking77 is not the full
130
- 77-label benchmark. No claim of general intelligence, Jev equivalence or
131
- calibration on unseen task families follows from the current validation scores.
132
-
133
- The harness implements local scoring, a loopback API, hosted Jev requests and
134
- response/timing comparisons. The local model identifies itself as OpenSysOne;
135
- compatibility with the request shape does not make it Jev. Local confidence is
136
- normalized entropy, not calibrated probability of correctness. Authenticated
137
- hosted Jev calls require `TYPESAFE_API_KEY` and have not been tested. Credentials
138
- and pretrained base weights are not part of this backup. Expanded-data snapshots
139
- may include the transformed training data, with upstream notices retained;
140
- raw upstream archives are referenced by pinned revision and checksum.
141
-
142
- <!-- opensysone-final:start -->
143
-
144
- ## Completed fleet deployment
145
-
146
- **This completed release supersedes the earlier training-stage Current status.** Descriptions of unfinished finalization and the training snapshot above are historical. `FINAL_MODEL.json` now identifies the verified calibrated release.
147
-
148
- Final artifact: [`final/20260916T194403396250Z-fleet/model.pt`](final/20260916T194403396250Z-fleet/model.pt). Integrity and provenance: [`backup_manifest.json`](final/20260916T194403396250Z-fleet/backup_manifest.json). Completed test/holdout results: [`metrics.json`](final/20260916T194403396250Z-fleet/evaluation/metrics.json).
149
-
150
- This contains trained adapter/head parameters and calibration metadata; pinned base weights are required separately. It scores explicit candidate decisions through the Jev-compatible harness. Public benchmark metrics do not establish general intelligence or universal calibration.
151
-
152
- Source revision: `6729461ccaad32e239c2148fa0c8f9ca23513a7a`. No license metadata has been changed.
153
-
154
- <!-- opensysone-final:end -->
 
13
 
14
  # OpenSysOne
15
 
16
+ **Inspired by [Jev](https://typesafe.ai/), TypeSafe.ai's System One model.**
17
+ Credit goes to the TypeSafe team for inspiring this project's exploration of
18
+ structured decisions with probabilities. OpenSysOne is an independent experimental implementation; API
19
+ compatibility does not establish Jev equivalence.
20
+
21
+ The **completed 4B release** scores a state, question and explicit candidate
22
+ answers, returning probabilities over those choices. Training, separate
23
+ calibration, final evaluation and local API verification completed on
24
+ 17 September 2026.
25
+
26
+ Start with the [model and reconstruction notes](model/README.md),
27
+ [results report](results/report.md), or [publication guide](docs/README.md).
28
+ The calibrated artifact is [model/model.pt](model/model.pt).
29
+ It contains custom OpenSysOne adapter/head weights and metadata. The pinned
30
+ Qwen3-4B-Instruct-2507 base is required separately; this is not a standalone
31
+ Transformers model or a standard PEFT adapter package.
32
+
33
+ ## Measured results
34
+
35
+ The full comparison uses the unchanged pretrained yes/no verifier, with a separate
36
+ temperature fitted for each model. Intervals are paired 95% source-group bootstrap
37
+ intervals for selected minus base accuracy.
38
+
39
+ | Evaluation | Decisions | Selected | Base verifier | Accuracy gain (95% interval) |
40
+ | --- | ---: | ---: | ---: | ---: |
41
+ | Known-family test | 2,042 | 92.90% | 84.48% | +8.42 pp [6.85, 9.89] |
42
+ | Social IQA family holdout | 768 | 72.92% | 70.31% | +2.60 pp [0.13, 5.34] |
43
+
44
+ On a separate matched 320-decision profile, selected accuracy was 89.06%, versus
45
+ 80.94% for the base verifier and 86.25% for a base model using one constrained
46
+ answer-label token. The selected scorer was **slower on all 12 profiled workloads**:
47
+ 1.11–1.17× the verifier latency and 2.18–15.58× the label baseline latency.
48
+ These are warm, serial FP32 measurements on one GB10, not concurrent-serving
49
+ throughput or comparisons with generated reasoning. See [tables and charts](results/).
50
+
51
+ ## Model and limits
52
+
53
+ The release uses rank-8 additive adapters and a scalar head: 16,517,633 trainable
54
+ parameters. The selected expanded branch's step 0 retains the refinement parent's
55
+ step-1,500 weights. The expanded branch's later step 159 was not selected; its
56
+ post-selection diagnostic gains are reported separately.
57
+
58
+ One temperature, 1.745822, was fitted on 510 separate known-family calibration
59
+ examples. Social IQA calibration remains limited: its top-label ECE is 8.30%.
60
+ No general intelligence, Jev-level quality or universal calibration claim follows.
61
+ Four-choice Banking77 is not the full 77-label task; benchmark grouping does not
62
+ rule out base-model pretraining overlap. The 1,024-token inference limit includes
63
+ the complete formatted candidate prompt, and longer inputs are rejected.
64
+
65
+ ## Files and provenance
66
+
67
+ - [model/](model/README.md): calibrated artifact, hash and pinned base requirements.
68
+ - [docs/](docs/README.md): layout and [reproduction guide](docs/reproduce.md).
69
+ - [source/](source/): complete committed project source, tests and usage guides.
70
+ - [results/](results/): final metrics, profiling report, tables and charts.
71
+ - [archive/](archive/README.md): index to preserved experiment history.
72
+
73
+ The original pointers remain authoritative:
74
+ [FINAL_MODEL.json](FINAL_MODEL.json) identifies the calibrated release,
75
+ [PROFILE_RESULTS.json](PROFILE_RESULTS.json) identifies verified profiling and
76
+ wrap-up evidence, and [CURRENT_SNAPSHOT.json](CURRENT_SNAPSHOT.json) identifies
77
+ the earlier training backup. Historical payload paths and hashes are preserved.
78
+
79
+ The existing `license: unknown` metadata is unchanged. Base-model and dataset
80
+ licenses remain separate; see the source's
81
+ [original data provenance](source/results/public-decisions-v1-manifest.json) and
82
+ [expanded data provenance](source/results/20260917-expanded-data/dataset-manifest.json).
83
+ Base weights and credentials are excluded. Hosted Jev calls require separate
84
+ authentication and were not exercised in this evaluation.