Publish verified OpenSysOne directory index and manifest pointer
Browse files- PUBLICATION.json +10 -0
- README.md +69 -139
PUBLICATION.json
ADDED
|
@@ -0,0 +1,10 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"accepted_lfs_attributes": null,
|
| 3 |
+
"format": "opensysone-publication-pointer-v1",
|
| 4 |
+
"manifest_path": "publication-manifest.json",
|
| 5 |
+
"manifest_sha256": "a92002e3605de0e9cdafe8691fca44e482f0465da26c189958838c4df01250b6",
|
| 6 |
+
"payload_commit": "75c3d90272aa383c2633df564a78b56370855b54",
|
| 7 |
+
"repo_id": "andyshu/opensysone",
|
| 8 |
+
"source_commit": "363ea4987da8e4a80ae8595e52f4f9db0866d4e7",
|
| 9 |
+
"verified_utc": "2026-09-17T10:04:57.045039+00:00"
|
| 10 |
+
}
|
README.md
CHANGED
|
@@ -13,142 +13,72 @@ tags:
|
|
| 13 |
|
| 14 |
# OpenSysOne
|
| 15 |
|
| 16 |
-
**
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
-
|
| 69 |
-
|
| 70 |
-
-
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
portability to a new operating system or a generic hosted inference service.
|
| 86 |
-
Model and optimizer files use PyTorch serialization; use the project's known,
|
| 87 |
-
hash-verified artifacts and its supplied reconstruction code.
|
| 88 |
-
|
| 89 |
-
The main 4B model uses FP32, rank-8 additive adapters (alpha 16) and a learned
|
| 90 |
-
scalar head initialized from pretrained yes-minus-no logits: 16,517,633 trainable
|
| 91 |
-
parameters. The 2B text decoder uses rank 16 / alpha 32, with 16,821,249 trainable
|
| 92 |
-
parameters. Training limits are 512 and 768 complete-chat tokens respectively;
|
| 93 |
-
1,024-token inference was separately verified. Training retains a 16 GiB per-job
|
| 94 |
-
CUDA allocation cap. BF16 did not pass the project's numerical invariance gate.
|
| 95 |
-
|
| 96 |
-
Pinned bases, recorded as Apache-2.0 in their provenance:
|
| 97 |
-
|
| 98 |
-
- `Qwen/Qwen3-4B-Instruct-2507` at
|
| 99 |
-
`cdbee75f17c01a7cc42f958dc650907174af0554`.
|
| 100 |
-
- `Qwen/Qwen3.5-2B` at
|
| 101 |
-
`15852e8c16360a2fea060d615a32b45270f8a8fc`.
|
| 102 |
-
|
| 103 |
-
## Data, evaluation and limitations
|
| 104 |
-
|
| 105 |
-
The original public training sources are SNLI, BoolQ, ARC and four-choice Banking77 routing.
|
| 106 |
-
Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences,
|
| 107 |
-
raw/split hashes and group/deduplication audits are included in
|
| 108 |
-
`source/results/public-decisions-v1-manifest.json`. That manifest credits
|
| 109 |
-
Stanford NLP (SNLI), Google (BoolQ), AllenAI (ARC/Social IQA) and PolyAI (Banking77),
|
| 110 |
-
and records their original dataset licences. The original repository's
|
| 111 |
-
`license: unknown` metadata is retained; no new licence for the project artifacts
|
| 112 |
-
is assigned by this backup.
|
| 113 |
-
|
| 114 |
-
The expanded candidate adds official training rows from HellaSwag (MIT), PIQA
|
| 115 |
-
(AFL-3.0 according to its creator's pinned README) and CommonsenseQA (MIT).
|
| 116 |
-
After the 512-token filter there are 80,765 training decisions, including all
|
| 117 |
-
40,915 original retained examples. All four original reserved split files and
|
| 118 |
-
their tokenized rows are unchanged. Another 383 retained new-source diagnostic
|
| 119 |
-
decisions are excluded from both training and checkpoint selection. Source pins,
|
| 120 |
-
credits, licences and transformations are in
|
| 121 |
-
`source/results/20260917-expanded-data/dataset-manifest.json`. The unchanged
|
| 122 |
-
selection set measures the original tasks; expansion alone is not evidence of
|
| 123 |
-
better performance on the three added tasks.
|
| 124 |
-
|
| 125 |
-
Validation selects checkpoints. Separate calibration fits one global temperature
|
| 126 |
-
only after selection is frozen. Final test/holdout comparisons use the unchanged
|
| 127 |
-
pretrained scorer, with a separately fitted base temperature and source-group
|
| 128 |
-
uncertainty. Frozen grouping and exact deduplication do not rule out pretraining
|
| 129 |
-
contamination or semantic duplicates. Four-choice Banking77 is not the full
|
| 130 |
-
77-label benchmark. No claim of general intelligence, Jev equivalence or
|
| 131 |
-
calibration on unseen task families follows from the current validation scores.
|
| 132 |
-
|
| 133 |
-
The harness implements local scoring, a loopback API, hosted Jev requests and
|
| 134 |
-
response/timing comparisons. The local model identifies itself as OpenSysOne;
|
| 135 |
-
compatibility with the request shape does not make it Jev. Local confidence is
|
| 136 |
-
normalized entropy, not calibrated probability of correctness. Authenticated
|
| 137 |
-
hosted Jev calls require `TYPESAFE_API_KEY` and have not been tested. Credentials
|
| 138 |
-
and pretrained base weights are not part of this backup. Expanded-data snapshots
|
| 139 |
-
may include the transformed training data, with upstream notices retained;
|
| 140 |
-
raw upstream archives are referenced by pinned revision and checksum.
|
| 141 |
-
|
| 142 |
-
<!-- opensysone-final:start -->
|
| 143 |
-
|
| 144 |
-
## Completed fleet deployment
|
| 145 |
-
|
| 146 |
-
**This completed release supersedes the earlier training-stage Current status.** Descriptions of unfinished finalization and the training snapshot above are historical. `FINAL_MODEL.json` now identifies the verified calibrated release.
|
| 147 |
-
|
| 148 |
-
Final artifact: [`final/20260916T194403396250Z-fleet/model.pt`](final/20260916T194403396250Z-fleet/model.pt). Integrity and provenance: [`backup_manifest.json`](final/20260916T194403396250Z-fleet/backup_manifest.json). Completed test/holdout results: [`metrics.json`](final/20260916T194403396250Z-fleet/evaluation/metrics.json).
|
| 149 |
-
|
| 150 |
-
This contains trained adapter/head parameters and calibration metadata; pinned base weights are required separately. It scores explicit candidate decisions through the Jev-compatible harness. Public benchmark metrics do not establish general intelligence or universal calibration.
|
| 151 |
-
|
| 152 |
-
Source revision: `6729461ccaad32e239c2148fa0c8f9ca23513a7a`. No license metadata has been changed.
|
| 153 |
-
|
| 154 |
-
<!-- opensysone-final:end -->
|
|
|
|
| 13 |
|
| 14 |
# OpenSysOne
|
| 15 |
|
| 16 |
+
**Inspired by [Jev](https://typesafe.ai/), TypeSafe.ai's System One model.**
|
| 17 |
+
Credit goes to the TypeSafe team for inspiring this project's exploration of
|
| 18 |
+
structured decisions with probabilities. OpenSysOne is an independent experimental implementation; API
|
| 19 |
+
compatibility does not establish Jev equivalence.
|
| 20 |
+
|
| 21 |
+
The **completed 4B release** scores a state, question and explicit candidate
|
| 22 |
+
answers, returning probabilities over those choices. Training, separate
|
| 23 |
+
calibration, final evaluation and local API verification completed on
|
| 24 |
+
17 September 2026.
|
| 25 |
+
|
| 26 |
+
Start with the [model and reconstruction notes](model/README.md),
|
| 27 |
+
[results report](results/report.md), or [publication guide](docs/README.md).
|
| 28 |
+
The calibrated artifact is [model/model.pt](model/model.pt).
|
| 29 |
+
It contains custom OpenSysOne adapter/head weights and metadata. The pinned
|
| 30 |
+
Qwen3-4B-Instruct-2507 base is required separately; this is not a standalone
|
| 31 |
+
Transformers model or a standard PEFT adapter package.
|
| 32 |
+
|
| 33 |
+
## Measured results
|
| 34 |
+
|
| 35 |
+
The full comparison uses the unchanged pretrained yes/no verifier, with a separate
|
| 36 |
+
temperature fitted for each model. Intervals are paired 95% source-group bootstrap
|
| 37 |
+
intervals for selected minus base accuracy.
|
| 38 |
+
|
| 39 |
+
| Evaluation | Decisions | Selected | Base verifier | Accuracy gain (95% interval) |
|
| 40 |
+
| --- | ---: | ---: | ---: | ---: |
|
| 41 |
+
| Known-family test | 2,042 | 92.90% | 84.48% | +8.42 pp [6.85, 9.89] |
|
| 42 |
+
| Social IQA family holdout | 768 | 72.92% | 70.31% | +2.60 pp [0.13, 5.34] |
|
| 43 |
+
|
| 44 |
+
On a separate matched 320-decision profile, selected accuracy was 89.06%, versus
|
| 45 |
+
80.94% for the base verifier and 86.25% for a base model using one constrained
|
| 46 |
+
answer-label token. The selected scorer was **slower on all 12 profiled workloads**:
|
| 47 |
+
1.11–1.17× the verifier latency and 2.18–15.58× the label baseline latency.
|
| 48 |
+
These are warm, serial FP32 measurements on one GB10, not concurrent-serving
|
| 49 |
+
throughput or comparisons with generated reasoning. See [tables and charts](results/).
|
| 50 |
+
|
| 51 |
+
## Model and limits
|
| 52 |
+
|
| 53 |
+
The release uses rank-8 additive adapters and a scalar head: 16,517,633 trainable
|
| 54 |
+
parameters. The selected expanded branch's step 0 retains the refinement parent's
|
| 55 |
+
step-1,500 weights. The expanded branch's later step 159 was not selected; its
|
| 56 |
+
post-selection diagnostic gains are reported separately.
|
| 57 |
+
|
| 58 |
+
One temperature, 1.745822, was fitted on 510 separate known-family calibration
|
| 59 |
+
examples. Social IQA calibration remains limited: its top-label ECE is 8.30%.
|
| 60 |
+
No general intelligence, Jev-level quality or universal calibration claim follows.
|
| 61 |
+
Four-choice Banking77 is not the full 77-label task; benchmark grouping does not
|
| 62 |
+
rule out base-model pretraining overlap. The 1,024-token inference limit includes
|
| 63 |
+
the complete formatted candidate prompt, and longer inputs are rejected.
|
| 64 |
+
|
| 65 |
+
## Files and provenance
|
| 66 |
+
|
| 67 |
+
- [model/](model/README.md): calibrated artifact, hash and pinned base requirements.
|
| 68 |
+
- [docs/](docs/README.md): layout and [reproduction guide](docs/reproduce.md).
|
| 69 |
+
- [source/](source/): complete committed project source, tests and usage guides.
|
| 70 |
+
- [results/](results/): final metrics, profiling report, tables and charts.
|
| 71 |
+
- [archive/](archive/README.md): index to preserved experiment history.
|
| 72 |
+
|
| 73 |
+
The original pointers remain authoritative:
|
| 74 |
+
[FINAL_MODEL.json](FINAL_MODEL.json) identifies the calibrated release,
|
| 75 |
+
[PROFILE_RESULTS.json](PROFILE_RESULTS.json) identifies verified profiling and
|
| 76 |
+
wrap-up evidence, and [CURRENT_SNAPSHOT.json](CURRENT_SNAPSHOT.json) identifies
|
| 77 |
+
the earlier training backup. Historical payload paths and hashes are preserved.
|
| 78 |
+
|
| 79 |
+
The existing `license: unknown` metadata is unchanged. Base-model and dataset
|
| 80 |
+
licenses remain separate; see the source's
|
| 81 |
+
[original data provenance](source/results/public-decisions-v1-manifest.json) and
|
| 82 |
+
[expanded data provenance](source/results/20260917-expanded-data/dataset-manifest.json).
|
| 83 |
+
Base weights and credentials are excluded. Hosted Jev calls require separate
|
| 84 |
+
authentication and were not exercised in this evaluation.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|