andyshu commited on
Commit
b213728
·
verified ·
1 Parent(s): 2d5c26a

Publish verified OpenSysOne training snapshot pointer and model card

Browse files
Files changed (2) hide show
  1. CURRENT_SNAPSHOT.json +14 -0
  2. README.md +107 -0
CURRENT_SNAPSHOT.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "repo_id": "andyshu/opensysone",
3
+ "snapshot_id": "20260917T022835Z-hf-backup",
4
+ "path": "snapshots/20260917T022835Z-hf-backup",
5
+ "payload_commit": "2d5c26ab7929b8c858aeea88842781fdbf79cd54",
6
+ "source_commit": "93602e6bbc494c1aa225f1b8a540fa99f5669771",
7
+ "training_source_commits": [
8
+ "4a60423c39d70f8d50472ce4f4f7fa4a4bd9fce1"
9
+ ],
10
+ "manifest_path": "publications/20260917T022835Z-hf-backup/93602e6bbc494c1aa225f1b8a540fa99f5669771/manifest.json",
11
+ "manifest_sha256": "f47912cbf5d81fad8a231eddf7dad0063fb8f96b9052890e9eb9bd1d532beff7",
12
+ "published_utc": "2026-09-17T02:53:53.038652+00:00",
13
+ "calibrated": false
14
+ }
README.md CHANGED
@@ -1,3 +1,110 @@
1
  ---
2
  license: unknown
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: unknown
3
+ pipeline_tag: text-classification
4
+ base_model:
5
+ - Qwen/Qwen3-4B-Instruct-2507
6
+ - Qwen/Qwen3.5-2B
7
+ tags:
8
+ - opensysone
9
+ - decision-scoring
10
+ - lora
11
+ - experimental
12
  ---
13
+
14
+ # OpenSysOne
15
+
16
+ OpenSysOne is an experimental natural-language decision scorer. It scores an
17
+ arbitrary state, question and candidate answer, then normalizes mutually exclusive
18
+ choices into probabilities. It does not generate chat responses.
19
+
20
+ This repository backs up the ongoing three-machine training experiment. It
21
+ contains custom adapter/head checkpoints, reproducible source, run provenance
22
+ and small result files. The full pretrained base weights are referenced by pinned
23
+ revision and are not duplicated here. **These are custom OpenSysOne artifacts,
24
+ not a drop-in Transformers or PEFT adapter repository.**
25
+
26
+ ## Current status
27
+
28
+ The snapshot is **training-stage, uncalibrated**. Final held-out evaluation and
29
+ calibrated deployment have not completed. The original campaign ends on
30
+ **17 September 2026 at 18:16:10 UTC**, with training stopping by 16:00 UTC.
31
+ `CURRENT_SNAPSHOT.json` identifies the backup and its source revision.
32
+ `FINAL_MODEL.json` will identify the separately verified calibrated release when
33
+ finalization and upload succeed; its absence means no final release is recorded.
34
+
35
+ Validation audit on 17 September, 02:10–02:15 UTC:
36
+
37
+ | Selected candidate | Step | Validation accuracy | Crossfit selection NLL |
38
+ | --- | ---: | ---: | ---: |
39
+ | Qwen3 4B, main run | 2,500 | 93.55% | 0.188640 |
40
+ | Qwen3 4B, lower learning rate | 2,500 | 92.58% | 0.218012 |
41
+ | Qwen3.5 2B | 2,000 | 89.84% | 0.255294 |
42
+
43
+ All use the same 512 validation decisions. Selection uses four source-group-
44
+ disjoint temperature-crossfit folds with fixed seed 431; smaller macro-family NLL
45
+ is better. This is **validation-selection evidence**, not a held-out quality claim.
46
+ The 2B run stopped cleanly at step 6,000 after eight evaluations without a new
47
+ best. Both 4B runs continued, and another lower-rate 4B refinement was started
48
+ from the main run's selected weights. See `source/NEXT_STEPS.md` for findings and
49
+ why independent trials were retained over unmeasured distributed training.
50
+
51
+ ## Contents and reconstruction
52
+
53
+ - `source/`: a committed source/documentation snapshot, including the training
54
+ scripts, tests, Jev-compatible harness, findings and operational handover.
55
+ - `snapshots/<id>/backup-manifest.json`: artifact roles, originating hosts/paths,
56
+ checkpoint steps, model/data/source provenance and SHA-256 checksums.
57
+ - `snapshots/<id>/artifacts/<candidate>/best.pt`: selected adapter/head weights
58
+ and reconstruction metadata. These initial backups have no fitted temperature.
59
+ - `snapshots/<id>/artifacts/<candidate>/checkpoint.pt`: resumable training state,
60
+ including Adam and random states. Its step may be later than the selected best.
61
+ - `final/<campaign>/`: final calibrated model and evaluation evidence, created only
62
+ after successful finalization and publication.
63
+
64
+ Keep `checkpoint.pt`, its sibling `best.pt` and matching validation evidence
65
+ together when resuming. Checkpoint metadata records the original local paths;
66
+ reconstruction needs the pinned base/data and matching project source. The current
67
+ harness works on the recorded GX10/Spark layout. This backup does not establish
68
+ portability to a new operating system or a generic hosted inference service.
69
+ Model and optimizer files use PyTorch serialization; use the project's known,
70
+ hash-verified artifacts and its supplied reconstruction code.
71
+
72
+ The main 4B model uses FP32, rank-8 additive adapters (alpha 16) and a learned
73
+ scalar head initialized from pretrained yes-minus-no logits: 16,517,633 trainable
74
+ parameters. The 2B text decoder uses rank 16 / alpha 32, with 16,821,249 trainable
75
+ parameters. Training limits are 512 and 768 complete-chat tokens respectively;
76
+ 1,024-token inference was separately verified. Training retains a 16 GiB per-job
77
+ CUDA allocation cap. BF16 did not pass the project's numerical invariance gate.
78
+
79
+ Pinned bases, recorded as Apache-2.0 in their provenance:
80
+
81
+ - `Qwen/Qwen3-4B-Instruct-2507` at
82
+ `cdbee75f17c01a7cc42f958dc650907174af0554`.
83
+ - `Qwen/Qwen3.5-2B` at
84
+ `15852e8c16360a2fea060d615a32b45270f8a8fc`.
85
+
86
+ ## Data, evaluation and limitations
87
+
88
+ Public training sources are SNLI, BoolQ, ARC and four-choice Banking77 routing.
89
+ Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences,
90
+ raw/split hashes and group/deduplication audits are included in
91
+ `source/results/public-decisions-v1-manifest.json`. That manifest credits
92
+ Stanford NLP (SNLI), Google (BoolQ), AllenAI (ARC/Social IQA) and PolyAI (Banking77),
93
+ and records their original dataset licences. The original repository's
94
+ `license: unknown` metadata is retained; no new licence for the project artifacts
95
+ is assigned by this backup.
96
+
97
+ Validation selects checkpoints. Separate calibration fits one global temperature
98
+ only after selection is frozen. Final test/holdout comparisons use the unchanged
99
+ pretrained scorer, with a separately fitted base temperature and source-group
100
+ uncertainty. Frozen grouping and exact deduplication do not rule out pretraining
101
+ contamination or semantic duplicates. Four-choice Banking77 is not the full
102
+ 77-label benchmark. No claim of general intelligence, Jev equivalence or
103
+ calibration on unseen task families follows from the current validation scores.
104
+
105
+ The harness implements local scoring, a loopback API, hosted Jev requests and
106
+ response/timing comparisons. The local model identifies itself as OpenSysOne;
107
+ compatibility with the request shape does not make it Jev. Local confidence is
108
+ normalized entropy, not calibrated probability of correctness. Authenticated
109
+ hosted Jev calls require `TYPESAFE_API_KEY` and have not been tested. Credentials,
110
+ pretrained base weights and raw training datasets are not part of this backup.