|
Download README.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 8.74 kB
-
https://huggingface.co/andyshu/opensysone/resolve/294f8ea1b877ac86188aa88eade4b81f5f190293/README.md
- Command line
-
hf download hf://andyshu/opensysone@294f8ea1b877ac86188aa88eade4b81f5f190293/README.md
-
curl -L -o README.md https://huggingface.co/andyshu/opensysone/resolve/294f8ea1b877ac86188aa88eade4b81f5f190293/README.md
8.74 kB
| license: unknown | |
| pipeline_tag: text-classification | |
| base_model: | |
| - Qwen/Qwen3-4B-Instruct-2507 | |
| - Qwen/Qwen3.5-2B | |
| tags: | |
| - opensysone | |
| - decision-scoring | |
| - lora | |
| - experimental | |
| # OpenSysOne | |
| **Credit:** OpenSysOne is inspired by [Jev](https://typesafe.ai/), TypeSafe.ai's | |
| System One model for structured decisions with probabilities. Credit goes to the | |
| TypeSafe team for motivating this project. OpenSysOne is an independent | |
| experimental implementation. | |
| OpenSysOne is an experimental natural-language decision scorer. It scores an | |
| arbitrary state, question and candidate answer, then normalizes mutually exclusive | |
| choices into probabilities. It does not generate chat responses. | |
| This repository backs up the ongoing three-machine training experiment. It | |
| contains custom adapter/head checkpoints, reproducible source, run provenance | |
| and small result files. The full pretrained base weights are referenced by pinned | |
| revision and are not duplicated here. **These are custom OpenSysOne artifacts, | |
| not a drop-in Transformers or PEFT adapter repository.** | |
| ## Current status | |
| The snapshot is **training-stage, uncalibrated**. Final held-out evaluation and | |
| calibrated deployment have not completed. The original campaign ends on | |
| **17 September 2026 at 18:16:10 UTC**, with training stopping by 16:00 UTC. | |
| `CURRENT_SNAPSHOT.json` identifies the backup and its source revision. | |
| `FINAL_MODEL.json` will identify the separately verified calibrated release when | |
| finalization and upload succeed; its absence means no final release is recorded. | |
| Validation audit on 17 September, 07:00 UTC: | |
| | Selected candidate | Step | Validation accuracy | Crossfit selection NLL | | |
| | --- | ---: | ---: | ---: | | |
| | Qwen3 4B, main run | 2,500 | 93.55% | 0.188640 | | |
| | Qwen3 4B, lower learning rate | 4,000 | 93.75% | 0.190020 | | |
| | Qwen3.5 2B | 2,000 | 89.84% | 0.255294 | | |
| | Qwen3 4B, refinement | 1,500 | 94.73% | 0.170150 | | |
| All use the same 512 validation decisions. Selection uses four source-group- | |
| disjoint temperature-crossfit folds with fixed seed 431; smaller macro-family NLL | |
| is better. This is **validation-selection evidence**, not a held-out quality claim. | |
| The 2B run stopped cleanly at step 6,000 after eight evaluations without a new | |
| best. At 07:07 UTC the original GX10 4B run stopped cleanly at step 4,380, | |
| preserving its selected step 2,500 and full resumable state. GX10 now verifies an | |
| expanded-data candidate from the refinement's frozen selected step 1,500, while | |
| both Spark 4B runs continue. See `source/EXPANDED_DATA.md` for its latest run | |
| state, verification evidence and the broader training mix. Independent trials | |
| retain the original absolute deadlines. | |
| ## Contents and reconstruction | |
| The [browser playground](source/PLAYGROUND.md) provides a context box, question, | |
| options list and probability bars, with a model selector. Its lightweight local | |
| server uses these custom checkpoints and the pinned base weights. It runs on | |
| GX10 over a loopback connection/SSH tunnel; this repository is a backup, not a | |
| hosted inference application. | |
| - `source/`: a committed source/documentation snapshot, including the training | |
| scripts, tests, Jev-compatible harness, findings and operational handover. | |
| - `snapshots/<id>/backup-manifest.json`: artifact roles, originating hosts/paths, | |
| checkpoint steps, model/data/source provenance and SHA-256 checksums. | |
| - `snapshots/<id>/artifacts/<candidate>/best.pt`: selected adapter/head weights | |
| and reconstruction metadata. These initial backups have no fitted temperature. | |
| - `snapshots/<id>/artifacts/<candidate>/checkpoint.pt`: resumable training state, | |
| including Adam and random states. Its step may be later than the selected best. | |
| - Expanded-data snapshots include transformed decision files and their source | |
| notices, split hashes, diagnostic exclusions and original dataset lineage. | |
| - `final/<campaign>/`: final calibrated model and evaluation evidence, created only | |
| after successful finalization and publication. | |
| Keep `checkpoint.pt`, its sibling `best.pt` and matching validation evidence | |
| together when resuming. Checkpoint metadata records the original local paths; | |
| reconstruction needs the pinned base/data and matching project source. The current | |
| harness works on the recorded GX10/Spark layout. This backup does not establish | |
| portability to a new operating system or a generic hosted inference service. | |
| Model and optimizer files use PyTorch serialization; use the project's known, | |
| hash-verified artifacts and its supplied reconstruction code. | |
| The main 4B model uses FP32, rank-8 additive adapters (alpha 16) and a learned | |
| scalar head initialized from pretrained yes-minus-no logits: 16,517,633 trainable | |
| parameters. The 2B text decoder uses rank 16 / alpha 32, with 16,821,249 trainable | |
| parameters. Training limits are 512 and 768 complete-chat tokens respectively; | |
| 1,024-token inference was separately verified. Training retains a 16 GiB per-job | |
| CUDA allocation cap. BF16 did not pass the project's numerical invariance gate. | |
| Pinned bases, recorded as Apache-2.0 in their provenance: | |
| - `Qwen/Qwen3-4B-Instruct-2507` at | |
| `cdbee75f17c01a7cc42f958dc650907174af0554`. | |
| - `Qwen/Qwen3.5-2B` at | |
| `15852e8c16360a2fea060d615a32b45270f8a8fc`. | |
| ## Data, evaluation and limitations | |
| The original public training sources are SNLI, BoolQ, ARC and four-choice Banking77 routing. | |
| Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences, | |
| raw/split hashes and group/deduplication audits are included in | |
| `source/results/public-decisions-v1-manifest.json`. That manifest credits | |
| Stanford NLP (SNLI), Google (BoolQ), AllenAI (ARC/Social IQA) and PolyAI (Banking77), | |
| and records their original dataset licences. The original repository's | |
| `license: unknown` metadata is retained; no new licence for the project artifacts | |
| is assigned by this backup. | |
| The expanded candidate adds official training rows from HellaSwag (MIT), PIQA | |
| (AFL-3.0 according to its creator's pinned README) and CommonsenseQA (MIT). | |
| After the 512-token filter there are 80,765 training decisions, including all | |
| 40,915 original retained examples. All four original reserved split files and | |
| their tokenized rows are unchanged. Another 383 retained new-source diagnostic | |
| decisions are excluded from both training and checkpoint selection. Source pins, | |
| credits, licences and transformations are in | |
| `source/results/20260917-expanded-data/dataset-manifest.json`. The unchanged | |
| selection set measures the original tasks; expansion alone is not evidence of | |
| better performance on the three added tasks. | |
| Validation selects checkpoints. Separate calibration fits one global temperature | |
| only after selection is frozen. Final test/holdout comparisons use the unchanged | |
| pretrained scorer, with a separately fitted base temperature and source-group | |
| uncertainty. Frozen grouping and exact deduplication do not rule out pretraining | |
| contamination or semantic duplicates. Four-choice Banking77 is not the full | |
| 77-label benchmark. No claim of general intelligence, Jev equivalence or | |
| calibration on unseen task families follows from the current validation scores. | |
| The harness implements local scoring, a loopback API, hosted Jev requests and | |
| response/timing comparisons. The local model identifies itself as OpenSysOne; | |
| compatibility with the request shape does not make it Jev. Local confidence is | |
| normalized entropy, not calibrated probability of correctness. Authenticated | |
| hosted Jev calls require `TYPESAFE_API_KEY` and have not been tested. Credentials | |
| and pretrained base weights are not part of this backup. Expanded-data snapshots | |
| may include the transformed training data, with upstream notices retained; | |
| raw upstream archives are referenced by pinned revision and checksum. | |
| <!-- opensysone-final:start --> | |
| ## Completed fleet deployment | |
| **This completed release supersedes the earlier training-stage Current status.** Descriptions of unfinished finalization and the training snapshot above are historical. `FINAL_MODEL.json` now identifies the verified calibrated release. | |
| Final artifact: [`final/20260916T194403396250Z-fleet/model.pt`](final/20260916T194403396250Z-fleet/model.pt). Integrity and provenance: [`backup_manifest.json`](final/20260916T194403396250Z-fleet/backup_manifest.json). Completed test/holdout results: [`metrics.json`](final/20260916T194403396250Z-fleet/evaluation/metrics.json). | |
| This contains trained adapter/head parameters and calibration metadata; pinned base weights are required separately. It scores explicit candidate decisions through the Jev-compatible harness. Public benchmark metrics do not establish general intelligence or universal calibration. | |
| Source revision: `6729461ccaad32e239c2148fa0c8f9ca23513a7a`. No license metadata has been changed. | |
| <!-- opensysone-final:end --> | |