opensysone / README.md
andyshu's picture
Credit TypeSafe Jev as the inspiration for OpenSysOne
fde4324 verified
|
Raw History Blame
8.74 kB
---
license: unknown
pipeline_tag: text-classification
base_model:
- Qwen/Qwen3-4B-Instruct-2507
- Qwen/Qwen3.5-2B
tags:
- opensysone
- decision-scoring
- lora
- experimental
---
# OpenSysOne
**Credit:** OpenSysOne is inspired by [Jev](https://typesafe.ai/), TypeSafe.ai's
System One model for structured decisions with probabilities. Credit goes to the
TypeSafe team for motivating this project. OpenSysOne is an independent
experimental implementation.
OpenSysOne is an experimental natural-language decision scorer. It scores an
arbitrary state, question and candidate answer, then normalizes mutually exclusive
choices into probabilities. It does not generate chat responses.
This repository backs up the ongoing three-machine training experiment. It
contains custom adapter/head checkpoints, reproducible source, run provenance
and small result files. The full pretrained base weights are referenced by pinned
revision and are not duplicated here. **These are custom OpenSysOne artifacts,
not a drop-in Transformers or PEFT adapter repository.**
## Current status
The snapshot is **training-stage, uncalibrated**. Final held-out evaluation and
calibrated deployment have not completed. The original campaign ends on
**17 September 2026 at 18:16:10 UTC**, with training stopping by 16:00 UTC.
`CURRENT_SNAPSHOT.json` identifies the backup and its source revision.
`FINAL_MODEL.json` will identify the separately verified calibrated release when
finalization and upload succeed; its absence means no final release is recorded.
Validation audit on 17 September, 07:00 UTC:
| Selected candidate | Step | Validation accuracy | Crossfit selection NLL |
| --- | ---: | ---: | ---: |
| Qwen3 4B, main run | 2,500 | 93.55% | 0.188640 |
| Qwen3 4B, lower learning rate | 4,000 | 93.75% | 0.190020 |
| Qwen3.5 2B | 2,000 | 89.84% | 0.255294 |
| Qwen3 4B, refinement | 1,500 | 94.73% | 0.170150 |
All use the same 512 validation decisions. Selection uses four source-group-
disjoint temperature-crossfit folds with fixed seed 431; smaller macro-family NLL
is better. This is **validation-selection evidence**, not a held-out quality claim.
The 2B run stopped cleanly at step 6,000 after eight evaluations without a new
best. At 07:07 UTC the original GX10 4B run stopped cleanly at step 4,380,
preserving its selected step 2,500 and full resumable state. GX10 now verifies an
expanded-data candidate from the refinement's frozen selected step 1,500, while
both Spark 4B runs continue. See `source/EXPANDED_DATA.md` for its latest run
state, verification evidence and the broader training mix. Independent trials
retain the original absolute deadlines.
## Contents and reconstruction
The [browser playground](source/PLAYGROUND.md) provides a context box, question,
options list and probability bars, with a model selector. Its lightweight local
server uses these custom checkpoints and the pinned base weights. It runs on
GX10 over a loopback connection/SSH tunnel; this repository is a backup, not a
hosted inference application.
- `source/`: a committed source/documentation snapshot, including the training
scripts, tests, Jev-compatible harness, findings and operational handover.
- `snapshots/<id>/backup-manifest.json`: artifact roles, originating hosts/paths,
checkpoint steps, model/data/source provenance and SHA-256 checksums.
- `snapshots/<id>/artifacts/<candidate>/best.pt`: selected adapter/head weights
and reconstruction metadata. These initial backups have no fitted temperature.
- `snapshots/<id>/artifacts/<candidate>/checkpoint.pt`: resumable training state,
including Adam and random states. Its step may be later than the selected best.
- Expanded-data snapshots include transformed decision files and their source
notices, split hashes, diagnostic exclusions and original dataset lineage.
- `final/<campaign>/`: final calibrated model and evaluation evidence, created only
after successful finalization and publication.
Keep `checkpoint.pt`, its sibling `best.pt` and matching validation evidence
together when resuming. Checkpoint metadata records the original local paths;
reconstruction needs the pinned base/data and matching project source. The current
harness works on the recorded GX10/Spark layout. This backup does not establish
portability to a new operating system or a generic hosted inference service.
Model and optimizer files use PyTorch serialization; use the project's known,
hash-verified artifacts and its supplied reconstruction code.
The main 4B model uses FP32, rank-8 additive adapters (alpha 16) and a learned
scalar head initialized from pretrained yes-minus-no logits: 16,517,633 trainable
parameters. The 2B text decoder uses rank 16 / alpha 32, with 16,821,249 trainable
parameters. Training limits are 512 and 768 complete-chat tokens respectively;
1,024-token inference was separately verified. Training retains a 16 GiB per-job
CUDA allocation cap. BF16 did not pass the project's numerical invariance gate.
Pinned bases, recorded as Apache-2.0 in their provenance:
- `Qwen/Qwen3-4B-Instruct-2507` at
`cdbee75f17c01a7cc42f958dc650907174af0554`.
- `Qwen/Qwen3.5-2B` at
`15852e8c16360a2fea060d615a32b45270f8a8fc`.
## Data, evaluation and limitations
The original public training sources are SNLI, BoolQ, ARC and four-choice Banking77 routing.
Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences,
raw/split hashes and group/deduplication audits are included in
`source/results/public-decisions-v1-manifest.json`. That manifest credits
Stanford NLP (SNLI), Google (BoolQ), AllenAI (ARC/Social IQA) and PolyAI (Banking77),
and records their original dataset licences. The original repository's
`license: unknown` metadata is retained; no new licence for the project artifacts
is assigned by this backup.
The expanded candidate adds official training rows from HellaSwag (MIT), PIQA
(AFL-3.0 according to its creator's pinned README) and CommonsenseQA (MIT).
After the 512-token filter there are 80,765 training decisions, including all
40,915 original retained examples. All four original reserved split files and
their tokenized rows are unchanged. Another 383 retained new-source diagnostic
decisions are excluded from both training and checkpoint selection. Source pins,
credits, licences and transformations are in
`source/results/20260917-expanded-data/dataset-manifest.json`. The unchanged
selection set measures the original tasks; expansion alone is not evidence of
better performance on the three added tasks.
Validation selects checkpoints. Separate calibration fits one global temperature
only after selection is frozen. Final test/holdout comparisons use the unchanged
pretrained scorer, with a separately fitted base temperature and source-group
uncertainty. Frozen grouping and exact deduplication do not rule out pretraining
contamination or semantic duplicates. Four-choice Banking77 is not the full
77-label benchmark. No claim of general intelligence, Jev equivalence or
calibration on unseen task families follows from the current validation scores.
The harness implements local scoring, a loopback API, hosted Jev requests and
response/timing comparisons. The local model identifies itself as OpenSysOne;
compatibility with the request shape does not make it Jev. Local confidence is
normalized entropy, not calibrated probability of correctness. Authenticated
hosted Jev calls require `TYPESAFE_API_KEY` and have not been tested. Credentials
and pretrained base weights are not part of this backup. Expanded-data snapshots
may include the transformed training data, with upstream notices retained;
raw upstream archives are referenced by pinned revision and checksum.
<!-- opensysone-final:start -->
## Completed fleet deployment
**This completed release supersedes the earlier training-stage Current status.** Descriptions of unfinished finalization and the training snapshot above are historical. `FINAL_MODEL.json` now identifies the verified calibrated release.
Final artifact: [`final/20260916T194403396250Z-fleet/model.pt`](final/20260916T194403396250Z-fleet/model.pt). Integrity and provenance: [`backup_manifest.json`](final/20260916T194403396250Z-fleet/backup_manifest.json). Completed test/holdout results: [`metrics.json`](final/20260916T194403396250Z-fleet/evaluation/metrics.json).
This contains trained adapter/head parameters and calibration metadata; pinned base weights are required separately. It scores explicit candidate decisions through the Jev-compatible harness. Public benchmark metrics do not establish general intelligence or universal calibration.
Source revision: `6729461ccaad32e239c2148fa0c8f9ca23513a7a`. No license metadata has been changed.
<!-- opensysone-final:end -->