--- license: unknown pipeline_tag: text-classification base_model: - Qwen/Qwen3-4B-Instruct-2507 - Qwen/Qwen3.5-2B tags: - opensysone - decision-scoring - lora - experimental --- # OpenSysOne OpenSysOne is an experimental natural-language decision scorer. It scores an arbitrary state, question and candidate answer, then normalizes mutually exclusive choices into probabilities. It does not generate chat responses. This repository backs up the ongoing three-machine training experiment. It contains custom adapter/head checkpoints, reproducible source, run provenance and small result files. The full pretrained base weights are referenced by pinned revision and are not duplicated here. **These are custom OpenSysOne artifacts, not a drop-in Transformers or PEFT adapter repository.** ## Current status The snapshot is **training-stage, uncalibrated**. Final held-out evaluation and calibrated deployment have not completed. The original campaign ends on **17 September 2026 at 18:16:10 UTC**, with training stopping by 16:00 UTC. `CURRENT_SNAPSHOT.json` identifies the backup and its source revision. `FINAL_MODEL.json` will identify the separately verified calibrated release when finalization and upload succeed; its absence means no final release is recorded. Validation audit on 17 September, 07:00 UTC: | Selected candidate | Step | Validation accuracy | Crossfit selection NLL | | --- | ---: | ---: | ---: | | Qwen3 4B, main run | 2,500 | 93.55% | 0.188640 | | Qwen3 4B, lower learning rate | 4,000 | 93.75% | 0.190020 | | Qwen3.5 2B | 2,000 | 89.84% | 0.255294 | | Qwen3 4B, refinement | 1,500 | 94.73% | 0.170150 | All use the same 512 validation decisions. Selection uses four source-group- disjoint temperature-crossfit folds with fixed seed 431; smaller macro-family NLL is better. This is **validation-selection evidence**, not a held-out quality claim. The 2B run stopped cleanly at step 6,000 after eight evaluations without a new best. At 07:07 UTC the original GX10 4B run stopped cleanly at step 4,380, preserving its selected step 2,500 and full resumable state. GX10 now verifies an expanded-data candidate from the refinement's frozen selected step 1,500, while both Spark 4B runs continue. See `source/EXPANDED_DATA.md` for its latest run state, verification evidence and the broader training mix. Independent trials retain the original absolute deadlines. ## Contents and reconstruction The [browser playground](source/PLAYGROUND.md) provides a context box, question, options list and probability bars, with a model selector. Its lightweight local server uses these custom checkpoints and the pinned base weights. It runs on GX10 over a loopback connection/SSH tunnel; this repository is a backup, not a hosted inference application. - `source/`: a committed source/documentation snapshot, including the training scripts, tests, Jev-compatible harness, findings and operational handover. - `snapshots//backup-manifest.json`: artifact roles, originating hosts/paths, checkpoint steps, model/data/source provenance and SHA-256 checksums. - `snapshots//artifacts//best.pt`: selected adapter/head weights and reconstruction metadata. These initial backups have no fitted temperature. - `snapshots//artifacts//checkpoint.pt`: resumable training state, including Adam and random states. Its step may be later than the selected best. - Expanded-data snapshots include transformed decision files and their source notices, split hashes, diagnostic exclusions and original dataset lineage. - `final//`: final calibrated model and evaluation evidence, created only after successful finalization and publication. Keep `checkpoint.pt`, its sibling `best.pt` and matching validation evidence together when resuming. Checkpoint metadata records the original local paths; reconstruction needs the pinned base/data and matching project source. The current harness works on the recorded GX10/Spark layout. This backup does not establish portability to a new operating system or a generic hosted inference service. Model and optimizer files use PyTorch serialization; use the project's known, hash-verified artifacts and its supplied reconstruction code. The main 4B model uses FP32, rank-8 additive adapters (alpha 16) and a learned scalar head initialized from pretrained yes-minus-no logits: 16,517,633 trainable parameters. The 2B text decoder uses rank 16 / alpha 32, with 16,821,249 trainable parameters. Training limits are 512 and 768 complete-chat tokens respectively; 1,024-token inference was separately verified. Training retains a 16 GiB per-job CUDA allocation cap. BF16 did not pass the project's numerical invariance gate. Pinned bases, recorded as Apache-2.0 in their provenance: - `Qwen/Qwen3-4B-Instruct-2507` at `cdbee75f17c01a7cc42f958dc650907174af0554`. - `Qwen/Qwen3.5-2B` at `15852e8c16360a2fea060d615a32b45270f8a8fc`. ## Data, evaluation and limitations The original public training sources are SNLI, BoolQ, ARC and four-choice Banking77 routing. Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences, raw/split hashes and group/deduplication audits are included in `source/results/public-decisions-v1-manifest.json`. That manifest credits Stanford NLP (SNLI), Google (BoolQ), AllenAI (ARC/Social IQA) and PolyAI (Banking77), and records their original dataset licences. The original repository's `license: unknown` metadata is retained; no new licence for the project artifacts is assigned by this backup. The expanded candidate adds official training rows from HellaSwag (MIT), PIQA (AFL-3.0 according to its creator's pinned README) and CommonsenseQA (MIT). After the 512-token filter there are 80,765 training decisions, including all 40,915 original retained examples. All four original reserved split files and their tokenized rows are unchanged. Another 383 retained new-source diagnostic decisions are excluded from both training and checkpoint selection. Source pins, credits, licences and transformations are in `source/results/20260917-expanded-data/dataset-manifest.json`. The unchanged selection set measures the original tasks; expansion alone is not evidence of better performance on the three added tasks. Validation selects checkpoints. Separate calibration fits one global temperature only after selection is frozen. Final test/holdout comparisons use the unchanged pretrained scorer, with a separately fitted base temperature and source-group uncertainty. Frozen grouping and exact deduplication do not rule out pretraining contamination or semantic duplicates. Four-choice Banking77 is not the full 77-label benchmark. No claim of general intelligence, Jev equivalence or calibration on unseen task families follows from the current validation scores. The harness implements local scoring, a loopback API, hosted Jev requests and response/timing comparisons. The local model identifies itself as OpenSysOne; compatibility with the request shape does not make it Jev. Local confidence is normalized entropy, not calibrated probability of correctness. Authenticated hosted Jev calls require `TYPESAFE_API_KEY` and have not been tested. Credentials and pretrained base weights are not part of this backup. Expanded-data snapshots may include the transformed training data, with upstream notices retained; raw upstream archives are referenced by pinned revision and checksum.