opensysone / README.md
andyshu's picture
Credit TypeSafe Jev as the inspiration for OpenSysOne
fde4324 verified
|
Raw History Blame
8.74 kB
metadata
license: unknown
pipeline_tag: text-classification
base_model:
  - Qwen/Qwen3-4B-Instruct-2507
  - Qwen/Qwen3.5-2B
tags:
  - opensysone
  - decision-scoring
  - lora
  - experimental

OpenSysOne

Credit: OpenSysOne is inspired by Jev, TypeSafe.ai's System One model for structured decisions with probabilities. Credit goes to the TypeSafe team for motivating this project. OpenSysOne is an independent experimental implementation.

OpenSysOne is an experimental natural-language decision scorer. It scores an arbitrary state, question and candidate answer, then normalizes mutually exclusive choices into probabilities. It does not generate chat responses.

This repository backs up the ongoing three-machine training experiment. It contains custom adapter/head checkpoints, reproducible source, run provenance and small result files. The full pretrained base weights are referenced by pinned revision and are not duplicated here. These are custom OpenSysOne artifacts, not a drop-in Transformers or PEFT adapter repository.

Current status

The snapshot is training-stage, uncalibrated. Final held-out evaluation and calibrated deployment have not completed. The original campaign ends on 17 September 2026 at 18:16:10 UTC, with training stopping by 16:00 UTC. CURRENT_SNAPSHOT.json identifies the backup and its source revision. FINAL_MODEL.json will identify the separately verified calibrated release when finalization and upload succeed; its absence means no final release is recorded.

Validation audit on 17 September, 07:00 UTC:

Selected candidate Step Validation accuracy Crossfit selection NLL
Qwen3 4B, main run 2,500 93.55% 0.188640
Qwen3 4B, lower learning rate 4,000 93.75% 0.190020
Qwen3.5 2B 2,000 89.84% 0.255294
Qwen3 4B, refinement 1,500 94.73% 0.170150

All use the same 512 validation decisions. Selection uses four source-group- disjoint temperature-crossfit folds with fixed seed 431; smaller macro-family NLL is better. This is validation-selection evidence, not a held-out quality claim. The 2B run stopped cleanly at step 6,000 after eight evaluations without a new best. At 07:07 UTC the original GX10 4B run stopped cleanly at step 4,380, preserving its selected step 2,500 and full resumable state. GX10 now verifies an expanded-data candidate from the refinement's frozen selected step 1,500, while both Spark 4B runs continue. See source/EXPANDED_DATA.md for its latest run state, verification evidence and the broader training mix. Independent trials retain the original absolute deadlines.

Contents and reconstruction

The browser playground provides a context box, question, options list and probability bars, with a model selector. Its lightweight local server uses these custom checkpoints and the pinned base weights. It runs on GX10 over a loopback connection/SSH tunnel; this repository is a backup, not a hosted inference application.

  • source/: a committed source/documentation snapshot, including the training scripts, tests, Jev-compatible harness, findings and operational handover.
  • snapshots/<id>/backup-manifest.json: artifact roles, originating hosts/paths, checkpoint steps, model/data/source provenance and SHA-256 checksums.
  • snapshots/<id>/artifacts/<candidate>/best.pt: selected adapter/head weights and reconstruction metadata. These initial backups have no fitted temperature.
  • snapshots/<id>/artifacts/<candidate>/checkpoint.pt: resumable training state, including Adam and random states. Its step may be later than the selected best.
  • Expanded-data snapshots include transformed decision files and their source notices, split hashes, diagnostic exclusions and original dataset lineage.
  • final/<campaign>/: final calibrated model and evaluation evidence, created only after successful finalization and publication.

Keep checkpoint.pt, its sibling best.pt and matching validation evidence together when resuming. Checkpoint metadata records the original local paths; reconstruction needs the pinned base/data and matching project source. The current harness works on the recorded GX10/Spark layout. This backup does not establish portability to a new operating system or a generic hosted inference service. Model and optimizer files use PyTorch serialization; use the project's known, hash-verified artifacts and its supplied reconstruction code.

The main 4B model uses FP32, rank-8 additive adapters (alpha 16) and a learned scalar head initialized from pretrained yes-minus-no logits: 16,517,633 trainable parameters. The 2B text decoder uses rank 16 / alpha 32, with 16,821,249 trainable parameters. Training limits are 512 and 768 complete-chat tokens respectively; 1,024-token inference was separately verified. Training retains a 16 GiB per-job CUDA allocation cap. BF16 did not pass the project's numerical invariance gate.

Pinned bases, recorded as Apache-2.0 in their provenance:

  • Qwen/Qwen3-4B-Instruct-2507 at cdbee75f17c01a7cc42f958dc650907174af0554.
  • Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc.

Data, evaluation and limitations

The original public training sources are SNLI, BoolQ, ARC and four-choice Banking77 routing. Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences, raw/split hashes and group/deduplication audits are included in source/results/public-decisions-v1-manifest.json. That manifest credits Stanford NLP (SNLI), Google (BoolQ), AllenAI (ARC/Social IQA) and PolyAI (Banking77), and records their original dataset licences. The original repository's license: unknown metadata is retained; no new licence for the project artifacts is assigned by this backup.

The expanded candidate adds official training rows from HellaSwag (MIT), PIQA (AFL-3.0 according to its creator's pinned README) and CommonsenseQA (MIT). After the 512-token filter there are 80,765 training decisions, including all 40,915 original retained examples. All four original reserved split files and their tokenized rows are unchanged. Another 383 retained new-source diagnostic decisions are excluded from both training and checkpoint selection. Source pins, credits, licences and transformations are in source/results/20260917-expanded-data/dataset-manifest.json. The unchanged selection set measures the original tasks; expansion alone is not evidence of better performance on the three added tasks.

Validation selects checkpoints. Separate calibration fits one global temperature only after selection is frozen. Final test/holdout comparisons use the unchanged pretrained scorer, with a separately fitted base temperature and source-group uncertainty. Frozen grouping and exact deduplication do not rule out pretraining contamination or semantic duplicates. Four-choice Banking77 is not the full 77-label benchmark. No claim of general intelligence, Jev equivalence or calibration on unseen task families follows from the current validation scores.

The harness implements local scoring, a loopback API, hosted Jev requests and response/timing comparisons. The local model identifies itself as OpenSysOne; compatibility with the request shape does not make it Jev. Local confidence is normalized entropy, not calibrated probability of correctness. Authenticated hosted Jev calls require TYPESAFE_API_KEY and have not been tested. Credentials and pretrained base weights are not part of this backup. Expanded-data snapshots may include the transformed training data, with upstream notices retained; raw upstream archives are referenced by pinned revision and checksum.

Completed fleet deployment

This completed release supersedes the earlier training-stage Current status. Descriptions of unfinished finalization and the training snapshot above are historical. FINAL_MODEL.json now identifies the verified calibrated release.

Final artifact: final/20260916T194403396250Z-fleet/model.pt. Integrity and provenance: backup_manifest.json. Completed test/holdout results: metrics.json.

This contains trained adapter/head parameters and calibration metadata; pinned base weights are required separately. It scores explicit candidate decisions through the Jev-compatible harness. Public benchmark metrics do not establish general intelligence or universal calibration.

Source revision: 6729461ccaad32e239c2148fa0c8f9ca23513a7a. No license metadata has been changed.