Download source/HF_MODEL_CARD.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 7.43 kB
-
https://huggingface.co/andyshu/opensysone/resolve/58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/HF_MODEL_CARD.md
- Command line
-
hf download hf://andyshu/opensysone@58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/HF_MODEL_CARD.md
-
curl -L -o HF_MODEL_CARD.md https://huggingface.co/andyshu/opensysone/resolve/58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/HF_MODEL_CARD.md
license: unknown
pipeline_tag: text-classification
base_model:
- Qwen/Qwen3-4B-Instruct-2507
- Qwen/Qwen3.5-2B
tags:
- opensysone
- decision-scoring
- lora
- experimental
OpenSysOne
OpenSysOne is an experimental natural-language decision scorer. It scores an arbitrary state, question and candidate answer, then normalizes mutually exclusive choices into probabilities. It does not generate chat responses.
This repository backs up the ongoing three-machine training experiment. It contains custom adapter/head checkpoints, reproducible source, run provenance and small result files. The full pretrained base weights are referenced by pinned revision and are not duplicated here. These are custom OpenSysOne artifacts, not a drop-in Transformers or PEFT adapter repository.
Current status
The snapshot is training-stage, uncalibrated. Final held-out evaluation and
calibrated deployment have not completed. The original campaign ends on
17 September 2026 at 18:16:10 UTC, with training stopping by 16:00 UTC.
CURRENT_SNAPSHOT.json identifies the backup and its source revision.
FINAL_MODEL.json will identify the separately verified calibrated release when
finalization and upload succeed; its absence means no final release is recorded.
Validation audit on 17 September, 07:00 UTC:
| Selected candidate | Step | Validation accuracy | Crossfit selection NLL |
|---|---|---|---|
| Qwen3 4B, main run | 2,500 | 93.55% | 0.188640 |
| Qwen3 4B, lower learning rate | 4,000 | 93.75% | 0.190020 |
| Qwen3.5 2B | 2,000 | 89.84% | 0.255294 |
| Qwen3 4B, refinement | 1,500 | 94.73% | 0.170150 |
All use the same 512 validation decisions. Selection uses four source-group-
disjoint temperature-crossfit folds with fixed seed 431; smaller macro-family NLL
is better. This is validation-selection evidence, not a held-out quality claim.
The 2B run stopped cleanly at step 6,000 after eight evaluations without a new
best. At 07:07 UTC the original GX10 4B run stopped cleanly at step 4,380,
preserving its selected step 2,500 and full resumable state. GX10 now verifies an
expanded-data candidate from the refinement's frozen selected step 1,500, while
both Spark 4B runs continue. See source/EXPANDED_DATA.md for its latest run
state, verification evidence and the broader training mix. Independent trials
retain the original absolute deadlines.
Contents and reconstruction
The browser playground provides a context box, question, options list and probability bars, with a model selector. Its lightweight local server uses these custom checkpoints and the pinned base weights. It runs on GX10 over a loopback connection/SSH tunnel; this repository is a backup, not a hosted inference application.
source/: a committed source/documentation snapshot, including the training scripts, tests, Jev-compatible harness, findings and operational handover.snapshots/<id>/backup-manifest.json: artifact roles, originating hosts/paths, checkpoint steps, model/data/source provenance and SHA-256 checksums.snapshots/<id>/artifacts/<candidate>/best.pt: selected adapter/head weights and reconstruction metadata. These initial backups have no fitted temperature.snapshots/<id>/artifacts/<candidate>/checkpoint.pt: resumable training state, including Adam and random states. Its step may be later than the selected best.- Expanded-data snapshots include transformed decision files and their source notices, split hashes, diagnostic exclusions and original dataset lineage.
final/<campaign>/: final calibrated model and evaluation evidence, created only after successful finalization and publication.
Keep checkpoint.pt, its sibling best.pt and matching validation evidence
together when resuming. Checkpoint metadata records the original local paths;
reconstruction needs the pinned base/data and matching project source. The current
harness works on the recorded GX10/Spark layout. This backup does not establish
portability to a new operating system or a generic hosted inference service.
Model and optimizer files use PyTorch serialization; use the project's known,
hash-verified artifacts and its supplied reconstruction code.
The main 4B model uses FP32, rank-8 additive adapters (alpha 16) and a learned scalar head initialized from pretrained yes-minus-no logits: 16,517,633 trainable parameters. The 2B text decoder uses rank 16 / alpha 32, with 16,821,249 trainable parameters. Training limits are 512 and 768 complete-chat tokens respectively; 1,024-token inference was separately verified. Training retains a 16 GiB per-job CUDA allocation cap. BF16 did not pass the project's numerical invariance gate.
Pinned bases, recorded as Apache-2.0 in their provenance:
Qwen/Qwen3-4B-Instruct-2507atcdbee75f17c01a7cc42f958dc650907174af0554.Qwen/Qwen3.5-2Bat15852e8c16360a2fea060d615a32b45270f8a8fc.
Data, evaluation and limitations
The original public training sources are SNLI, BoolQ, ARC and four-choice Banking77 routing.
Social IQA is a wholly untrained task-family holdout. Frozen source pins, licences,
raw/split hashes and group/deduplication audits are included in
source/results/public-decisions-v1-manifest.json. That manifest credits
Stanford NLP (SNLI), Google (BoolQ), AllenAI (ARC/Social IQA) and PolyAI (Banking77),
and records their original dataset licences. The original repository's
license: unknown metadata is retained; no new licence for the project artifacts
is assigned by this backup.
The expanded candidate adds official training rows from HellaSwag (MIT), PIQA
(AFL-3.0 according to its creator's pinned README) and CommonsenseQA (MIT).
After the 512-token filter there are 80,765 training decisions, including all
40,915 original retained examples. All four original reserved split files and
their tokenized rows are unchanged. Another 383 retained new-source diagnostic
decisions are excluded from both training and checkpoint selection. Source pins,
credits, licences and transformations are in
source/results/20260917-expanded-data/dataset-manifest.json. The unchanged
selection set measures the original tasks; expansion alone is not evidence of
better performance on the three added tasks.
Validation selects checkpoints. Separate calibration fits one global temperature only after selection is frozen. Final test/holdout comparisons use the unchanged pretrained scorer, with a separately fitted base temperature and source-group uncertainty. Frozen grouping and exact deduplication do not rule out pretraining contamination or semantic duplicates. Four-choice Banking77 is not the full 77-label benchmark. No claim of general intelligence, Jev equivalence or calibration on unseen task families follows from the current validation scores.
The harness implements local scoring, a loopback API, hosted Jev requests and
response/timing comparisons. The local model identifies itself as OpenSysOne;
compatibility with the request shape does not make it Jev. Local confidence is
normalized entropy, not calibrated probability of correctness. Authenticated
hosted Jev calls require TYPESAFE_API_KEY and have not been tested. Credentials
and pretrained base weights are not part of this backup. Expanded-data snapshots
may include the transformed training data, with upstream notices retained;
raw upstream archives are referenced by pinned revision and checksum.