opensysone / source /README.md
andyshu's picture
Back up verified OpenSysOne training snapshot and pinned source
cc3f990 verified
|
Raw History Blame
3.74 kB
# OpenSysOne
Try the trained models in the [browser playground](PLAYGROUND.md): enter context,
a question and possible answers, then compare their probabilities across models.
Experimental decision scorer: pretrained Transformer, arbitrary natural-language
candidates and scalar logits, with a Jev-compatible API harness.
Read [HANDOVER.md](HANDOVER.md) to continue on GX10 and [PLAN.md](PLAN.md) for the
current plan. [FLEET_RUN.md](FLEET_RUN.md) records the three-machine allocation,
process controls and deadline. [RESULTS.md](RESULTS.md) records measured outcomes.
The original proposal is preserved in [RESEARCH_BRIEF.md](RESEARCH_BRIEF.md).
The active public-data experiment compares pinned posttrained Qwen3.5-2B and
Qwen3-4B-Instruct-2507. Ordinary low-rank decoder adapters and a scalar head train
in FP32; the head starts from pretrained yes-minus-no token logits. Frozen data
covers entailment, passage questions, science and banking intent, with Social IQA
reserved as an entirely unseen task family. Checkpoint selection uses validation
only; separate calibration and untouched evaluation follow training.
Use the existing isolated experiment environment on GX10:
```bash
cd ~/projects/opensysone
~/ai/envs/opensysone/bin/python -m unittest discover -s tests -v
~/ai/envs/opensysone/bin/python scripts/campaign_status.py
```
Artifacts, including optimizer/RNG state, are written to `~/ai/opensysone/runs/`.
Pretrained weights are under `~/ai/models/opensysone`; frozen public data is under
`~/ai/opensysone/data`. The isolated environment reuses torch/Transformers read-only
and adds pyarrow 25.0.1. The shared Python environment is not modified. Real-model
checkpoint reload, longest-input training gradients and authenticated loopback
HTTP are checked before the 24-hour campaign. The campaign checkpoints on a
cadence independent of evaluation, reserves at least two hours for final evaluation and
automatically starts the calibrated local API. The fleet runs independent 4B and
2B candidates, selects on the same 512 validation decisions, and evaluates the
selected artifact after training. Exact paths and state are in the
handover. These experiments do not establish Jev-level intelligence.
[JEV_HARNESS.md](JEV_HARNESS.md) documents local inference, hosted Jev calls,
response/timing comparisons, bearer authentication and Mac SSH tunneling. Hosted
calls use `TYPESAFE_API_KEY`; local inference does not make remote calls.
The original pinned Qwen2.5-0.5B synthetic smoke, partial tuning and verified
shared-prefix caching remain available in `smoke.py` and `decision_model.py`:
```bash
bash scripts/run_smoke.sh
```
The new scorers use bounded full forwards; cache branching has not been verified
for them. CPU tests cover adapters, categorical gradients, restoration, API shapes
and the old model's cache correctness.
Investigate the preserved BF16 checkpoint without changing its weights:
```bash
free -b
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv
bash scripts/run_precision.sh --run ~/ai/opensysone/runs/20260916T154714Z
cat ~/ai/opensysone/runs/LAST_PRECISION_RUN
```
The precision wrapper uses the smoke's lock, 25-minute timeout and 16 GiB CUDA
allocation cap. Each fresh run stores provenance and raw comparisons under
`artifacts/`, plus `run.log` and `exit_code`. Exit 0 means the diagnosis completed;
read the numerical comparisons to determine whether a precision passes.
See [NEXT_STEPS.md](NEXT_STEPS.md) for the September 17 findings, active refinement
and assessment of joint work on the two Sparks.
Current fleet status: `python3 scripts/fleet_status.py` (or `--json`).
Hugging Face backup and final-publication controls: [HUGGINGFACE.md](HUGGINGFACE.md).