opensysone / source /README.md
andyshu's picture
Back up verified OpenSysOne training snapshot and pinned source
cc3f990 verified
|
Raw History Blame
3.74 kB

OpenSysOne

Try the trained models in the browser playground: enter context, a question and possible answers, then compare their probabilities across models.

Experimental decision scorer: pretrained Transformer, arbitrary natural-language candidates and scalar logits, with a Jev-compatible API harness.

Read HANDOVER.md to continue on GX10 and PLAN.md for the current plan. FLEET_RUN.md records the three-machine allocation, process controls and deadline. RESULTS.md records measured outcomes. The original proposal is preserved in RESEARCH_BRIEF.md.

The active public-data experiment compares pinned posttrained Qwen3.5-2B and Qwen3-4B-Instruct-2507. Ordinary low-rank decoder adapters and a scalar head train in FP32; the head starts from pretrained yes-minus-no token logits. Frozen data covers entailment, passage questions, science and banking intent, with Social IQA reserved as an entirely unseen task family. Checkpoint selection uses validation only; separate calibration and untouched evaluation follow training.

Use the existing isolated experiment environment on GX10:

cd ~/projects/opensysone
~/ai/envs/opensysone/bin/python -m unittest discover -s tests -v
~/ai/envs/opensysone/bin/python scripts/campaign_status.py

Artifacts, including optimizer/RNG state, are written to ~/ai/opensysone/runs/. Pretrained weights are under ~/ai/models/opensysone; frozen public data is under ~/ai/opensysone/data. The isolated environment reuses torch/Transformers read-only and adds pyarrow 25.0.1. The shared Python environment is not modified. Real-model checkpoint reload, longest-input training gradients and authenticated loopback HTTP are checked before the 24-hour campaign. The campaign checkpoints on a cadence independent of evaluation, reserves at least two hours for final evaluation and automatically starts the calibrated local API. The fleet runs independent 4B and 2B candidates, selects on the same 512 validation decisions, and evaluates the selected artifact after training. Exact paths and state are in the handover. These experiments do not establish Jev-level intelligence.

JEV_HARNESS.md documents local inference, hosted Jev calls, response/timing comparisons, bearer authentication and Mac SSH tunneling. Hosted calls use TYPESAFE_API_KEY; local inference does not make remote calls.

The original pinned Qwen2.5-0.5B synthetic smoke, partial tuning and verified shared-prefix caching remain available in smoke.py and decision_model.py:

bash scripts/run_smoke.sh

The new scorers use bounded full forwards; cache branching has not been verified for them. CPU tests cover adapters, categorical gradients, restoration, API shapes and the old model's cache correctness.

Investigate the preserved BF16 checkpoint without changing its weights:

free -b
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv
bash scripts/run_precision.sh --run ~/ai/opensysone/runs/20260916T154714Z
cat ~/ai/opensysone/runs/LAST_PRECISION_RUN

The precision wrapper uses the smoke's lock, 25-minute timeout and 16 GiB CUDA allocation cap. Each fresh run stores provenance and raw comparisons under artifacts/, plus run.log and exit_code. Exit 0 means the diagnosis completed; read the numerical comparisons to determine whether a precision passes.

See NEXT_STEPS.md for the September 17 findings, active refinement and assessment of joint work on the two Sparks.

Current fleet status: python3 scripts/fleet_status.py (or --json).

Hugging Face backup and final-publication controls: HUGGINGFACE.md.