File size: 3,842 Bytes
92b5c9a
 
a2af77f
 
92b5c9a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
# Fine-tune on your data

Already using System One requests? [Convert the same state/questions/criteria into training rows](SYSTEM_ONE.md#fine-tune-with-the-same-inputs) with separate hard or soft targets and shared source-component IDs.

The included CLI trains the existing Choice, Noul and Score paths without adding new parameters. It supports hard or soft labels, deterministic scheduling, full-input admission, DEV checkpoint selection and optimizer/RNG resume. Token embedding, embedding normalization and type embedding remain frozen. The other 486 parameter tensors are trainable when their type is present.

Provide your own TRAIN and DEV JSONL. The CLI rejects shared IDs, shared normalized full inputs and same-source components where supplied. It does not prove semantic independence; use an appropriate development split and do not train on a release/test set.

| Type | Target |
|---|---|
| Choice | `{"choice_id":"candidate-id"}` or `{"probabilities":[0.2,0.8]}` |
| Noul | `{"probability":0.8}`; hard labels are 0 or 1 |
| Score | `{"probabilities":[0.1,0.2,0.7]}` in increasing-value level order; hard labels are one-hot |

Every row retains `id`, full `state_text` and a complete typed `question`. Probability vectors sum to one. A scalar Score is not silently converted into a distribution.

```bash
TRAIN=/data/train.jsonl DEV=/data/dev.jsonl OUTPUT=/runs/my-decision \
ROCR_VISIBLE_DEVICES=0 bash examples/finetune.sh
```

`examples/finetune.sh` is the editable example configuration; it uses actual supported flags, not a nonexistent `--config` option. Its four-epoch recipe uses logical batch 64 / microbatch 8, BF16 autocast with FP32 parameters/loss, AdamW, encoder LR 2.5e-5 / head LR 1e-4, clip 1, CE plus 0.1 ordinal RPS for Score, and a cosine floor of 1e-6 after 10% warmup. Logical tails use their actual row count. These are example defaults, not a guarantee of improvement.

The default selector minimizes equal-type macro soft NLL at step 0 and epoch ends. The separate `--selection hard-accuracy` policy requires an explicit `hard_target_id` on every DEV row, uses row hard accuracy then soft NLL then the earlier step, and never derives hard gold from soft targets. Noul hard selection uses `p_yes >= .5`. Record the selector and cadence used in any model comparison; these defaults are starting points rather than a guarantee of task quality.

Successful runs export the selected native bundle and report its manifest in `COMPLETE.json`. Selecting step 0 means no adopted adaptation. Immutable checkpoints include model, AdamW and RNG state; allow substantial disk space. For an interrupted run, use exactly the same original arguments and output plus `--resume <checkpoint> --resume-sha256 <hash>`. Read `LATEST.json` for the last completely saved boundary. `--stop-after-step N` is an optional clean pause; it does not restart or shorten the LR horizon.

The included source underwent a bounded AMD continuous-versus-resume check and independent native reload. That interface evidence is separate from dataset quality or long training stability. Validate the quality of your selected model on an appropriately isolated evaluation set; no automatic release or adoption is performed.

## Lex checkpoint provenance

The editable example above is a downstream fine-tuning recipe, not the exact Lex release recipe. Lex was refit from Kai on all 6,000 original typed-decisions TRAIN decisions using eight epochs, fixed final step 750 and blended original hard/soft cross-entropy. It used no final-refit DEV selector, no warmup and no RPS term. See [METHODS.md](METHODS.md) and [TRAINING_PROVENANCE.json](TRAINING_PROVENANCE.json). The public CLI does not provide a named `blend50` switch; explicitly form normalized blended targets if using that objective. Reusing Lex as an initializer is a new experiment.