Download FINETUNING.md from vllm-sr/Decision-1.0-Lex-0.6B: direct link, hf CLI and curl.
- Browser
- Download file 3.84 kB
-
https://huggingface.co/vllm-sr/Decision-1.0-Lex-0.6B/resolve/e4a96e95ad08867f7114e1a3ed4d42dca4e2723e/FINETUNING.md
- Command line
-
hf download hf://vllm-sr/Decision-1.0-Lex-0.6B@e4a96e95ad08867f7114e1a3ed4d42dca4e2723e/FINETUNING.md
-
curl -L -o FINETUNING.md https://huggingface.co/vllm-sr/Decision-1.0-Lex-0.6B/resolve/e4a96e95ad08867f7114e1a3ed4d42dca4e2723e/FINETUNING.md
Fine-tune on your data
Already using System One requests? Convert the same state/questions/criteria into training rows with separate hard or soft targets and shared source-component IDs.
The included CLI trains the existing Choice, Noul and Score paths without adding new parameters. It supports hard or soft labels, deterministic scheduling, full-input admission, DEV checkpoint selection and optimizer/RNG resume. Token embedding, embedding normalization and type embedding remain frozen. The other 486 parameter tensors are trainable when their type is present.
Provide your own TRAIN and DEV JSONL. The CLI rejects shared IDs, shared normalized full inputs and same-source components where supplied. It does not prove semantic independence; use an appropriate development split and do not train on a release/test set.
| Type | Target |
|---|---|
| Choice | {"choice_id":"candidate-id"} or {"probabilities":[0.2,0.8]} |
| Noul | {"probability":0.8}; hard labels are 0 or 1 |
| Score | {"probabilities":[0.1,0.2,0.7]} in increasing-value level order; hard labels are one-hot |
Every row retains id, full state_text and a complete typed question. Probability vectors sum to one. A scalar Score is not silently converted into a distribution.
TRAIN=/data/train.jsonl DEV=/data/dev.jsonl OUTPUT=/runs/my-decision \
ROCR_VISIBLE_DEVICES=0 bash examples/finetune.sh
examples/finetune.sh is the editable example configuration; it uses actual supported flags, not a nonexistent --config option. Its four-epoch recipe uses logical batch 64 / microbatch 8, BF16 autocast with FP32 parameters/loss, AdamW, encoder LR 2.5e-5 / head LR 1e-4, clip 1, CE plus 0.1 ordinal RPS for Score, and a cosine floor of 1e-6 after 10% warmup. Logical tails use their actual row count. These are example defaults, not a guarantee of improvement.
The default selector minimizes equal-type macro soft NLL at step 0 and epoch ends. The separate --selection hard-accuracy policy requires an explicit hard_target_id on every DEV row, uses row hard accuracy then soft NLL then the earlier step, and never derives hard gold from soft targets. Noul hard selection uses p_yes >= .5. Record the selector and cadence used in any model comparison; these defaults are starting points rather than a guarantee of task quality.
Successful runs export the selected native bundle and report its manifest in COMPLETE.json. Selecting step 0 means no adopted adaptation. Immutable checkpoints include model, AdamW and RNG state; allow substantial disk space. For an interrupted run, use exactly the same original arguments and output plus --resume <checkpoint> --resume-sha256 <hash>. Read LATEST.json for the last completely saved boundary. --stop-after-step N is an optional clean pause; it does not restart or shorten the LR horizon.
The included source underwent a bounded AMD continuous-versus-resume check and independent native reload. That interface evidence is separate from dataset quality or long training stability. Validate the quality of your selected model on an appropriately isolated evaluation set; no automatic release or adoption is performed.
Lex checkpoint provenance
The editable example above is a downstream fine-tuning recipe, not the exact Lex release recipe. Lex was refit from Kai on all 6,000 original typed-decisions TRAIN decisions using eight epochs, fixed final step 750 and blended original hard/soft cross-entropy. It used no final-refit DEV selector, no warmup and no RPS term. See METHODS.md and TRAINING_PROVENANCE.json. The public CLI does not provide a named blend50 switch; explicitly form normalized blended targets if using that objective. Reusing Lex as an initializer is a new experiment.