APUS-OpenJev-v1 / 4B-5949 /training.md
gump2049's picture
Publish complete APUS-OpenJev-v1 models and Technical Report v1.1
68e5880 verified
|
Raw History Blame Contribute Delete
2.69 kB

This release: 4B, step 5949

This is the completed 5949-step SFT endpoint.

Training provenance

Both families use the registered 5949-record SFT curriculum (3898 parent groups), with 5949 maximum steps and one epoch. Each checkpoint's recorded training step, epoch, and original trainer-state hash are preserved in the separate LoRA archive manifest at fixed revision 1e5557f923746031f8187b7daf299b9bee41cb3c (repository access required). This merged package's release-manifest.json inventories inference artifacts and does not contain that per-checkpoint trainer-state record; intermediate checkpoints did not finish the whole schedule. Randomized loader order means the step number alone is not a verified count of unique examples seen at an intermediate checkpoint.

Source Scheduled records Parent groups
Mind2Web browser Choice 1798 671
Mind2Web browser TYPE 158 125
HelpSteer3 principle 1300 1300
SGD 513 10
GoEmotions independent-attribute Score 500 297
BoolQ 600 600
MNLI 1000 1000
Local counterfactual 80 20

Parent groups can overlap across browser Choice/TYPE. The 4B run initialized from a 427-record pilot adapter (weight SHA256 0047e5f1f0c98f17da93def94032041997592e1a609ddcab58de41f4d30e3a38), while the inspected 9B config records no origin adapter. Do not add pilot records to the registered schedule as if all were independent.

Decision objective: 0.5 CE(low) + 0.5 CE(high) + 0.1 KL(P_high.detach || P_low). TYPE examples use full-depth text cross-entropy. LoRA r=8, alpha=16, dropout=0; learning rate 1e-4; batch size 1; seed 20260920; training max length 6144. The recorded compiled schedule maximum is 5845 tokens. Training settings and data/schedule hashes are retained in the separate LoRA checkpoint depth configuration at fixed archive revision 1e5557f923746031f8187b7daf299b9bee41cb3c. The merged package's depth_config.json contains only portable inference and identity metadata; it is not the full training configuration.

This is SFT with a within-model distillation term; it does not prove RLCD, online reinforcement learning or teacher-model OPD occurred. The declared public sources contain multiple licensing regimes (including CC-BY, CC-BY-SA and mixed-source material); separate provenance and redistribution review remains necessary. Raw datasets are not part of this upload.