{ "experiment": "auto-200m-2", "authorization": "User requested auto-200m-2: a ModernBERT-base student trained on the RTX PRO 6000 in a temporary job on the datasets used for auto-0.4b-2, with 64k context, aiming for accuracy close to auto at a smaller size; publish publicly as ProCreations/auto-200m-2.", "initialization": "answerdotai/ModernBERT-base (revision 8949b909), sequence classification with CLS pooling, global-attention RoPE theta chosen by a masked-LM sweep on training text, max_position_embeddings 65536.", "data": "ProCreations/auto-1b-data d265bbf7 cleaned exactly as auto-0.4b-2 (711,985 rows, 516.7M tokens) plus the 120,000 archived cross-paired rows labelled only by the Auto 3B teacher; validation partitions identical to auto-0.4b-2; benchmark ProCreations/approve-or-deny a38b6259 never used for training, pilots or checkpoint selection.", "loss": "(1-alpha)*CE(label) + alpha*T^2*KL(teacher/T || student/T); teacher = cached Auto 3B logits (train rows and augmented rows); augmented rows use the teacher argmax as label.", "optimizer": "Full-parameter fused AdamW, FP32 master weights, BF16 autocast, betas (0.9, 0.95), weight decay 0.01, grad clip 1.0; 3% warmup, cosine to 10% of peak.", "batching": "Length-bucketed microbatches up to 65,536 padded tokens / 128 rows; optimizer step at >=128 rows or >=131,072 tokens; every row at full length (no truncation).", "pilot": {"rows": 120000, "lrs": [3e-05, 6e-05, 0.0001], "epochs": 1, "alpha": 0.5, "temperature": 2.0, "criterion": "lowest NLL on the 7,824-row validation selection split restricted to inputs <= 4,096 tokens"}, "runs": [ {"name": "p1_short", "start_from": "init", "indices": "short_all", "epochs": 3, "lr": "pilot", "alpha": 0.5, "temperature": 2.0, "evals_per_epoch": 3}, {"name": "p2_long", "start_from": "final:p1_short", "indices": "long_mix", "short_replay": 1.0, "aug_replay": 0.25, "epochs": 2, "lr": "pilot*0.4", "alpha": 0.5, "temperature": 2.0, "evals_per_epoch": 3} ], "selection": "Candidates = p2_long exports and uniform weight averages of its last 2/3/4 exports. Ranked on the 7,824-row validation selection split by NLL; the single top candidate is frozen before any benchmark evaluation. Audit partition and benchmark are evaluated once for the frozen candidate.", "selection_amendment": "2026-09-26 02:45 UTC, during p2 and before any candidate benchmark or audit evaluation: p2 raised validation NLL on the (mostly short) selection split (p1 final 0.0720; p2 step 537 0.0815, step 1074 0.0899), so the candidate pool also includes the p1_short final export and p1-final/p2-final weight averages (0.3/0.5/0.7). Same criterion: lowest validation NLL, one frozen candidate.", "second_look": "2026-09-26 04:12 UTC, after the frozen candidate (p1_short-step-16058, never trained on >4,096-token rows) scored 2,890/3,000 (53 FA / 57 FD): at the maintainer's request, ONE long-trained candidate, p2_long-step-2685 (best long-trained candidate by validation NLL, 0.0726), is also benchmarked. Decision rule fixed before its result: publish it only if its benchmark accuracy is higher; if equal, the one with fewer false approvals; otherwise publish the frozen candidate. Both results are disclosed.", "reference": "ProCreations/auto-0.4b-2 revision 5937dd01 re-evaluated with the same code; must reproduce 2,910/3,000 (36 false approvals / 54 false denials)." }