jaredpalmer commited on
Commit
957b91e
·
verified ·
1 Parent(s): 485ace8

Kev-4B: real-document delta (round 8; documents-v1 locked test 0.804 -> 0.904)

Browse files
Files changed (8) hide show
  1. README.md +49 -13
  2. adapter_model.safetensors +1 -1
  3. head.pt +2 -2
  4. provenance.json +46 -29
  5. result.json +0 -0
  6. train.log +94 -46
  7. training_config.json +15 -8
  8. training_metrics.json +8 -8
README.md CHANGED
@@ -30,6 +30,11 @@ metrics:
30
  model-index:
31
  - name: Kev-4B
32
  results:
 
 
 
 
 
33
  - task: { type: text-classification, name: typed decision (choice / noul / score) }
34
  dataset: { type: mixed, name: "decision-v7 development (1,204 records; ten trained public sources + programmatic policy data)" }
35
  metrics:
@@ -51,20 +56,51 @@ model-index:
51
 
52
  Kev-4B is a **decision model**: one document (the *state*) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation. It is a LoRA adapter (r=16, 33.8M trainable parameters) plus a pointer head on `Qwen/Qwen3.5-4B-Base` (revision `1001bb4d`), serving TypeSafe's public `/v1/systemone` contract.
53
 
54
- **The recommended Kev.** The best accuracy per byte: out of domain 0.797 on the development partition and **0.837 on the locked test**, Brier 0.255 on the test, held-out rule pairs 0.77–0.78. This checkpoint is the `decision-v7` recipe (trial `q35-4b-s23/00-trial-0`, seed 2, selected on development accuracy) followed by a 9-minute **delta fine-tune** (`--init_from`, lr 2e-5, one epoch) on 1,425 additional records — date-bearing policy cases rendered with explicit day counts, and evidence-free cases with uniform targets — mixed with 2,000 replayed training records. Against the pre-delta checkpoint on the locked test: +1.0 pp [−0.1, +2.1], Brier 0.266 → 0.255, `deadline` 0.65 → 0.75.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
55
 
56
- - Hub: `jaredpalmer/kev-4b` (this repo; trial `night2-4b-du/00-trial-0`). The pre-delta checkpoint is at revision `v7-base`; the Qwen3 generation at `qwen3` ([its card](kev-4b-qwen3.md)).
57
- - Demo: [huggingface.co/spaces/jaredpalmer/kev](https://huggingface.co/spaces/jaredpalmer/kev) runs Kev-4B and Kev-0.8B on ZeroGPU with the same encoder and API code as `kev.serve`.
58
- - Code, suites, every trial with hashes and paired bootstraps: [github.com/jaredpalmer/kev](https://github.com/jaredpalmer/kev) — `PLAN_Qwen35.md` (the port and this experiment), `PLAN.md`, `runs/leaderboard.md`
59
 
60
- ## Results (same frozen items for every row)
61
 
62
  | | Kev-4B (Qwen3) | Kev-4B before the delta (`v7-base`) | **Kev-4B, raw logits** | **Kev-4B as served (T = 2.14)** | Jev |
63
  |---|---|---|---|---|---|
64
  | in-distribution accuracy (decision-v7 dev, 1,204 records) | 0.854 | 0.877 | 0.872 | 0.872 | 0.845 |
65
  | out-of-domain accuracy (transfer-v4 dev, 764 records) | 0.790 | 0.794 | **0.797** | 0.797 | 0.857 |
66
  | out-of-domain Brier | 0.328 | 0.316 | 0.299 | **0.264** | 0.211 |
67
- | out-of-domain ECE | 0.102 | 0.130 | 0.122 | **0.041** | 0.049 |
68
  | confident errors out of domain (p ≥ 0.9 and wrong) | 8.2% | 8.2% | 6.9% | **2.6%** | 3.7% |
69
  | coverage at ≤ 5% error (share of decisions automatable) | 0.31 | 0.54 | 0.54 | 0.57 | 0.70 |
70
  | held-out policy structures, both siblings correct | 0.73 | 0.78 | 0.78 | 0.78 | 0.86 |
@@ -80,18 +116,18 @@ Per-source out-of-domain accuracy (Kev-4B / Jev): QNLI 0.91 / 0.93, SciQ 0.97 /
80
 
81
  **What the delta cost.** MMLU-Pro fell 0.500 → 0.490 and scienthoon's ECE rose 0.086 → 0.116; coverage at ≤ 5% error was unchanged (0.54 development, 0.67 → 0.68 locked test) and confident errors fell (8.2% → 6.9%). The pre-registered criteria for the delta (`PLAN.md`, "Tonight's autoresearch") were met for dates and for the unknowable-confidence behaviour; the coverage criterion asked for +5 pp and got 0; the locked read decided promotion.
82
 
83
- **Newer evaluation columns** (`transfer-v9` development, Kev-4B / Jev): MMLU-Pro (10-way) 0.490 / 0.840; state buried among unrelated records 0.68 / 0.70; unknowable share at ≥ 0.9 confidence 0.00 / 0.09 (intact controls 0.94).
84
 
85
- **External suites** (same items as their published Jev numbers): SemIf's authored 144 — 0.896 before the delta (live Jev 0.965; SemIf's untrained Qwen3.5-4B 0.813); scienthoon's 900 tickets — queue 0.918, angry 0.790, ECE 0.116 (Jev 0.897, 0.914, 0.105). On ekzhang's 1,000-question MMLU-Pro sample the pre-delta checkpoint scores 0.468 (Jev 0.829).
86
 
87
- ## How it was built
88
 
89
  - **Base model**: Qwen3.5-4B-Base, a hybrid of 24 Gated DeltaNet (linear attention) layers and 8 full-attention layers. Because the recurrent layers cannot honour a block-causal mask, questions run as separate causal rows that continue from the shared state (`kev/model.py: forward_rows_batch`); isolation is exact by construction (together vs alone within 1e-5) and on attention-only models this form is bit-identical to the packed one.
90
- - **Recipe**: `decision-v7`, two epochs, LoRA r=16 (attention, MLP and DeltaNet projections), lr 5e-5 — the same data and settings as every other Kev, so the Qwen3 → Qwen3.5 difference is the base (`PLAN_Qwen35.md` §10: locked test +7.3 pp [+2.8, +11.7] over Kev-8B).
91
  - **Delta**: `kev.train --init_from jaredpalmer/kev-4b@v7-base --data evals/night2/dates_unknowable.jsonl --replay 2000 --lr 2e-5 --epochs 1`. The 1,425 new records are generated (no public dataset): 900 date-bearing policy cases, a third rendered plainly, a third with a relational day-count sentence, a third with a `date_facts` field; 255 cases with the deciding sentence removed and a uniform soft target over the options, plus their 270 intact controls. Record hashes are in `evals/night2/manifest.json`; the source checkpoint's hashes are in `training_config.json`.
92
  - Why a delta and not a retrain: it is a controlled change (one fixed checkpoint, one data addition, 9 minutes), and the results section shows exactly what it moved.
93
 
94
- ## Known limits
95
 
96
  - Use [Kev-9B](kev-9b.md) when accuracy and calibration matter more than memory: 0.852 vs 0.837 out of domain on the locked test, Brier 0.237 vs 0.255.
97
 
@@ -102,11 +138,11 @@ Per-source out-of-domain accuracy (Kev-4B / Jev): QNLI 0.91 / 0.93, SciQ 0.97 /
102
  - The raw logits are over-confident out of domain; the built-in temperature (T = 2.14) fixes most of it without changing any answer. `KEV_TEMPERATURE=1.0` gives the raw values. Coverage at a 5% error budget is 0.54–0.68 against Jev's 0.70.
103
  - 4B bf16 needs ~9 GB of GPU memory for serving; training took 56 min on one H100 (peak 24.6 GB).
104
 
105
- ## Training
106
 
107
  Frozen suite `evals/v7/decision-v7`: 10,000 public records (1,000 per source), 896 policy minimal-pair records over nine template families, 1,680 records from 60 randomly generated rule structures in four rendering styles. Two epochs, LoRA r=16 α=32 on `q/k/v/o_proj`, `gate/up/down_proj`, `in_proj_qkv/z/a/b`, `out_proj`; pointer head from scratch; cross-entropy on the option distribution; lr 5e-5 (OneCycle), effective batch 8, bf16 autocast with fp32 master weights, gradient checkpointing; option permutation, none-of-the-above insertion, distractors, none minimal pairs on 25% of Choice records. Then the delta described above (one epoch, lr 2e-5, 3,937 records seen, 9 minutes on one H100). No Jev outputs were used for training.
108
 
109
- ## Evaluation protocol
110
 
111
  Development partitions select models; the locked test partition is read at most once per candidate (`runs/locked/kev-4b-night2-du-ungated/`; the pre-delta read is `runs/locked/kev-4b-q35/`). Every number carries suite hash, code hashes and git commit in `result.json`. Untrained-base baselines use zero-shot letter logits on the same items (`scripts/base_mmlu_probe.py`).
112
 
 
30
  model-index:
31
  - name: Kev-4B
32
  results:
33
+ - task: { type: text-classification, name: typed decision, real documents, locked test }
34
+ dataset: { type: mixed, name: "documents-v1 test (936 questions on CFPB complaint narratives; read once)" }
35
+ metrics:
36
+ - { type: accuracy, value: 0.904 }
37
+ - { type: brier_score, value: 0.156 }
38
  - task: { type: text-classification, name: typed decision (choice / noul / score) }
39
  dataset: { type: mixed, name: "decision-v7 development (1,204 records; ten trained public sources + programmatic policy data)" }
40
  metrics:
 
56
 
57
  Kev-4B is a **decision model**: one document (the *state*) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation. It is a LoRA adapter (r=16, 33.8M trainable parameters) plus a pointer head on `Qwen/Qwen3.5-4B-Base` (revision `1001bb4d`), serving TypeSafe's public `/v1/systemone` contract.
58
 
59
+ **This version (2026-09-24): real-document delta.** The previous Kev-4B plus one epoch (lr 2e-5) on `documents-v1` train: 5,219 real US consumer-finance complaint narratives (CFPB, 2015-2024, up to ~7k tokens) with 7,488 questions (which product, which main issue), labels kept only where two open-weight teachers agreed with the consumer's own filing, mixed with 2,000 replayed `decision-v7` records. On complaint narratives it has never seen, accuracy goes from 0.804 to **0.904** on the locked test (+9.9 pp [+7.5, +12.4], 936 questions) and from 0.811 to **0.891** on a private held-out set (`documents-v2`, 953 questions); on the development split it scores 0.895 against Jev's 0.868. Everything else is unchanged within noise: locked out-of-domain test 0.835 (previous 0.837), served Brier 0.233 (0.232).
60
+
61
+ **Read this before relying on the documents numbers.** The gain is measured **in distribution**: training and every documents suite share one source (CFPB complaints) and the same two question templates. It shows Kev-4B learns real long documents from a few thousand labelled examples; it does not show the same gain on other kinds of documents. Evaluation labels are AI-adjudicated (a unanimous three-model judge panel, or two agreeing adjudications) and human spot-checked (47/50 and 50/50).
62
+
63
+ - Hub: `jaredpalmer/kev-4b` (this repo; trial `r8-small/00-trial-0`, `PLAN.md` round 8). The previous version is at tag `night2-du-release`; the pre-delta v7 checkpoint at `v7-base`; the Qwen3 generation at `qwen3` ([its card](kev-4b-qwen3.md)).
64
+ - Code, suites, every trial with hashes and paired bootstraps: [github.com/jaredpalmer/kev](https://github.com/jaredpalmer/kev). The numbers below are in `runs/release/kev-4b-r8.json`.
65
+
66
+ ## Results (as served: each checkpoint at its own fitted temperature)
67
+
68
+ | | **Kev-4B (this version, T = 2.96)** | previous Kev-4B (T = 2.14) | Jev |
69
+ |---|---|---|---|
70
+ | **real documents**, locked test (`documents-v1`, 936 questions) | **0.904** | 0.804 | – |
71
+ | real documents, private held-out (`documents-v2`, 953) | **0.891** | 0.811 | – |
72
+ | real documents, development (920) | **0.895** | 0.811 | 0.868 |
73
+ | real documents, Brier (locked test) | **0.156** | 0.286 | – |
74
+ | in-distribution accuracy (decision-v7 dev, 1,264 questions) | 0.873 | 0.872 | 0.845 |
75
+ | out-of-domain accuracy (transfer-v4 dev) | 0.802 | 0.797 | 0.857 |
76
+ | out-of-domain Brier / ECE | 0.265 / 0.043 | 0.264 / 0.040 | 0.211 / 0.049 |
77
+ | confident errors out of domain (p ≥ 0.9 and wrong) | 2.9% | 2.6% | 3.7% |
78
+ | coverage at ≤ 5% error | 0.552 | 0.573 | 0.70 |
79
+ | held-out policy structures, both siblings correct | 0.781 | 0.781 | 0.86 |
80
+ | unknowable items answered at ≥ 0.9 (transfer-v9) | 0.00 | 0.00 | 0.09 |
81
+ | MMLU-Pro (transfer-v9 dev, 10-way) | 0.515 | 0.490 | 0.840 |
82
+ | **locked test**, out-of-domain accuracy / Brier | **0.835 / 0.233** | 0.837 / 0.232 | – |
83
+ | **locked test**, in-distribution accuracy | 0.875 | 0.871 | – |
84
+ | SemIf (144 authored decisions) | 0.882 | 0.889 | – |
85
+ | scienthoon (873 support tickets) | 0.723 | 0.696 | – |
86
+ | WANLI-v2 (1,002 NLI pairs) | 0.691 | 0.699 | – |
87
+ | TypeSafe (89 answered rows) | 0.843 | 0.843 | – |
88
+
89
+ Paired against the previous version (record-clustered bootstrap, 95 %): documents dev +8.4 pp [+6.0, +10.6], locked test +9.9 [+7.5, +12.4], private held-out +8.0 [+5.6, +10.3]; scienthoon +2.6 [+0.8, +4.4]; SemIf −0.7 [−2.8, +1.4]; WANLI-v2 −0.8 [−2.0, +0.4]; TypeSafe identical on all 89 rows. The release was decided by a rule registered before any read (`PLAN.md`, round 8): documents development lower bound > 0, short-state and external guards sized to what each suite can resolve, then one read of the untouched documents test and one locked read. A first seed (round 7) gave the same documents gain and failed only per-suite lower bounds on the two smallest suites; it was not released.
90
+
91
+ **Calibration.** The delta sharpened the raw logits (fitted temperature 2.14 → 2.96; raw out-of-domain Brier 0.327, raw locked Brier 0.278). As served, calibration is unchanged. `KEV_TEMPERATURE=1.0` gives the raw values.
92
+
93
+ ## Previous version: `night2-du` (2026-09-21), kept at tag `night2-du-release`
94
 
95
+ **At its release, the recommended Kev.** The best accuracy per byte: out of domain 0.797 on the development partition and **0.837 on the locked test**, Brier 0.255 on the test, held-out rule pairs 0.77–0.78. This checkpoint is the `decision-v7` recipe (trial `q35-4b-s23/00-trial-0`, seed 2, selected on development accuracy) followed by a 9-minute **delta fine-tune** (`--init_from`, lr 2e-5, one epoch) on 1,425 additional records — date-bearing policy cases rendered with explicit day counts, and evidence-free cases with uniform targets — mixed with 2,000 replayed training records. Against the pre-delta checkpoint on the locked test: +1.0 pp [−0.1, +2.1], Brier 0.266 → 0.255, `deadline` 0.65 → 0.75.
 
 
96
 
 
97
 
98
  | | Kev-4B (Qwen3) | Kev-4B before the delta (`v7-base`) | **Kev-4B, raw logits** | **Kev-4B as served (T = 2.14)** | Jev |
99
  |---|---|---|---|---|---|
100
  | in-distribution accuracy (decision-v7 dev, 1,204 records) | 0.854 | 0.877 | 0.872 | 0.872 | 0.845 |
101
  | out-of-domain accuracy (transfer-v4 dev, 764 records) | 0.790 | 0.794 | **0.797** | 0.797 | 0.857 |
102
  | out-of-domain Brier | 0.328 | 0.316 | 0.299 | **0.264** | 0.211 |
103
+ | out-of-domain ECE | 0.102 | 0.130 | 0.122 | **0.040** | 0.049 |
104
  | confident errors out of domain (p ≥ 0.9 and wrong) | 8.2% | 8.2% | 6.9% | **2.6%** | 3.7% |
105
  | coverage at ≤ 5% error (share of decisions automatable) | 0.31 | 0.54 | 0.54 | 0.57 | 0.70 |
106
  | held-out policy structures, both siblings correct | 0.73 | 0.78 | 0.78 | 0.78 | 0.86 |
 
116
 
117
  **What the delta cost.** MMLU-Pro fell 0.500 → 0.490 and scienthoon's ECE rose 0.086 → 0.116; coverage at ≤ 5% error was unchanged (0.54 development, 0.67 → 0.68 locked test) and confident errors fell (8.2% → 6.9%). The pre-registered criteria for the delta (`PLAN.md`, "Tonight's autoresearch") were met for dates and for the unknowable-confidence behaviour; the coverage criterion asked for +5 pp and got 0; the locked read decided promotion.
118
 
119
+ **Newer evaluation columns** (`transfer-v9` development, Kev-4B / Jev): MMLU-Pro (10-way) 0.490 / 0.840; state buried among unrelated records 0.67 / 0.70; unknowable share at ≥ 0.9 confidence 0.00 / 0.09 (intact controls 0.94).
120
 
121
+ **External suites** (same items as their published Jev numbers): SemIf's authored 144 — 0.896 before the delta (live Jev 0.965; SemIf's untrained Qwen3.5-4B 0.813); scienthoon's 900 tickets — queue 0.918, angry 0.790, ECE 0.116 (Jev 0.897, 0.914, 0.105). On ekzhang's 1,000-question MMLU-Pro sample the shipped checkpoint scores 0.468 over all 1,000 questions (8 exceed the state limit and count as wrong; live Jev 0.835 on the same items, ekzhang reports 0.829). On SemIf's pinned third-party selections (`evals/external/{wanli,typesafe}-v1`): WANLI-256 accuracy 0.695 (live Jev 0.758); TypeSafe-102 equal-case agreement / total-variation distance 0.856 / 0.231 over the 89 rows within the 8,192-token serving context (13 rejected), 0.770 / 0.308 over all 102 with rejected rows scored as wrong (live Jev 0.891 / 0.125; published TypeSafe answers 0.883 / 0.127); plain accuracy on the answered rows 0.843, coverage at <= 5% error 0.02 (Jev 0.892, 0.84). The shipped temperature is fitted in distribution and does not transfer to every workload. On WANLI, a single temperature fitted on the workload's own labelled rows (`python -m kev.calibrate`, group-disjoint out-of-fold) lowers ECE from 0.166 as shipped to 0.052 (workload T 3.91 against the shipped 2.14). Accuracy is unchanged and coverage at <= 5% error does not improve. On TypeSafe the shipped temperature already fits and refitting does not help (ECE 0.158 as shipped, 0.175 out of fold).
122
 
123
+ ### How it was built
124
 
125
  - **Base model**: Qwen3.5-4B-Base, a hybrid of 24 Gated DeltaNet (linear attention) layers and 8 full-attention layers. Because the recurrent layers cannot honour a block-causal mask, questions run as separate causal rows that continue from the shared state (`kev/model.py: forward_rows_batch`); isolation is exact by construction (together vs alone within 1e-5) and on attention-only models this form is bit-identical to the packed one.
126
+ - **Recipe**: `decision-v7`, two epochs, LoRA r=16 (attention, MLP and DeltaNet projections), lr 5e-5 — the same data and settings as every other Kev, so the Qwen3 → Qwen3.5 difference is the base (`PLAN.md`, Qwen3.5 port §10: locked test +7.3 pp [+2.8, +11.7] over Kev-8B).
127
  - **Delta**: `kev.train --init_from jaredpalmer/kev-4b@v7-base --data evals/night2/dates_unknowable.jsonl --replay 2000 --lr 2e-5 --epochs 1`. The 1,425 new records are generated (no public dataset): 900 date-bearing policy cases, a third rendered plainly, a third with a relational day-count sentence, a third with a `date_facts` field; 255 cases with the deciding sentence removed and a uniform soft target over the options, plus their 270 intact controls. Record hashes are in `evals/night2/manifest.json`; the source checkpoint's hashes are in `training_config.json`.
128
  - Why a delta and not a retrain: it is a controlled change (one fixed checkpoint, one data addition, 9 minutes), and the results section shows exactly what it moved.
129
 
130
+ ### Known limits
131
 
132
  - Use [Kev-9B](kev-9b.md) when accuracy and calibration matter more than memory: 0.852 vs 0.837 out of domain on the locked test, Brier 0.237 vs 0.255.
133
 
 
138
  - The raw logits are over-confident out of domain; the built-in temperature (T = 2.14) fixes most of it without changing any answer. `KEV_TEMPERATURE=1.0` gives the raw values. Coverage at a 5% error budget is 0.54–0.68 against Jev's 0.70.
139
  - 4B bf16 needs ~9 GB of GPU memory for serving; training took 56 min on one H100 (peak 24.6 GB).
140
 
141
+ ### Training
142
 
143
  Frozen suite `evals/v7/decision-v7`: 10,000 public records (1,000 per source), 896 policy minimal-pair records over nine template families, 1,680 records from 60 randomly generated rule structures in four rendering styles. Two epochs, LoRA r=16 α=32 on `q/k/v/o_proj`, `gate/up/down_proj`, `in_proj_qkv/z/a/b`, `out_proj`; pointer head from scratch; cross-entropy on the option distribution; lr 5e-5 (OneCycle), effective batch 8, bf16 autocast with fp32 master weights, gradient checkpointing; option permutation, none-of-the-above insertion, distractors, none minimal pairs on 25% of Choice records. Then the delta described above (one epoch, lr 2e-5, 3,937 records seen, 9 minutes on one H100). No Jev outputs were used for training.
144
 
145
+ ### Evaluation protocol
146
 
147
  Development partitions select models; the locked test partition is read at most once per candidate (`runs/locked/kev-4b-night2-du-ungated/`; the pre-delta read is `runs/locked/kev-4b-q35/`). Every number carries suite hash, code hashes and git commit in `result.json`. Untrained-base baselines use zero-shot letter logits on the same items (`scripts/base_mmlu_probe.py`).
148
 
adapter_model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:9797de69a42188e411b17b7b4fcb66a23374dcebc21d71a7a66f836b5d34df2b
3
  size 129924032
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c2c99077b58cd4a604304b743b7db5391187ef1d58e575f90c554c744aa23d65
3
  size 129924032
head.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:d8f796da36ff7bd7c0fb9496b452139bb7851af4fc82b07b500b682d3f721d6a
3
- size 5248767
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9a3051c3676f0be29d885935b6f50c287f7acf04e3cb29a6fed82ca3e5540b7e
3
+ size 5249727
provenance.json CHANGED
@@ -1,11 +1,11 @@
1
  {
2
  "config": {
3
  "epochs": 1,
4
- "seed": 1,
5
  "lr": 2e-05,
6
  "lora": 16,
7
- "accum": 2,
8
- "batch": 4,
9
  "perm_kl": 0.0,
10
  "perm_frac": 0.3,
11
  "ord_w": 0.0,
@@ -18,45 +18,62 @@
18
  "head_lr": 0.0,
19
  "weight_decay": 0.01,
20
  "anchor_w": 0.0,
 
 
 
21
  "dtype": "bf16",
22
  "checkpointing": 1,
23
  "replay": 2000,
 
 
24
  "base": "Qwen/Qwen3.5-4B-Base",
25
  "base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
26
- "init_from": "jaredpalmer/kev-4b",
27
- "data": "evals/night2/dates_unknowable.jsonl"
28
  },
29
- "config_sha256": "c155e290391f9a75d667a1e9c7f60d4b8e596a56900753fbf90d3a57d59f5488",
30
  "suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
31
  "source_hashes": {
32
  "kev/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
33
- "kev/anchors.py": "089d8a5493502bb26f540eb1c5e681780ca0bb01276073d4e1733211f15d0e10",
34
- "kev/api.py": "c9eebbdb6c625563d33f030299ce0fbe8b50493dd1e92592df9514c30f5720af",
35
- "kev/autoresearch.py": "0a8aa6b57c1c9ba93d25b2cf631b03aed2e686e374c1148002eec48765b7b1fc",
36
- "kev/benchmark.py": "4192ec3b26b065452f84bde38a091e6854a28fe185d2e0f39a0c91b2efe69df7",
37
- "kev/compare.py": "bd0445f021de59e35c7bff9304e39dd2e1211e3877a453575594ae7b81b0ada4",
 
 
38
  "kev/composition.py": "f335ed17e18e0a544893db5e22b9a059e6ce1b2e14dbb863ac7d7bbf8f3e0536",
39
  "kev/contrastive.py": "cbb979aa5d40265ad0e64695f94b281d91751fa405811ddfa8212fede111edcf",
40
- "kev/data.py": "997c31d737c2a130ade49edd6534aa47d910786c98af883745b1a97ec858a704",
41
- "kev/evaluate.py": "6f95c52f757ecffc887710e1b354f06c69c35336421d3e75eae84d251ddd120c",
42
- "kev/experiment.py": "a968eeede91385af2ae2fd98deaabf5461912752ae5ec71e8aba87223681eb26",
43
- "kev/jev.py": "e0213782359ba2f95adbf045ddaf0a008b08b4162bf6a9b66d4ce55fc91a51cb",
44
- "kev/model.py": "a17d57a52da3fe6ebef7146ce548bd3cedc09e21cc2b7738d9077771f5f16989",
45
- "kev/plot.py": "d689c7dd18cde7f9ea77cb4f55c50ff7da1880b21cb2ecf244348e45842a1382",
46
- "kev/publish.py": "c0b3efac3cfe93990d1846cfd306cf762a9d03e9e58380d13c2961a5ddae4619",
47
- "kev/serve.py": "d095bbee3a210fa8807d1a7b073b56181c93aa69b00f2e243a323ab0a19fa8c9",
48
- "kev/study_v3.py": "9fc44d44dad09a4f1ee29a7bcb2eb3c7aa373d09186d0e9533666b69ec95c401",
49
- "kev/suite.py": "44858f9a99df086a47d1ae36141631fab6fb6a8293a9e1d0c6ad397ddcfaf94a",
50
- "kev/train.py": "259a7f7d369045c2aacbfa166699470966b93c788f71a77e481f45866b26e42b",
51
- "kev/transfer_v9.py": "588aa2ff3ec0823c2e31733bef9a3748b263849c966ef6a80ebd686539c605e9",
52
- "modal_app.py": "0f5887718def578723383ece7a888ac90b521a641ba7e13f2e5d55da028e6065",
53
- "pyproject.toml": "52da5eea3efc6f2b1c0589acebad62e56a214294bb02c1a4218c93efd4af3182",
54
- "uv.lock": "18b3e5ea0f25d2e8546fab81f16cb965ae05c3289fffaaa1ce27d114adee47f3"
 
 
 
 
55
  },
56
- "git_commit": "99d0fa97b35de619f14913fa5a12a7ea9ec64903",
57
  "platform": "Linux-4.19.0-gvisor-x86_64-with-glibc2.36",
58
  "torch": "2.8.0+cu128",
59
  "device": "cuda",
60
- "gpu": "NVIDIA H100 80GB HBM3",
61
- "legacy_checkpoint": false
 
 
 
 
 
 
 
62
  }
 
1
  {
2
  "config": {
3
  "epochs": 1,
4
+ "seed": 2,
5
  "lr": 2e-05,
6
  "lora": 16,
7
+ "accum": 4,
8
+ "batch": 2,
9
  "perm_kl": 0.0,
10
  "perm_frac": 0.3,
11
  "ord_w": 0.0,
 
18
  "head_lr": 0.0,
19
  "weight_decay": 0.01,
20
  "anchor_w": 0.0,
21
+ "label_smoothing": 0.0,
22
+ "brier_w": 0.0,
23
+ "focal_gamma": 0.0,
24
  "dtype": "bf16",
25
  "checkpointing": 1,
26
  "replay": 2000,
27
+ "max_state": 7552,
28
+ "data": "evals/documents-v1/train.jsonl",
29
  "base": "Qwen/Qwen3.5-4B-Base",
30
  "base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
31
+ "init_from": "jaredpalmer/kev-4b"
 
32
  },
33
+ "config_sha256": "6fb0132214a793f590cce025ce669b85433bab47a63522c369c226f0fb34ff3c",
34
  "suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
35
  "source_hashes": {
36
  "kev/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
37
+ "kev/anchors.py": "6963eafdb276db5a6c94d939eae675c761555553b8d0f963ffd803448beaabbb",
38
+ "kev/api.py": "7bffacfb762c626b8bc2f670f350295af5ccb0883dbe90239c7d2f8e5ef56582",
39
+ "kev/autoresearch.py": "f7d7fabb9c0df7065bee3fec4aa4bee7028c5d1565847aefe2d74e3d7d40ed8a",
40
+ "kev/benchmark.py": "9c442efd8b6b852c574a646192c61ee91afa1306e3b2978bb6391ae1b91a3fa0",
41
+ "kev/calibrate.py": "c1eff4744fabd349e8abca86777a7aa0cbea44904c8223d3b61fca3d43741519",
42
+ "kev/checkpoint.py": "7cd4fbedfcfd11c2997490df2bd88eaa12af60e5ad60da2e5390e6df1ac1f626",
43
+ "kev/compare.py": "3606dbf98bb305d66420158498cb837304d2edb8dab01cdaf2ec42964bef3435",
44
  "kev/composition.py": "f335ed17e18e0a544893db5e22b9a059e6ce1b2e14dbb863ac7d7bbf8f3e0536",
45
  "kev/contrastive.py": "cbb979aa5d40265ad0e64695f94b281d91751fa405811ddfa8212fede111edcf",
46
+ "kev/data.py": "ef1be396e7218ba5c94f0f67de2ceff23a39656b57199cb3aa3192cde53e8320",
47
+ "kev/device.py": "d1677fd98ec0979c7284546306e34e0d09298ca4042fc39eb2b32a74f4e975c4",
48
+ "kev/evaluate.py": "999264f2837dcfbf2ec93601aa4745e701698ab850a43ad89324672a6604965c",
49
+ "kev/experiment.py": "644adedf43fbd10566830fa6094e16dd001b8cad29d9b9da82f0e0f245ca06b1",
50
+ "kev/jev.py": "acd4cc3f1844e438cc83a8d409c15ef78a5ef64d39ea10f583d75a6c7646b243",
51
+ "kev/metrics.py": "dba8d90550999edcd642581fb4da0367bd4b39d08471815184fa732d11e137b9",
52
+ "kev/mlx_model.py": "f582428796faf6bf962b772a69a6227ac3259ccbd3215c0c8b55ce28dbbd0909",
53
+ "kev/model.py": "387f281bc0b72da1620a8d2fcf508996fc30b9d8e342211ecd67c6d930f9cbaa",
54
+ "kev/plot.py": "0874bcff2885d8155a1cceade0de8a275a7163d3c3fd6aa0c294e8ccf5ff7e02",
55
+ "kev/predictors.py": "b2a66dd9f0f0a8026bc94603383bacfdd622e75d1c935173997175606626a9b3",
56
+ "kev/publish.py": "8abc9bcf4a01365697b05cf3c5ad0013462bd92d005f5616954e40ab1100c7aa",
57
+ "kev/serve.py": "c210b1f4598e64b49c399f73603ed372f4314c165ceda50e1340c0466ec1c735",
58
+ "kev/study_v3.py": "7891150aa623e4479c7c789d4dc4662f186132b37b7bac1e1f072b9e08ee13e3",
59
+ "kev/suite.py": "d524663fe2656d873c548a6216a2ddfebaf7299650f07ba31ae24e4c453398f1",
60
+ "kev/train.py": "68428ed43b3e362e52d4650dc5296ab88767061f7998a525715e1b306eeef474",
61
+ "kev/transfer_v9.py": "0902409742151250a28af1fd8f42258b70fcc161f3deed3ee35fe3f473f3c763",
62
+ "modal_app.py": "28318ef72dbeba7855f3c6b05339ee8c0f2e57562d21ede72352fba98d0e6ae3",
63
+ "pyproject.toml": "7c17fbe9efcadda3eb488b59adf1b6db5938cd33af4d3bd389ce4359b73fcc7d",
64
+ "uv.lock": "a9922dbb89acdef78299fd2b4a8c3f7f0fa1b2bc08b55595b6926fa785a9c466"
65
  },
66
+ "git_commit": "ea271bf9dd0f2a280b1ef096a7d237388484620c",
67
  "platform": "Linux-4.19.0-gvisor-x86_64-with-glibc2.36",
68
  "torch": "2.8.0+cu128",
69
  "device": "cuda",
70
+ "gpu": "NVIDIA H200",
71
+ "legacy_checkpoint": false,
72
+ "measured_checkpoint": {
73
+ "requested": "/runs/r8-small/00-trial-0/checkpoint",
74
+ "resolved": "/runs/r8-small/00-trial-0/checkpoint",
75
+ "head_sha256": "19fdcdfb65380b51e90458edeb7e8dd7554532f357379c4cda44152cbbcd4195",
76
+ "adapter_sha256": "c2c99077b58cd4a604304b743b7db5391187ef1d58e575f90c554c744aa23d65",
77
+ "inference_temperature": 1.0
78
+ }
79
  }
result.json CHANGED
The diff for this file is too large to render. See raw diff
 
train.log CHANGED
@@ -1,49 +1,97 @@
1
  Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
2
- delta: warm start from /__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/f780095d6511e6304868ff46ca2b74ee3eba29d8: 496 adapter tensors and the pointer head loaded
3
  device=cuda trainable params=33.8M
4
- replay: 2000 of 12576 suite training records mixed with 1425 from evals/night2/dates_unknowable.jsonl
5
- 3425 training requests (holdout=[]), questions by type {'noul': 1268, 'score': 1431, 'choice': 1214}
6
  [transformers] `causal_conv1d_fn` is falling back to its reference PyTorch implementation because `causal_conv1d` is not installed. This is correct but much slower; install `causal_conv1d` for the optimized kernel.
7
- ep0 step 10/429 loss 0.241 kl 0.000 anchor 0.000 1.518s/rec
8
- ep0 step 20/429 loss 0.354 kl 0.000 anchor 0.000 0.907s/rec
9
- ep0 step 30/429 loss 0.241 kl 0.000 anchor 0.000 0.645s/rec
10
- ep0 step 40/429 loss 0.470 kl 0.000 anchor 0.000 0.515s/rec
11
- ep0 step 50/429 loss 0.382 kl 0.000 anchor 0.000 0.435s/rec
12
- ep0 step 60/429 loss 0.248 kl 0.000 anchor 0.000 0.374s/rec
13
- ep0 step 70/429 loss 0.315 kl 0.000 anchor 0.000 0.334s/rec
14
- ep0 step 80/429 loss 0.367 kl 0.000 anchor 0.000 0.301s/rec
15
- ep0 step 90/429 loss 0.264 kl 0.000 anchor 0.000 0.278s/rec
16
- ep0 step 100/429 loss 0.339 kl 0.000 anchor 0.000 0.260s/rec
17
- ep0 step 110/429 loss 0.491 kl 0.000 anchor 0.000 0.246s/rec
18
- ep0 step 120/429 loss 0.277 kl 0.000 anchor 0.000 0.231s/rec
19
- ep0 step 130/429 loss 0.388 kl 0.000 anchor 0.000 0.222s/rec
20
- ep0 step 140/429 loss 0.147 kl 0.000 anchor 0.000 0.224s/rec
21
- ep0 step 150/429 loss 0.143 kl 0.000 anchor 0.000 0.216s/rec
22
- ep0 step 160/429 loss 0.387 kl 0.000 anchor 0.000 0.206s/rec
23
- ep0 step 170/429 loss 0.234 kl 0.000 anchor 0.000 0.198s/rec
24
- ep0 step 180/429 loss 0.321 kl 0.000 anchor 0.000 0.193s/rec
25
- ep0 step 190/429 loss 0.292 kl 0.000 anchor 0.000 0.187s/rec
26
- ep0 step 200/429 loss 0.296 kl 0.000 anchor 0.000 0.183s/rec
27
- ep0 step 210/429 loss 0.160 kl 0.000 anchor 0.000 0.180s/rec
28
- ep0 step 220/429 loss 0.302 kl 0.000 anchor 0.000 0.175s/rec
29
- ep0 step 230/429 loss 0.202 kl 0.000 anchor 0.000 0.171s/rec
30
- ep0 step 240/429 loss 0.313 kl 0.000 anchor 0.000 0.167s/rec
31
- ep0 step 250/429 loss 0.355 kl 0.000 anchor 0.000 0.164s/rec
32
- ep0 step 260/429 loss 0.260 kl 0.000 anchor 0.000 0.162s/rec
33
- ep0 step 270/429 loss 0.155 kl 0.000 anchor 0.000 0.158s/rec
34
- ep0 step 280/429 loss 0.308 kl 0.000 anchor 0.000 0.156s/rec
35
- ep0 step 290/429 loss 0.285 kl 0.000 anchor 0.000 0.154s/rec
36
- ep0 step 300/429 loss 0.173 kl 0.000 anchor 0.000 0.152s/rec
37
- ep0 step 310/429 loss 0.122 kl 0.000 anchor 0.000 0.150s/rec
38
- ep0 step 320/429 loss 0.300 kl 0.000 anchor 0.000 0.148s/rec
39
- ep0 step 330/429 loss 0.324 kl 0.000 anchor 0.000 0.146s/rec
40
- ep0 step 340/429 loss 0.268 kl 0.000 anchor 0.000 0.145s/rec
41
- ep0 step 350/429 loss 0.203 kl 0.000 anchor 0.000 0.143s/rec
42
- ep0 step 360/429 loss 0.180 kl 0.000 anchor 0.000 0.142s/rec
43
- ep0 step 370/429 loss 0.399 kl 0.000 anchor 0.000 0.140s/rec
44
- ep0 step 380/429 loss 0.162 kl 0.000 anchor 0.000 0.138s/rec
45
- ep0 step 390/429 loss 0.336 kl 0.000 anchor 0.000 0.137s/rec
46
- ep0 step 400/429 loss 0.195 kl 0.000 anchor 0.000 0.137s/rec
47
- ep0 step 410/429 loss 0.167 kl 0.000 anchor 0.000 0.135s/rec
48
- ep0 step 420/429 loss 0.167 kl 0.000 anchor 0.000 0.134s/rec
49
- saved /runs/night2-4b-du/00-trial-0/checkpoint
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
2
+ delta: warm start from /__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/485ace8703592fcf405488b262449990824cfed1: 496 adapter tensors and the pointer head loaded
3
  device=cuda trainable params=33.8M
4
+ replay: 2000 of 12576 suite training records mixed with 5219 from evals/documents-v1/train.jsonl
5
+ 7219 training requests (holdout=[]), questions by type {'choice': 8557, 'score': 580, 'noul': 837}
6
  [transformers] `causal_conv1d_fn` is falling back to its reference PyTorch implementation because `causal_conv1d` is not installed. This is correct but much slower; install `causal_conv1d` for the optimized kernel.
7
+ ep0 step 10/903 loss 0.574 kl 0.000 anchor 0.000 2.168s/rec
8
+ ep0 step 20/903 loss 0.461 kl 0.000 anchor 0.000 1.289s/rec
9
+ ep0 step 30/903 loss 0.478 kl 0.000 anchor 0.000 1.091s/rec
10
+ ep0 step 40/903 loss 0.427 kl 0.000 anchor 0.000 0.865s/rec
11
+ ep0 step 50/903 loss 0.417 kl 0.000 anchor 0.000 0.726s/rec
12
+ ep0 step 60/903 loss 0.334 kl 0.000 anchor 0.000 0.640s/rec
13
+ ep0 step 70/903 loss 0.464 kl 0.000 anchor 0.000 0.570s/rec
14
+ ep0 step 80/903 loss 0.365 kl 0.000 anchor 0.000 0.550s/rec
15
+ ep0 step 90/903 loss 0.221 kl 0.000 anchor 0.000 0.510s/rec
16
+ ep0 step 100/903 loss 0.299 kl 0.000 anchor 0.000 0.480s/rec
17
+ ep0 step 110/903 loss 0.673 kl 0.000 anchor 0.000 0.453s/rec
18
+ ep0 step 120/903 loss 0.232 kl 0.000 anchor 0.000 0.435s/rec
19
+ ep0 step 130/903 loss 0.207 kl 0.000 anchor 0.000 0.417s/rec
20
+ ep0 step 140/903 loss 0.257 kl 0.000 anchor 0.000 0.406s/rec
21
+ ep0 step 150/903 loss 0.363 kl 0.000 anchor 0.000 0.393s/rec
22
+ ep0 step 160/903 loss 0.418 kl 0.000 anchor 0.000 0.381s/rec
23
+ ep0 step 170/903 loss 0.426 kl 0.000 anchor 0.000 0.371s/rec
24
+ ep0 step 180/903 loss 0.244 kl 0.000 anchor 0.000 0.365s/rec
25
+ ep0 step 190/903 loss 0.142 kl 0.000 anchor 0.000 0.357s/rec
26
+ ep0 step 200/903 loss 0.196 kl 0.000 anchor 0.000 0.349s/rec
27
+ ep0 step 210/903 loss 0.224 kl 0.000 anchor 0.000 0.343s/rec
28
+ ep0 step 220/903 loss 0.290 kl 0.000 anchor 0.000 0.337s/rec
29
+ ep0 step 230/903 loss 0.248 kl 0.000 anchor 0.000 0.331s/rec
30
+ ep0 step 240/903 loss 0.176 kl 0.000 anchor 0.000 0.325s/rec
31
+ ep0 step 250/903 loss 0.135 kl 0.000 anchor 0.000 0.321s/rec
32
+ ep0 step 260/903 loss 0.328 kl 0.000 anchor 0.000 0.317s/rec
33
+ ep0 step 270/903 loss 0.398 kl 0.000 anchor 0.000 0.314s/rec
34
+ ep0 step 280/903 loss 0.311 kl 0.000 anchor 0.000 0.310s/rec
35
+ ep0 step 290/903 loss 0.075 kl 0.000 anchor 0.000 0.308s/rec
36
+ ep0 step 300/903 loss 0.155 kl 0.000 anchor 0.000 0.304s/rec
37
+ ep0 step 310/903 loss 0.367 kl 0.000 anchor 0.000 0.303s/rec
38
+ ep0 step 320/903 loss 0.358 kl 0.000 anchor 0.000 0.300s/rec
39
+ ep0 step 330/903 loss 0.253 kl 0.000 anchor 0.000 0.298s/rec
40
+ ep0 step 340/903 loss 0.400 kl 0.000 anchor 0.000 0.295s/rec
41
+ ep0 step 350/903 loss 0.087 kl 0.000 anchor 0.000 0.293s/rec
42
+ ep0 step 360/903 loss 0.247 kl 0.000 anchor 0.000 0.290s/rec
43
+ ep0 step 370/903 loss 0.225 kl 0.000 anchor 0.000 0.289s/rec
44
+ ep0 step 380/903 loss 0.171 kl 0.000 anchor 0.000 0.286s/rec
45
+ ep0 step 390/903 loss 0.402 kl 0.000 anchor 0.000 0.284s/rec
46
+ ep0 step 400/903 loss 0.151 kl 0.000 anchor 0.000 0.283s/rec
47
+ ep0 step 410/903 loss 0.175 kl 0.000 anchor 0.000 0.281s/rec
48
+ ep0 step 420/903 loss 0.200 kl 0.000 anchor 0.000 0.278s/rec
49
+ ep0 step 430/903 loss 0.460 kl 0.000 anchor 0.000 0.276s/rec
50
+ ep0 step 440/903 loss 0.176 kl 0.000 anchor 0.000 0.274s/rec
51
+ ep0 step 450/903 loss 0.124 kl 0.000 anchor 0.000 0.273s/rec
52
+ ep0 step 460/903 loss 0.139 kl 0.000 anchor 0.000 0.272s/rec
53
+ ep0 step 470/903 loss 0.285 kl 0.000 anchor 0.000 0.271s/rec
54
+ ep0 step 480/903 loss 0.172 kl 0.000 anchor 0.000 0.270s/rec
55
+ ep0 step 490/903 loss 0.296 kl 0.000 anchor 0.000 0.269s/rec
56
+ ep0 step 500/903 loss 0.254 kl 0.000 anchor 0.000 0.268s/rec
57
+ ep0 step 510/903 loss 0.147 kl 0.000 anchor 0.000 0.266s/rec
58
+ ep0 step 520/903 loss 0.168 kl 0.000 anchor 0.000 0.266s/rec
59
+ ep0 step 530/903 loss 0.166 kl 0.000 anchor 0.000 0.265s/rec
60
+ ep0 step 540/903 loss 0.222 kl 0.000 anchor 0.000 0.264s/rec
61
+ ep0 step 550/903 loss 0.178 kl 0.000 anchor 0.000 0.267s/rec
62
+ ep0 step 560/903 loss 0.183 kl 0.000 anchor 0.000 0.266s/rec
63
+ ep0 step 570/903 loss 0.383 kl 0.000 anchor 0.000 0.266s/rec
64
+ ep0 step 580/903 loss 0.316 kl 0.000 anchor 0.000 0.265s/rec
65
+ ep0 step 590/903 loss 0.082 kl 0.000 anchor 0.000 0.264s/rec
66
+ ep0 step 600/903 loss 0.176 kl 0.000 anchor 0.000 0.263s/rec
67
+ ep0 step 610/903 loss 0.064 kl 0.000 anchor 0.000 0.263s/rec
68
+ ep0 step 620/903 loss 0.108 kl 0.000 anchor 0.000 0.263s/rec
69
+ ep0 step 630/903 loss 0.181 kl 0.000 anchor 0.000 0.263s/rec
70
+ ep0 step 640/903 loss 0.073 kl 0.000 anchor 0.000 0.262s/rec
71
+ ep0 step 650/903 loss 0.152 kl 0.000 anchor 0.000 0.261s/rec
72
+ ep0 step 660/903 loss 0.142 kl 0.000 anchor 0.000 0.260s/rec
73
+ ep0 step 670/903 loss 0.452 kl 0.000 anchor 0.000 0.259s/rec
74
+ ep0 step 680/903 loss 0.165 kl 0.000 anchor 0.000 0.259s/rec
75
+ ep0 step 690/903 loss 0.126 kl 0.000 anchor 0.000 0.258s/rec
76
+ ep0 step 700/903 loss 0.177 kl 0.000 anchor 0.000 0.257s/rec
77
+ ep0 step 710/903 loss 0.200 kl 0.000 anchor 0.000 0.259s/rec
78
+ ep0 step 720/903 loss 0.218 kl 0.000 anchor 0.000 0.258s/rec
79
+ ep0 step 730/903 loss 0.173 kl 0.000 anchor 0.000 0.257s/rec
80
+ ep0 step 740/903 loss 0.233 kl 0.000 anchor 0.000 0.257s/rec
81
+ ep0 step 750/903 loss 0.123 kl 0.000 anchor 0.000 0.256s/rec
82
+ ep0 step 760/903 loss 0.070 kl 0.000 anchor 0.000 0.256s/rec
83
+ ep0 step 770/903 loss 0.339 kl 0.000 anchor 0.000 0.255s/rec
84
+ ep0 step 780/903 loss 0.338 kl 0.000 anchor 0.000 0.255s/rec
85
+ ep0 step 790/903 loss 0.131 kl 0.000 anchor 0.000 0.255s/rec
86
+ ep0 step 800/903 loss 0.110 kl 0.000 anchor 0.000 0.254s/rec
87
+ ep0 step 810/903 loss 0.110 kl 0.000 anchor 0.000 0.256s/rec
88
+ ep0 step 820/903 loss 0.100 kl 0.000 anchor 0.000 0.255s/rec
89
+ ep0 step 830/903 loss 0.390 kl 0.000 anchor 0.000 0.255s/rec
90
+ ep0 step 840/903 loss 0.119 kl 0.000 anchor 0.000 0.255s/rec
91
+ ep0 step 850/903 loss 0.233 kl 0.000 anchor 0.000 0.255s/rec
92
+ ep0 step 860/903 loss 0.226 kl 0.000 anchor 0.000 0.254s/rec
93
+ ep0 step 870/903 loss 0.121 kl 0.000 anchor 0.000 0.253s/rec
94
+ ep0 step 880/903 loss 0.125 kl 0.000 anchor 0.000 0.254s/rec
95
+ ep0 step 890/903 loss 0.604 kl 0.000 anchor 0.000 0.253s/rec
96
+ ep0 step 900/903 loss 0.379 kl 0.000 anchor 0.000 0.253s/rec
97
+ saved /runs/r8-small/00-trial-0/checkpoint
training_config.json CHANGED
@@ -7,21 +7,26 @@
7
  "head_lr": 0.0,
8
  "weight_decay": 0.01,
9
  "lora": 16,
10
- "accum": 2,
11
  "holdout": "",
12
  "perm_kl": 0.0,
13
  "perm_frac": 0.3,
14
  "ord_w": 0.0,
 
 
 
15
  "suite": "/root/evals/v7/decision-v7",
16
  "train_sources": "",
17
  "device": "cuda",
18
- "batch": 4,
19
  "dtype": "bf16",
 
20
  "checkpointing": 1,
21
  "option_isolation": 0,
22
  "special_embeddings": 0,
23
  "head_dim": 256,
24
  "lora_targets": "all",
 
25
  "base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
26
  "p_none": 0.1,
27
  "p_none_distract": 0.12,
@@ -32,19 +37,21 @@
32
  "anchor": "",
33
  "anchor_w": 0.0,
34
  "anchor_sources": "",
35
- "out": "/runs/night2-4b-du/00-trial-0/checkpoint",
36
- "data": "evals/night2/dates_unknowable.jsonl",
 
37
  "replay": 2000,
38
  "init_from": "jaredpalmer/kev-4b",
39
- "seed": 1
40
  },
41
  "suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
42
  "base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
43
  "init_source": {
44
  "init_from": "jaredpalmer/kev-4b",
45
- "resolved": "/__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/f780095d6511e6304868ff46ca2b74ee3eba29d8",
46
- "adapter_sha256": "20e798d1670df3c680d9b18aa1cad85d7bd6f25d57b2eed1a1621673d04f62a9",
47
- "head_sha256": "0976cae95ae97515c11fc324713e29218d39556a20f6d8bcc7d0c8a16738ea27"
 
48
  },
49
  "ordinal_objective": "ranked_probability_score",
50
  "holdout": []
 
7
  "head_lr": 0.0,
8
  "weight_decay": 0.01,
9
  "lora": 16,
10
+ "accum": 4,
11
  "holdout": "",
12
  "perm_kl": 0.0,
13
  "perm_frac": 0.3,
14
  "ord_w": 0.0,
15
+ "label_smoothing": 0.0,
16
+ "brier_w": 0.0,
17
+ "focal_gamma": 0.0,
18
  "suite": "/root/evals/v7/decision-v7",
19
  "train_sources": "",
20
  "device": "cuda",
21
+ "batch": 2,
22
  "dtype": "bf16",
23
+ "weights_dtype": "fp32",
24
  "checkpointing": 1,
25
  "option_isolation": 0,
26
  "special_embeddings": 0,
27
  "head_dim": 256,
28
  "lora_targets": "all",
29
+ "lora_placement": "full",
30
  "base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
31
  "p_none": 0.1,
32
  "p_none_distract": 0.12,
 
37
  "anchor": "",
38
  "anchor_w": 0.0,
39
  "anchor_sources": "",
40
+ "out": "/runs/r8-small/00-trial-0/checkpoint",
41
+ "data": "evals/documents-v1/train.jsonl",
42
+ "max_state": 7552,
43
  "replay": 2000,
44
  "init_from": "jaredpalmer/kev-4b",
45
+ "seed": 2
46
  },
47
  "suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
48
  "base_revision": "1001bb4d826a52d1f399e183466143f4da7b741b",
49
  "init_source": {
50
  "init_from": "jaredpalmer/kev-4b",
51
+ "resolved": "/__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-4b/snapshots/485ace8703592fcf405488b262449990824cfed1",
52
+ "adapter_sha256": "9797de69a42188e411b17b7b4fcb66a23374dcebc21d71a7a66f836b5d34df2b",
53
+ "head_sha256": "d8f796da36ff7bd7c0fb9496b452139bb7851af4fc82b07b500b682d3f721d6a",
54
+ "adapter_tensors": 496
55
  },
56
  "ordinal_objective": "ranked_probability_score",
57
  "holdout": []
training_metrics.json CHANGED
@@ -1,14 +1,14 @@
1
  {
2
- "wall_seconds": 528.0673320293427,
3
- "records_seen": 3937,
4
- "requested_records": 3425,
5
  "truncated_records": 0,
6
  "rejected_records": 0,
7
- "optimizer_steps": 429,
8
- "forward_tokens": 640184,
9
- "peak_device_bytes": 24556173824,
10
  "device": "cuda",
11
  "dtype": "bf16",
12
- "batch": 4,
13
- "peak_rss_bytes": 31547326464
14
  }
 
1
  {
2
+ "wall_seconds": 2574.7955305576324,
3
+ "records_seen": 10179,
4
+ "requested_records": 7219,
5
  "truncated_records": 0,
6
  "rejected_records": 0,
7
+ "optimizer_steps": 903,
8
+ "forward_tokens": 6749150,
9
+ "peak_device_bytes": 39839947776,
10
  "device": "cuda",
11
  "dtype": "bf16",
12
+ "batch": 2,
13
+ "peak_rss_bytes": 31563161600
14
  }