jaredpalmer commited on
Commit
9a45d25
·
verified ·
1 Parent(s): 54f4f87

Kev-0.8B round 15: documents + skills in one delta

Browse files
Files changed (8) hide show
  1. README.md +102 -13
  2. adapter_model.safetensors +1 -1
  3. head.pt +2 -2
  4. provenance.json +49 -30
  5. result.json +0 -0
  6. train.log +285 -46
  7. training_config.json +16 -10
  8. training_metrics.json +8 -8
README.md CHANGED
@@ -30,29 +30,118 @@ metrics:
30
  model-index:
31
  - name: Kev-0.8B
32
  results:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
33
  - task: { type: text-classification, name: typed decision (choice / noul / score) }
34
- dataset: { type: mixed, name: "decision-v7 development (1,204 records; ten trained public sources + programmatic policy data)" }
35
  metrics:
36
- - { type: accuracy, value: 0.825 }
37
- - { type: expected_calibration_error, value: 0.110, name: "ECE, raw probabilities" }
38
  - task: { type: text-classification, name: typed decision, out-of-domain }
39
- dataset: { type: mixed, name: "transfer-v4 development (764 records; six never-trained sources + held-out policy structures)" }
 
 
 
 
 
40
  metrics:
41
- - { type: accuracy, value: 0.652 }
42
- - { type: brier_score, value: 0.499 }
43
  ---
44
 
45
  # Kev-0.8B
46
 
47
  Kev-0.8B is a **decision model**: one document (the *state*) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation. It is a LoRA adapter (r=16, 11.3M trainable parameters) plus a pointer head on `Qwen/Qwen3.5-0.8B-Base` (revision `dc7cdfe2`), serving TypeSafe's public `/v1/systemone` contract.
48
 
49
- **The small member of the Kev family.** Same data and recipe as the 0.6B it replaces, on the Qwen3.5 base: in-distribution 0.825 (Kev-0.6B 0.801), out of domain 0.652 (0.620), and it is the first small Kev that learns any rule composition (held-out pairs 0.42 vs 0.08). Three seeds of the base recipe: transfer 0.622 / 0.634 / **0.643**; this checkpoint is seed 2 (selected on development accuracy) followed by a 9-minute **delta fine-tune** on 1,425 generated records (date-bearing policy cases with explicit day counts; evidence-free cases with uniform targets) mixed with 2,000 replayed training records — the same delta as Kev-4B and Kev-9B. Locked test against the pre-delta checkpoint: out of domain 0.668 → **0.684** (+2.2 pp [−0.8, +5.5]), Brier 0.473 → 0.460. Out of domain it is still a sub-1B model: use Kev-4B for accuracy; use this one where memory rules the 4B out, and measure on your own data.
 
 
 
 
 
 
50
 
51
- - Hub: `jaredpalmer/kev-0.8b` (this repo; trial `night2-08b-du2/00-trial-0`). The pre-delta checkpoint is at revision `v7-base`.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52
  - Demo: [huggingface.co/spaces/jaredpalmer/kev](https://huggingface.co/spaces/jaredpalmer/kev) runs Kev-4B and Kev-0.8B on ZeroGPU with the same encoder and API code as `kev.serve`.
53
- - Code, suites, results, and the full research log: [github.com/jaredpalmer/kev](https://github.com/jaredpalmer/kev) — `PLAN_Qwen35.md`, `PLAN.md`, `runs/leaderboard.md`
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
54
 
55
- ## Results (same frozen items for every row)
56
 
57
  | | Kev-0.6B (Qwen3) | **Kev-0.8B** | Kev-4B | Kev-9B | Jev |
58
  |---|---|---|---|---|---|
@@ -72,7 +161,7 @@ Paired against Kev-0.6B on the same items (record-clustered bootstrap), before t
72
 
73
  **Locked test, read once per checkpoint** (`runs/locked/kev-08b-night2-du-ungated/`; pre-delta `runs/locked/kev-08b-q35-ungated/`): in-distribution **0.834** (Brier 0.268, ECE 0.100), out-of-domain **0.684** (Brier 0.460, ECE 0.154, confident errors 8.7%, held-out pairs 0.45). Pre-delta: 0.827 / 0.668; Kev-0.6B on the same test items: 0.808 / 0.642.
74
 
75
- ## Known limits
76
 
77
  - **Out of domain it is a sub-1B model.** Knowledge (MMLU 0.41) and paraphrase (PAWS 0.59) are near the untrained base; the same recipe reaches 0.79 at 4B and 0.81 at 9B on these items.
78
  - **Slow on a Mac for its size.** The DeltaNet kernels have no MPS implementation; a five-question request takes ~0.33 s in bf16 on an M5 (Kev-0.6B: 0.12 s). On CUDA with `flash-linear-attention` it is fast.
@@ -80,11 +169,11 @@ Paired against Kev-0.6B on the same items (record-clustered bootstrap), before t
80
  - Ordinal hedging on date arithmetic (`deadline` 0.38): collapses to the middle level. `KEV_DATE_FACTS=1` (day counts appended to the state) helps the larger models more than this one.
81
  - Confident-error rate out of domain is 9.9% for the raw logits; the built-in temperature (T = 2.41, fitted on the in-distribution development rows and stored in `head.pt`) brings it to 0.3% and ECE from 0.179 to 0.054 without changing any answer. `KEV_TEMPERATURE=1.0` gives the raw values. Probabilities are usable in-domain; treat them as advisory elsewhere.
82
 
83
- ## Training
84
 
85
  Frozen suite `evals/v7/decision-v7`: 10,000 public records (1,000 per source), 896 policy minimal-pair records over nine template families, 1,680 records from 60 randomly generated rule structures in four rendering styles. Two epochs, LoRA r=16 α=32 on attention, MLP and DeltaNet projections; pointer head from scratch; cross-entropy on the option distribution; lr 1e-4 (OneCycle), batch 8, bf16 autocast with fp32 master weights; option permutation, none-of-the-above insertion, distractors, none minimal pairs on 25% of Choice records; ~20 min on one H100. Then the delta: `--init_from jaredpalmer/kev-0.8b@v7-base --data evals/night2/dates_unknowable.jsonl --replay 2000 --lr 4e-5 --epochs 1`, 9 minutes. No Jev outputs were used for training.
86
 
87
- ## Evaluation protocol
88
 
89
  Development partitions select models; the locked test partition is read at most once per candidate. Every number carries suite hash, code hashes and git commit in `result.json`.
90
 
 
30
  model-index:
31
  - name: Kev-0.8B
32
  results:
33
+ - task: { type: text-classification, name: typed decision, real documents, locked test }
34
+ dataset: { type: mixed, name: "documents-v1 test (936 questions on CFPB complaint narratives; read once)" }
35
+ metrics:
36
+ - { type: accuracy, value: 0.851 }
37
+ - { type: brier_score, value: 0.244 }
38
+ - task: { type: text-classification, name: typed decision, skill records, locked test }
39
+ dataset: { type: mixed, name: "hard-v1 test (1,088 questions; programmatic labels, held-out templates; read once)" }
40
+ metrics:
41
+ - { type: accuracy, value: 0.665 }
42
+ - { type: brier_score, value: 0.460 }
43
+ - task: { type: text-classification, name: typed decision, developer tooling, locked test }
44
+ dataset: { type: mixed, name: "devtools-v1 test (1,071 questions; six public developer-tooling sources; read once)" }
45
+ metrics:
46
+ - { type: accuracy, value: 0.637 }
47
+ - { type: brier_score, value: 0.442 }
48
  - task: { type: text-classification, name: typed decision (choice / noul / score) }
49
+ dataset: { type: mixed, name: "decision-v7 development (1,264 questions; ten trained public sources + programmatic policy data)" }
50
  metrics:
51
+ - { type: accuracy, value: 0.827 }
52
+ - { type: expected_calibration_error, value: 0.033, name: "ECE, as served" }
53
  - task: { type: text-classification, name: typed decision, out-of-domain }
54
+ dataset: { type: mixed, name: "transfer-v4 development (656 questions; six never-trained sources + held-out policy structures)" }
55
+ metrics:
56
+ - { type: accuracy, value: 0.648 }
57
+ - { type: brier_score, value: 0.430 }
58
+ - task: { type: text-classification, name: typed decision, out-of-domain, locked test }
59
+ dataset: { type: mixed, name: "transfer-v4 test (read once)" }
60
  metrics:
61
+ - { type: accuracy, value: 0.697 }
62
+ - { type: brier_score, value: 0.397 }
63
  ---
64
 
65
  # Kev-0.8B
66
 
67
  Kev-0.8B is a **decision model**: one document (the *state*) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation. It is a LoRA adapter (r=16, 11.3M trainable parameters) plus a pointer head on `Qwen/Qwen3.5-0.8B-Base` (revision `dc7cdfe2`), serving TypeSafe's public `/v1/systemone` contract.
68
 
69
+ **This version (2026-09-24): documents and skills in one delta.** The `night2-du` Kev-0.8B (below) plus one epoch (lr 2e-5, seed 1) on 16,539 records trained together, mixed with 6,000 replayed `decision-v7` records. They combine three sets. `documents-v1` train holds 5,219 real US consumer-finance complaint narratives (CFPB, up to ~7k tokens) with 7,488 questions (which product, which main issue), the set Kev-4B's round-8 delta used. `hard-v1` train holds 6,000 programmatically labelled records in seven skill families: long policy documents, trade-offs, probability, multi-hop, dates and arithmetic, judging a proposed answer, and missing-fact abstention. `devtools-v1` train holds 5,320 developer-tooling decisions from four public datasets. Every read below was taken once, on held-out splits:
70
+
71
+ - **Real documents.** The locked test goes from 0.608 to **0.851** (+24.4 pp [+21.3, +27.6], 936 questions; Brier 0.528 → 0.244). A private held-out set (`documents-v2`, 953 questions) goes from 0.616 to **0.848**.
72
+ - **Skills.** The `hard-v1` test goes from 0.396 to **0.665** (+26.9 pp [+23.4, +30.4], 1,088 questions). The `devtools-v1` test goes from 0.472 to **0.637** (+16.4 pp [+13.1, +19.4], 1,071 questions).
73
+ - **Everything else.** The locked out-of-domain test moved 0.684 → 0.697 (+1.2 pp [−1.1, +3.7]), with served Brier 0.412 → 0.397. On JevBench's public items, which no training or selection step saw, accuracy over all items goes from 0.597 to 0.636. The hard tier moves only from 0.333 to 0.360, and that change is within noise.
74
+
75
+ It is still a sub-1B model. On the development splits it trails Jev everywhere it can be compared: documents 0.842 vs 0.868, hard-v1 0.594 vs 0.777, devtools-v1 0.602 vs 0.713, out of domain 0.648 vs 0.857.
76
 
77
+ **Read this before relying on these numbers.**
78
+
79
+ - **All three gains are measured in distribution.** Training and every documents suite share one source (CFPB complaints) and the same two question templates. Evaluation labels are AI-adjudicated (a unanimous three-model judge panel, or two agreeing adjudications) and human spot-checked (47/50 and 50/50). `hard-v1` is generated: its labels are computed by each family's solver, and the test split holds out *templates* (0-3 train, 4 development, 5 test) of the same seven generators, so a test item is a new wording of a trained skill. JevBench's public hard tier is the out-of-distribution check, and there the paired gain is +2.7 pp [−1.8, +7.2]: not distinguishable from zero.
80
+ - **devtools-v1 labels are the public datasets' own, not adjudicated for this suite.** Some are human (CodeReviewer: whether a reviewer commented on the hunk; Aegis: human safety labels), some heuristic or by construction (CommitPackFT: the commit type is the first verb of the subject; FlakeFlagger: the test both passed and failed over reruns; When2Call: built by NVIDIA's pipeline). Before training, every model was near chance on CodeReviewer and FlakeFlagger, Jev included (see the [Kev-4B card](kev-4b.md)). After training on the same sources, this version gets 0.500 → 0.667 on CodeReviewer development. On FlakeFlagger it moves 0.500 → 0.520 on development but 0.507 → 0.813 on test. The two splits disagree that widely on a 150-question source, so treat the FlakeFlagger number as unexplained, not as a skill.
81
+ - **One eval-only source got worse.** When2Call, which is never trained on and asks whether to call a tool, ask for a missing parameter or decline, fell from 0.260 to 0.167 on development and from 0.233 to 0.133 on test, below the one-in-four rate of guessing among its four options. Prompt injection, also eval-only, is flat (0.527 → 0.547; Jev 0.893). Do not use this checkpoint for tool-call routing.
82
+ - **Other reads that went down.** Out-of-domain development accuracy moved 0.652 → 0.648, and coverage at ≤ 5 % error 0.229 → 0.145. TypeSafe's 89 answered rows moved 0.629 → 0.596 (−3.4 pp [−13.5, +5.0], 3 questions).
83
+
84
+ **Why it took until round 15.** Six registered rounds of 0.8B deltas came before this one.
85
+
86
+ - **Rounds 7, 8 and 9: documents alone.** They trained four 0.8B documents deltas. Each gained +20.8 to +22.7 pp on documents, and each failed the short-state guard, among other guards, on the 656-question transfer-v4 development panel. That panel cannot tell a cost of about 1 pp from one of 2 pp.
87
+ - **Round 11: a bigger panel.** It judged fresh seeds on a pooled short-state panel: the 656 transfer-v4 development questions plus 1,150 transfer-r3 test questions. The documents delta passed there (round 11). Separately, a skills-only delta passed (round 12).
88
+ - **Round 13: stacking failed.** It trained the skills data on top of the round-11 documents checkpoint, the path Kev-4B took. The second delta eroded the first: seed 1's documents lower bound was −2.03 pp against a −2 pp floor, and seed 2's short-state accuracy was −1.6 pp [−3.2, 0.0].
89
+ - **Round 15: joint training passed.** Before any training it registered one epoch on all three sets together, from the released checkpoint. Both primaries had to hold (documents development, and hard-v1 + devtools-v1 development pooled), with round 12's guards on the pooled short-state panel.
90
+ - This checkpoint, arm (a), scored documents +21.0 pp [+18.0, +24.0] and skills +18.0 [+15.8, +20.0], with short state −0.1 [−1.4, +1.3] and pooled externals +2.1 [+0.6, +3.5].
91
+ - The second seed at the same settings, arm (c), also passed (documents +20.2 [+16.9, +23.5], skills +16.7 [+14.5, +18.9], short state −0.3 [−1.7, +1.1]).
92
+ - A larger step, arm (b) at lr 4e-5, failed the short-state and scienthoon guards.
93
+ - Confirmation then required the documents-v1 test and the pooled hard-v1 + devtools-v1 tests to have lower bounds above zero (skills tests pooled +21.7 pp [+19.5, +24.0]), and one locked read (accuracy ≥ parent − 1 pp, served Brier ≤ parent + 0.005). All passed.
94
+ - The in-trial screening gate "held-out pairs ≥ 70 %" fails at this size, as it did for the released parent (0.422 for both), which is why the locked read is named `kev-08b-r15-ungated`.
95
+
96
+ - Hub: `jaredpalmer/kev-0.8b` (this repo; trial `r15-08b/00-trial-0`; the registration and every read are in `PLAN.md` rounds 9, 11, 12, 13 and 15 on the `research/overnight-r6` branch). The previous version is at tag `night2-du-release`; the pre-delta v7 checkpoint at `v7-base`.
97
  - Demo: [huggingface.co/spaces/jaredpalmer/kev](https://huggingface.co/spaces/jaredpalmer/kev) runs Kev-4B and Kev-0.8B on ZeroGPU with the same encoder and API code as `kev.serve`.
98
+ - Code, suites, results, and the full research log: [github.com/jaredpalmer/kev](https://github.com/jaredpalmer/kev) — `PLAN.md`, `runs/leaderboard.md`. The numbers below are in `runs/release/kev-08b-r15.json`; JevBench in `runs/jevbench-public/kev-08b-r15/`.
99
+
100
+ ## Results (as served: each checkpoint at its own fitted temperature)
101
+
102
+ | | **Kev-0.8B (this version, T = 2.35)** | `night2-du` Kev-0.8B (T = 2.41) | Jev |
103
+ |---|---|---|---|
104
+ | **real documents**, locked test (`documents-v1`, 936 questions) | **0.851** | 0.608 | – |
105
+ | real documents, private held-out (`documents-v2`, 953) | **0.848** | 0.616 | – |
106
+ | real documents, development (920) | **0.842** | 0.633 | 0.868 |
107
+ | real documents, Brier (locked test) | **0.244** | 0.528 | – |
108
+ | **hard-v1**, test (1,088, read once) | **0.665** | 0.396 | – |
109
+ | hard-v1, development (1,083) | **0.594** | 0.350 | 0.777 |
110
+ | hard-v1 ECE, test / development | 0.125 / 0.112 | 0.138 / 0.140 | – / 0.035 |
111
+ | **devtools-v1**, test (1,071, read once) | **0.637** | 0.472 | – |
112
+ | devtools-v1, development (1,072) | **0.602** | 0.487 | 0.713 |
113
+ | in-distribution accuracy (decision-v7 dev, 1,264 questions) | 0.827 | 0.825 | 0.845 |
114
+ | out-of-domain accuracy (transfer-v4 dev, 656) | 0.648 | 0.652 | 0.857 |
115
+ | out-of-domain Brier / ECE | 0.430 / 0.049 | 0.430 / 0.054 | 0.211 / 0.049 |
116
+ | confident errors out of domain (p ≥ 0.9 and wrong) | 0.2% | 0.3% | 3.7% |
117
+ | coverage at ≤ 5% error | 0.145 | 0.229 | 0.70 |
118
+ | held-out policy structures, both siblings correct | 0.422 | 0.422 | 0.86 |
119
+ | unknowable items answered at ≥ 0.9 (transfer-v9) | 0.00 | 0.00 | 0.09 |
120
+ | MMLU-Pro (transfer-v9 dev, 10-way) | 0.230 | 0.185 | 0.840 |
121
+ | **locked test**, out-of-domain accuracy / Brier | **0.697 / 0.397** | 0.684 / 0.412 | – |
122
+ | **locked test**, in-distribution accuracy | 0.838 | 0.834 | – |
123
+ | SemIf (144 authored decisions) | 0.722 | 0.701 | – |
124
+ | scienthoon (873 support tickets) | 0.534 | 0.520 | – |
125
+ | WANLI-v2 (1,002 NLI pairs) | 0.602 | 0.570 | – |
126
+ | TypeSafe (89 answered rows) | 0.596 | 0.629 | – |
127
+ | JevBench public items, all 231 (report only) | 0.636 | 0.597 | – |
128
+ | JevBench public, hard tier (111): accuracy / ECE | 0.360 / 0.181 | 0.333 / 0.245 | – |
129
+
130
+ Jev's devtools-v1 figure is over all 1,074 development questions; Kev's rows drop the reused CodeReviewer id (2 questions), as on the Kev-4B card.
131
+
132
+ Paired against the `night2-du` version (record-clustered bootstrap, 95 %): documents development +21.0 pp [+18.0, +24.0], locked test +24.4 [+21.3, +27.6], private held-out +23.2 [+19.9, +26.4]; hard-v1 development +24.4 [+20.9, +28.0], test +26.9 [+23.4, +30.4]; devtools-v1 development +11.5 [+8.9, +14.0], test +16.4 [+13.1, +19.4]; SemIf +2.1 [−3.5, +7.6]; scienthoon +1.4 [−1.5, +4.1]; WANLI-v2 +3.2 [+1.7, +4.8]; TypeSafe −3.4 [−13.5, +5.0]; locked out-of-domain test +1.2 [−1.1, +3.7].
133
+
134
+ **JevBench public items, report only.** `runs/jevbench-public/kev-08b-r15` holds the unchanged harness run against this checkpoint. Over all 231 public items accuracy goes 0.597 → 0.636, most of it on the standard tier: 0.736 → 0.819, with 6 items newly right and none newly wrong. On the hard tier it goes 0.333 → 0.360: paired over the 111 hard items that is +2.7 pp [−1.8, +7.2], with 5 newly right and 2 newly wrong (exact McNemar p = 0.453). Hard-tier ECE falls 0.245 → 0.181. For comparison, Kev-4B's skills delta gained +9.0 pp on the same hard items. The skill data moved the 0.8B far less out of distribution than in distribution.
135
+
136
+ **Calibration.** The fitted temperature moved 2.41 → 2.35 (raw out-of-domain Brier 0.481 on development, 0.416 on the locked test). As served, out-of-domain calibration is unchanged. `KEV_TEMPERATURE=1.0` gives the raw values.
137
+
138
+ ## Previous version: `night2-du` (2026-09-21) at tag `night2-du-release`
139
+
140
+ **The small member of the Kev family.** Same data and recipe as the 0.6B it replaces, on the Qwen3.5 base: in-distribution 0.825 (Kev-0.6B 0.801), out of domain 0.652 (0.620), and it is the first small Kev that learns any rule composition (held-out pairs 0.42 vs 0.08). Three seeds of the base recipe: transfer 0.622 / 0.634 / **0.643**; this checkpoint is seed 2 (selected on development accuracy) followed by a 9-minute **delta fine-tune** on 1,425 generated records (date-bearing policy cases with explicit day counts; evidence-free cases with uniform targets) mixed with 2,000 replayed training records — the same delta as Kev-4B and Kev-9B. Locked test against the pre-delta checkpoint: out of domain 0.668 → **0.684** (+2.2 pp [−0.8, +5.5]), Brier 0.473 → 0.460. Out of domain it is still a sub-1B model: use Kev-4B for accuracy; use this one where memory rules the 4B out, and measure on your own data.
141
+
142
+ - Trial `night2-08b-du2/00-trial-0`. The pre-delta checkpoint is at revision `v7-base`.
143
 
144
+ ### Results (same frozen items for every row)
145
 
146
  | | Kev-0.6B (Qwen3) | **Kev-0.8B** | Kev-4B | Kev-9B | Jev |
147
  |---|---|---|---|---|---|
 
161
 
162
  **Locked test, read once per checkpoint** (`runs/locked/kev-08b-night2-du-ungated/`; pre-delta `runs/locked/kev-08b-q35-ungated/`): in-distribution **0.834** (Brier 0.268, ECE 0.100), out-of-domain **0.684** (Brier 0.460, ECE 0.154, confident errors 8.7%, held-out pairs 0.45). Pre-delta: 0.827 / 0.668; Kev-0.6B on the same test items: 0.808 / 0.642.
163
 
164
+ ### Known limits
165
 
166
  - **Out of domain it is a sub-1B model.** Knowledge (MMLU 0.41) and paraphrase (PAWS 0.59) are near the untrained base; the same recipe reaches 0.79 at 4B and 0.81 at 9B on these items.
167
  - **Slow on a Mac for its size.** The DeltaNet kernels have no MPS implementation; a five-question request takes ~0.33 s in bf16 on an M5 (Kev-0.6B: 0.12 s). On CUDA with `flash-linear-attention` it is fast.
 
169
  - Ordinal hedging on date arithmetic (`deadline` 0.38): collapses to the middle level. `KEV_DATE_FACTS=1` (day counts appended to the state) helps the larger models more than this one.
170
  - Confident-error rate out of domain is 9.9% for the raw logits; the built-in temperature (T = 2.41, fitted on the in-distribution development rows and stored in `head.pt`) brings it to 0.3% and ECE from 0.179 to 0.054 without changing any answer. `KEV_TEMPERATURE=1.0` gives the raw values. Probabilities are usable in-domain; treat them as advisory elsewhere.
171
 
172
+ ### Training
173
 
174
  Frozen suite `evals/v7/decision-v7`: 10,000 public records (1,000 per source), 896 policy minimal-pair records over nine template families, 1,680 records from 60 randomly generated rule structures in four rendering styles. Two epochs, LoRA r=16 α=32 on attention, MLP and DeltaNet projections; pointer head from scratch; cross-entropy on the option distribution; lr 1e-4 (OneCycle), batch 8, bf16 autocast with fp32 master weights; option permutation, none-of-the-above insertion, distractors, none minimal pairs on 25% of Choice records; ~20 min on one H100. Then the delta: `--init_from jaredpalmer/kev-0.8b@v7-base --data evals/night2/dates_unknowable.jsonl --replay 2000 --lr 4e-5 --epochs 1`, 9 minutes. No Jev outputs were used for training.
175
 
176
+ ### Evaluation protocol
177
 
178
  Development partitions select models; the locked test partition is read at most once per candidate. Every number carries suite hash, code hashes and git commit in `result.json`.
179
 
adapter_model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:c81d5716f0af7622d8d2b97013c333d48263ca01113f9a7cf4526e96f6ac0b26
3
  size 43338624
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9b908623acb162118575f4e7a94524f9c139c335be4bfb74d6cfceca01e1885a
3
  size 43338624
head.pt CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:39f4343ccccc65e583bbfff0de0e11bfedb849fcfaaf94b50ac4f2b73bc79c65
3
- size 2103103
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f400bd12802b2b105ae45d6b03774a158a3db4fccff42413734ddca2e5c920b6
3
+ size 2103999
provenance.json CHANGED
@@ -2,10 +2,10 @@
2
  "config": {
3
  "epochs": 1,
4
  "seed": 1,
5
- "lr": 4e-05,
6
  "lora": 16,
7
- "accum": 1,
8
- "batch": 8,
9
  "perm_kl": 0.0,
10
  "perm_frac": 0.3,
11
  "ord_w": 0.0,
@@ -18,44 +18,63 @@
18
  "head_lr": 0.0,
19
  "weight_decay": 0.01,
20
  "anchor_w": 0.0,
 
 
 
21
  "dtype": "bf16",
22
- "replay": 2000,
 
 
 
23
  "base": "Qwen/Qwen3.5-0.8B-Base",
24
  "base_revision": "dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68",
25
- "init_from": "jaredpalmer/kev-0.8b",
26
- "data": "evals/night2/dates_unknowable.jsonl"
27
  },
28
- "config_sha256": "0e130e9f338c8a4c911c6dc37493b36e6be7ca972cbafbb42466043b6daa7a4b",
29
  "suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
30
  "source_hashes": {
31
  "kev/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
32
- "kev/anchors.py": "089d8a5493502bb26f540eb1c5e681780ca0bb01276073d4e1733211f15d0e10",
33
- "kev/api.py": "c9eebbdb6c625563d33f030299ce0fbe8b50493dd1e92592df9514c30f5720af",
34
- "kev/autoresearch.py": "0a8aa6b57c1c9ba93d25b2cf631b03aed2e686e374c1148002eec48765b7b1fc",
35
- "kev/benchmark.py": "4192ec3b26b065452f84bde38a091e6854a28fe185d2e0f39a0c91b2efe69df7",
36
- "kev/compare.py": "bd0445f021de59e35c7bff9304e39dd2e1211e3877a453575594ae7b81b0ada4",
 
 
37
  "kev/composition.py": "f335ed17e18e0a544893db5e22b9a059e6ce1b2e14dbb863ac7d7bbf8f3e0536",
38
  "kev/contrastive.py": "cbb979aa5d40265ad0e64695f94b281d91751fa405811ddfa8212fede111edcf",
39
- "kev/data.py": "997c31d737c2a130ade49edd6534aa47d910786c98af883745b1a97ec858a704",
40
- "kev/evaluate.py": "1b20e3f9edf417aa8dae924b1526e52f74b710cadf7213c5ec68334f6e7f8fe1",
41
- "kev/experiment.py": "ba034856f79bd500d8b2eeeb7382e4db3fccb30b437f9f3cf55cf1a7c565a6c9",
42
- "kev/jev.py": "e0213782359ba2f95adbf045ddaf0a008b08b4162bf6a9b66d4ce55fc91a51cb",
43
- "kev/model.py": "a17d57a52da3fe6ebef7146ce548bd3cedc09e21cc2b7738d9077771f5f16989",
44
- "kev/plot.py": "d689c7dd18cde7f9ea77cb4f55c50ff7da1880b21cb2ecf244348e45842a1382",
45
- "kev/publish.py": "8f23cda767008756e0f169249eb44577db76bd6c1e12e9cf64c302bad04c8f16",
46
- "kev/serve.py": "d095bbee3a210fa8807d1a7b073b56181c93aa69b00f2e243a323ab0a19fa8c9",
47
- "kev/study_v3.py": "9fc44d44dad09a4f1ee29a7bcb2eb3c7aa373d09186d0e9533666b69ec95c401",
48
- "kev/suite.py": "44858f9a99df086a47d1ae36141631fab6fb6a8293a9e1d0c6ad397ddcfaf94a",
49
- "kev/train.py": "c2a3770072b73dd4c0769c8f188ddedbb402f2f8195df6e10b11e5112cab734b",
50
- "kev/transfer_v9.py": "588aa2ff3ec0823c2e31733bef9a3748b263849c966ef6a80ebd686539c605e9",
51
- "modal_app.py": "d3700b2914be5ef6baa6d588a967d3a242903e1ed96d0f5cc525a8908aa58340",
52
- "pyproject.toml": "52da5eea3efc6f2b1c0589acebad62e56a214294bb02c1a4218c93efd4af3182",
53
- "uv.lock": "18b3e5ea0f25d2e8546fab81f16cb965ae05c3289fffaaa1ce27d114adee47f3"
 
 
 
 
 
54
  },
55
- "git_commit": "19dcae9b6e3e1a48200c5825aad9fc200d31e20a",
56
  "platform": "Linux-4.19.0-gvisor-x86_64-with-glibc2.36",
57
  "torch": "2.8.0+cu128",
58
  "device": "cuda",
59
- "gpu": "NVIDIA H100 80GB HBM3",
60
- "legacy_checkpoint": false
 
 
 
 
 
 
 
61
  }
 
2
  "config": {
3
  "epochs": 1,
4
  "seed": 1,
5
+ "lr": 2e-05,
6
  "lora": 16,
7
+ "accum": 2,
8
+ "batch": 4,
9
  "perm_kl": 0.0,
10
  "perm_frac": 0.3,
11
  "ord_w": 0.0,
 
18
  "head_lr": 0.0,
19
  "weight_decay": 0.01,
20
  "anchor_w": 0.0,
21
+ "label_smoothing": 0.0,
22
+ "brier_w": 0.0,
23
+ "focal_gamma": 0.0,
24
  "dtype": "bf16",
25
+ "checkpointing": 1,
26
+ "replay": 6000,
27
+ "max_state": 7552,
28
+ "data": "evals/round15/joint/train.jsonl",
29
  "base": "Qwen/Qwen3.5-0.8B-Base",
30
  "base_revision": "dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68",
31
+ "init_from": "jaredpalmer/kev-0.8b"
 
32
  },
33
+ "config_sha256": "6b30b3d48d805636aae8797c94bdce8ae8b7af866a06b6ac65bbc83ebb14d95a",
34
  "suite_sha256": "a8f50e481b7d90b97da049e0ff6a01cee2f1ed204aed61a8265af0edbb5514d2",
35
  "source_hashes": {
36
  "kev/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
37
+ "kev/anchors.py": "6963eafdb276db5a6c94d939eae675c761555553b8d0f963ffd803448beaabbb",
38
+ "kev/api.py": "7bffacfb762c626b8bc2f670f350295af5ccb0883dbe90239c7d2f8e5ef56582",
39
+ "kev/autoresearch.py": "f7d7fabb9c0df7065bee3fec4aa4bee7028c5d1565847aefe2d74e3d7d40ed8a",
40
+ "kev/benchmark.py": "10a892d76cb4c200e94d0a68ee74054aa1870d7f484c0530a531cacfa75032ed",
41
+ "kev/calibrate.py": "c1eff4744fabd349e8abca86777a7aa0cbea44904c8223d3b61fca3d43741519",
42
+ "kev/checkpoint.py": "f3edb4d159c0aa12d7006fc599fe6ebc3f2c288b3b69890c289f7dac451dd03d",
43
+ "kev/compare.py": "3606dbf98bb305d66420158498cb837304d2edb8dab01cdaf2ec42964bef3435",
44
  "kev/composition.py": "f335ed17e18e0a544893db5e22b9a059e6ce1b2e14dbb863ac7d7bbf8f3e0536",
45
  "kev/contrastive.py": "cbb979aa5d40265ad0e64695f94b281d91751fa405811ddfa8212fede111edcf",
46
+ "kev/cuda_graphs.py": "12f1953fea8f77c3c9dff29468202fbb37545145f93d3618194565a89949fd55",
47
+ "kev/data.py": "e77a2fffeb8ee7e05b118893be8e38b8358aef97c3ca2e143cdce208201b92f9",
48
+ "kev/device.py": "d1677fd98ec0979c7284546306e34e0d09298ca4042fc39eb2b32a74f4e975c4",
49
+ "kev/evaluate.py": "999264f2837dcfbf2ec93601aa4745e701698ab850a43ad89324672a6604965c",
50
+ "kev/experiment.py": "644adedf43fbd10566830fa6094e16dd001b8cad29d9b9da82f0e0f245ca06b1",
51
+ "kev/jev.py": "acd4cc3f1844e438cc83a8d409c15ef78a5ef64d39ea10f583d75a6c7646b243",
52
+ "kev/metrics.py": "dba8d90550999edcd642581fb4da0367bd4b39d08471815184fa732d11e137b9",
53
+ "kev/mlx_model.py": "f582428796faf6bf962b772a69a6227ac3259ccbd3215c0c8b55ce28dbbd0909",
54
+ "kev/model.py": "2634ffe7747d69473fb6596942c248e5df2586df3610bd1adb47e7e9acd99f96",
55
+ "kev/plot.py": "0874bcff2885d8155a1cceade0de8a275a7163d3c3fd6aa0c294e8ccf5ff7e02",
56
+ "kev/predictors.py": "4d02906c285d78bcbd91908f4fc86577ccc95c320d704817f6444aabbc4fcd4f",
57
+ "kev/publish.py": "8abc9bcf4a01365697b05cf3c5ad0013462bd92d005f5616954e40ab1100c7aa",
58
+ "kev/serve.py": "57d379bbcdeb3e4d4c3fdde5471e4dddeaf7710b8ed61ecad3466f3b67e6c610",
59
+ "kev/study_v3.py": "7891150aa623e4479c7c789d4dc4662f186132b37b7bac1e1f072b9e08ee13e3",
60
+ "kev/suite.py": "32b2882d1ea0fe5d03ba4d676a6180d580545f157e36fa715e7e89169ec99130",
61
+ "kev/train.py": "68428ed43b3e362e52d4650dc5296ab88767061f7998a525715e1b306eeef474",
62
+ "kev/transfer_v9.py": "0902409742151250a28af1fd8f42258b70fcc161f3deed3ee35fe3f473f3c763",
63
+ "modal_app.py": "7c6bc859f7d3e7011539a8f5e9e90390288cc72b5d80f911c8841c068f4d1136",
64
+ "pyproject.toml": "7c17fbe9efcadda3eb488b59adf1b6db5938cd33af4d3bd389ce4359b73fcc7d",
65
+ "uv.lock": "a9922dbb89acdef78299fd2b4a8c3f7f0fa1b2bc08b55595b6926fa785a9c466"
66
  },
67
+ "git_commit": "45923b7a3460b6d36358e2e143455902c1eb856b",
68
  "platform": "Linux-4.19.0-gvisor-x86_64-with-glibc2.36",
69
  "torch": "2.8.0+cu128",
70
  "device": "cuda",
71
+ "gpu": "NVIDIA H200",
72
+ "legacy_checkpoint": false,
73
+ "measured_checkpoint": {
74
+ "requested": "/runs/r15-08b/00-trial-0/checkpoint",
75
+ "resolved": "/runs/r15-08b/00-trial-0/checkpoint",
76
+ "head_sha256": "bc77488acd55d782a34ec875885c8647aa4afbb446fb1c4c3bb35fbc94ce865f",
77
+ "adapter_sha256": "9b908623acb162118575f4e7a94524f9c139c335be4bfb74d6cfceca01e1885a",
78
+ "inference_temperature": 1.0
79
+ }
80
  }
result.json CHANGED
The diff for this file is too large to render. See raw diff
 
train.log CHANGED
@@ -1,49 +1,288 @@
1
  Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
2
- delta: warm start from /__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-0.8b/snapshots/c917edefdfd72b3e9ba71455584700acc70595f6: 372 adapter tensors and the pointer head loaded
3
  device=cuda trainable params=11.3M
4
- replay: 2000 of 12576 suite training records mixed with 1425 from evals/night2/dates_unknowable.jsonl
5
- 3425 training requests (holdout=[]), questions by type {'noul': 1268, 'score': 1431, 'choice': 1214}
6
  [transformers] `causal_conv1d_fn` is falling back to its reference PyTorch implementation because `causal_conv1d` is not installed. This is correct but much slower; install `causal_conv1d` for the optimized kernel.
7
- ep0 step 10/429 loss 0.364 kl 0.000 anchor 0.000 1.620s/rec
8
- ep0 step 20/429 loss 0.487 kl 0.000 anchor 0.000 0.946s/rec
9
- ep0 step 30/429 loss 0.310 kl 0.000 anchor 0.000 0.650s/rec
10
- ep0 step 40/429 loss 0.476 kl 0.000 anchor 0.000 0.506s/rec
11
- ep0 step 50/429 loss 0.393 kl 0.000 anchor 0.000 0.416s/rec
12
- ep0 step 60/429 loss 0.418 kl 0.000 anchor 0.000 0.346s/rec
13
- ep0 step 70/429 loss 0.449 kl 0.000 anchor 0.000 0.301s/rec
14
- ep0 step 80/429 loss 0.461 kl 0.000 anchor 0.000 0.266s/rec
15
- ep0 step 90/429 loss 0.257 kl 0.000 anchor 0.000 0.240s/rec
16
- ep0 step 100/429 loss 0.734 kl 0.000 anchor 0.000 0.222s/rec
17
- ep0 step 110/429 loss 0.427 kl 0.000 anchor 0.000 0.205s/rec
18
- ep0 step 120/429 loss 0.269 kl 0.000 anchor 0.000 0.190s/rec
19
- ep0 step 130/429 loss 0.658 kl 0.000 anchor 0.000 0.177s/rec
20
- ep0 step 140/429 loss 0.192 kl 0.000 anchor 0.000 0.181s/rec
21
- ep0 step 150/429 loss 0.390 kl 0.000 anchor 0.000 0.171s/rec
22
- ep0 step 160/429 loss 0.528 kl 0.000 anchor 0.000 0.161s/rec
23
- ep0 step 170/429 loss 0.385 kl 0.000 anchor 0.000 0.153s/rec
24
- ep0 step 180/429 loss 0.368 kl 0.000 anchor 0.000 0.146s/rec
25
- ep0 step 190/429 loss 0.362 kl 0.000 anchor 0.000 0.140s/rec
26
- ep0 step 200/429 loss 0.365 kl 0.000 anchor 0.000 0.135s/rec
27
- ep0 step 210/429 loss 0.267 kl 0.000 anchor 0.000 0.130s/rec
28
- ep0 step 220/429 loss 0.463 kl 0.000 anchor 0.000 0.125s/rec
29
- ep0 step 230/429 loss 0.235 kl 0.000 anchor 0.000 0.121s/rec
30
- ep0 step 240/429 loss 0.228 kl 0.000 anchor 0.000 0.117s/rec
31
- ep0 step 250/429 loss 0.267 kl 0.000 anchor 0.000 0.113s/rec
32
- ep0 step 260/429 loss 0.424 kl 0.000 anchor 0.000 0.110s/rec
33
- ep0 step 270/429 loss 0.231 kl 0.000 anchor 0.000 0.107s/rec
34
- ep0 step 280/429 loss 0.409 kl 0.000 anchor 0.000 0.104s/rec
35
- ep0 step 290/429 loss 0.401 kl 0.000 anchor 0.000 0.102s/rec
36
- ep0 step 300/429 loss 0.402 kl 0.000 anchor 0.000 0.099s/rec
37
- ep0 step 310/429 loss 0.264 kl 0.000 anchor 0.000 0.097s/rec
38
- ep0 step 320/429 loss 0.337 kl 0.000 anchor 0.000 0.095s/rec
39
- ep0 step 330/429 loss 0.552 kl 0.000 anchor 0.000 0.093s/rec
40
- ep0 step 340/429 loss 0.436 kl 0.000 anchor 0.000 0.091s/rec
41
- ep0 step 350/429 loss 0.309 kl 0.000 anchor 0.000 0.089s/rec
42
- ep0 step 360/429 loss 0.239 kl 0.000 anchor 0.000 0.088s/rec
43
- ep0 step 370/429 loss 0.449 kl 0.000 anchor 0.000 0.087s/rec
44
- ep0 step 380/429 loss 0.207 kl 0.000 anchor 0.000 0.085s/rec
45
- ep0 step 390/429 loss 0.430 kl 0.000 anchor 0.000 0.084s/rec
46
- ep0 step 400/429 loss 0.216 kl 0.000 anchor 0.000 0.082s/rec
47
- ep0 step 410/429 loss 0.219 kl 0.000 anchor 0.000 0.081s/rec
48
- ep0 step 420/429 loss 0.284 kl 0.000 anchor 0.000 0.080s/rec
49
- saved /runs/night2-08b-du2/00-trial-0/checkpoint
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
2
+ delta: warm start from /__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-0.8b/snapshots/54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8: 372 adapter tensors and the pointer head loaded
3
  device=cuda trainable params=11.3M
4
+ replay: 6000 of 12576 suite training records mixed with 16539 from evals/round15/joint/train.jsonl
5
+ 22539 training requests (holdout=[]), questions by type {'choice': 19496, 'noul': 9796, 'score': 1931}
6
  [transformers] `causal_conv1d_fn` is falling back to its reference PyTorch implementation because `causal_conv1d` is not installed. This is correct but much slower; install `causal_conv1d` for the optimized kernel.
7
+ ep0 step 10/2818 loss 1.465 kl 0.000 anchor 0.000 0.180s/rec
8
+ ep0 step 20/2818 loss 1.563 kl 0.000 anchor 0.000 0.145s/rec
9
+ ep0 step 30/2818 loss 1.416 kl 0.000 anchor 0.000 0.129s/rec
10
+ ep0 step 40/2818 loss 1.453 kl 0.000 anchor 0.000 0.130s/rec
11
+ ep0 step 50/2818 loss 0.963 kl 0.000 anchor 0.000 0.126s/rec
12
+ ep0 step 60/2818 loss 0.991 kl 0.000 anchor 0.000 0.123s/rec
13
+ ep0 step 70/2818 loss 1.169 kl 0.000 anchor 0.000 0.118s/rec
14
+ ep0 step 80/2818 loss 1.124 kl 0.000 anchor 0.000 0.114s/rec
15
+ ep0 step 90/2818 loss 0.815 kl 0.000 anchor 0.000 0.113s/rec
16
+ ep0 step 100/2818 loss 0.975 kl 0.000 anchor 0.000 0.110s/rec
17
+ ep0 step 110/2818 loss 0.679 kl 0.000 anchor 0.000 0.108s/rec
18
+ ep0 step 120/2818 loss 0.823 kl 0.000 anchor 0.000 0.106s/rec
19
+ ep0 step 130/2818 loss 0.765 kl 0.000 anchor 0.000 0.103s/rec
20
+ ep0 step 140/2818 loss 0.788 kl 0.000 anchor 0.000 0.102s/rec
21
+ ep0 step 150/2818 loss 0.958 kl 0.000 anchor 0.000 0.103s/rec
22
+ ep0 step 160/2818 loss 0.748 kl 0.000 anchor 0.000 0.105s/rec
23
+ ep0 step 170/2818 loss 0.727 kl 0.000 anchor 0.000 0.104s/rec
24
+ ep0 step 180/2818 loss 0.706 kl 0.000 anchor 0.000 0.103s/rec
25
+ ep0 step 190/2818 loss 0.745 kl 0.000 anchor 0.000 0.105s/rec
26
+ ep0 step 200/2818 loss 0.738 kl 0.000 anchor 0.000 0.104s/rec
27
+ ep0 step 210/2818 loss 0.797 kl 0.000 anchor 0.000 0.104s/rec
28
+ ep0 step 220/2818 loss 0.680 kl 0.000 anchor 0.000 0.105s/rec
29
+ ep0 step 230/2818 loss 0.779 kl 0.000 anchor 0.000 0.105s/rec
30
+ ep0 step 240/2818 loss 0.825 kl 0.000 anchor 0.000 0.103s/rec
31
+ ep0 step 250/2818 loss 0.897 kl 0.000 anchor 0.000 0.103s/rec
32
+ ep0 step 260/2818 loss 0.752 kl 0.000 anchor 0.000 0.102s/rec
33
+ ep0 step 270/2818 loss 0.773 kl 0.000 anchor 0.000 0.103s/rec
34
+ ep0 step 280/2818 loss 0.691 kl 0.000 anchor 0.000 0.103s/rec
35
+ ep0 step 290/2818 loss 0.816 kl 0.000 anchor 0.000 0.104s/rec
36
+ ep0 step 300/2818 loss 0.791 kl 0.000 anchor 0.000 0.104s/rec
37
+ ep0 step 310/2818 loss 0.798 kl 0.000 anchor 0.000 0.104s/rec
38
+ ep0 step 320/2818 loss 0.664 kl 0.000 anchor 0.000 0.104s/rec
39
+ ep0 step 330/2818 loss 0.781 kl 0.000 anchor 0.000 0.104s/rec
40
+ ep0 step 340/2818 loss 0.904 kl 0.000 anchor 0.000 0.103s/rec
41
+ ep0 step 350/2818 loss 0.650 kl 0.000 anchor 0.000 0.103s/rec
42
+ ep0 step 360/2818 loss 0.727 kl 0.000 anchor 0.000 0.103s/rec
43
+ ep0 step 370/2818 loss 0.637 kl 0.000 anchor 0.000 0.102s/rec
44
+ ep0 step 380/2818 loss 0.550 kl 0.000 anchor 0.000 0.103s/rec
45
+ ep0 step 390/2818 loss 0.501 kl 0.000 anchor 0.000 0.102s/rec
46
+ ep0 step 400/2818 loss 0.629 kl 0.000 anchor 0.000 0.103s/rec
47
+ ep0 step 410/2818 loss 0.874 kl 0.000 anchor 0.000 0.103s/rec
48
+ ep0 step 420/2818 loss 0.721 kl 0.000 anchor 0.000 0.102s/rec
49
+ ep0 step 430/2818 loss 0.669 kl 0.000 anchor 0.000 0.102s/rec
50
+ ep0 step 440/2818 loss 0.710 kl 0.000 anchor 0.000 0.102s/rec
51
+ ep0 step 450/2818 loss 0.585 kl 0.000 anchor 0.000 0.101s/rec
52
+ ep0 step 460/2818 loss 0.867 kl 0.000 anchor 0.000 0.101s/rec
53
+ ep0 step 470/2818 loss 0.801 kl 0.000 anchor 0.000 0.101s/rec
54
+ ep0 step 480/2818 loss 0.708 kl 0.000 anchor 0.000 0.102s/rec
55
+ ep0 step 490/2818 loss 0.692 kl 0.000 anchor 0.000 0.101s/rec
56
+ ep0 step 500/2818 loss 0.883 kl 0.000 anchor 0.000 0.101s/rec
57
+ ep0 step 510/2818 loss 0.592 kl 0.000 anchor 0.000 0.102s/rec
58
+ ep0 step 520/2818 loss 0.717 kl 0.000 anchor 0.000 0.102s/rec
59
+ ep0 step 530/2818 loss 0.792 kl 0.000 anchor 0.000 0.101s/rec
60
+ ep0 step 540/2818 loss 0.683 kl 0.000 anchor 0.000 0.102s/rec
61
+ ep0 step 550/2818 loss 0.760 kl 0.000 anchor 0.000 0.102s/rec
62
+ ep0 step 560/2818 loss 0.628 kl 0.000 anchor 0.000 0.102s/rec
63
+ ep0 step 570/2818 loss 0.715 kl 0.000 anchor 0.000 0.102s/rec
64
+ ep0 step 580/2818 loss 0.643 kl 0.000 anchor 0.000 0.102s/rec
65
+ ep0 step 590/2818 loss 0.675 kl 0.000 anchor 0.000 0.102s/rec
66
+ ep0 step 600/2818 loss 0.463 kl 0.000 anchor 0.000 0.102s/rec
67
+ ep0 step 610/2818 loss 0.412 kl 0.000 anchor 0.000 0.102s/rec
68
+ ep0 step 620/2818 loss 0.669 kl 0.000 anchor 0.000 0.103s/rec
69
+ ep0 step 630/2818 loss 0.732 kl 0.000 anchor 0.000 0.102s/rec
70
+ ep0 step 640/2818 loss 0.566 kl 0.000 anchor 0.000 0.102s/rec
71
+ ep0 step 650/2818 loss 0.528 kl 0.000 anchor 0.000 0.102s/rec
72
+ ep0 step 660/2818 loss 0.743 kl 0.000 anchor 0.000 0.102s/rec
73
+ ep0 step 670/2818 loss 0.580 kl 0.000 anchor 0.000 0.102s/rec
74
+ ep0 step 680/2818 loss 0.639 kl 0.000 anchor 0.000 0.102s/rec
75
+ ep0 step 690/2818 loss 0.618 kl 0.000 anchor 0.000 0.102s/rec
76
+ ep0 step 700/2818 loss 0.425 kl 0.000 anchor 0.000 0.102s/rec
77
+ ep0 step 710/2818 loss 0.660 kl 0.000 anchor 0.000 0.102s/rec
78
+ ep0 step 720/2818 loss 0.746 kl 0.000 anchor 0.000 0.102s/rec
79
+ ep0 step 730/2818 loss 0.765 kl 0.000 anchor 0.000 0.102s/rec
80
+ ep0 step 740/2818 loss 0.730 kl 0.000 anchor 0.000 0.102s/rec
81
+ ep0 step 750/2818 loss 0.477 kl 0.000 anchor 0.000 0.102s/rec
82
+ ep0 step 760/2818 loss 0.661 kl 0.000 anchor 0.000 0.102s/rec
83
+ ep0 step 770/2818 loss 0.707 kl 0.000 anchor 0.000 0.103s/rec
84
+ ep0 step 780/2818 loss 0.632 kl 0.000 anchor 0.000 0.103s/rec
85
+ ep0 step 790/2818 loss 0.592 kl 0.000 anchor 0.000 0.102s/rec
86
+ ep0 step 800/2818 loss 0.632 kl 0.000 anchor 0.000 0.103s/rec
87
+ ep0 step 810/2818 loss 0.589 kl 0.000 anchor 0.000 0.103s/rec
88
+ ep0 step 820/2818 loss 0.721 kl 0.000 anchor 0.000 0.103s/rec
89
+ ep0 step 830/2818 loss 0.556 kl 0.000 anchor 0.000 0.104s/rec
90
+ ep0 step 840/2818 loss 0.714 kl 0.000 anchor 0.000 0.103s/rec
91
+ ep0 step 850/2818 loss 0.714 kl 0.000 anchor 0.000 0.103s/rec
92
+ ep0 step 860/2818 loss 0.756 kl 0.000 anchor 0.000 0.103s/rec
93
+ ep0 step 870/2818 loss 0.691 kl 0.000 anchor 0.000 0.103s/rec
94
+ ep0 step 880/2818 loss 0.665 kl 0.000 anchor 0.000 0.103s/rec
95
+ ep0 step 890/2818 loss 0.689 kl 0.000 anchor 0.000 0.104s/rec
96
+ ep0 step 900/2818 loss 0.470 kl 0.000 anchor 0.000 0.104s/rec
97
+ ep0 step 910/2818 loss 0.609 kl 0.000 anchor 0.000 0.104s/rec
98
+ ep0 step 920/2818 loss 0.581 kl 0.000 anchor 0.000 0.104s/rec
99
+ ep0 step 930/2818 loss 0.567 kl 0.000 anchor 0.000 0.104s/rec
100
+ ep0 step 940/2818 loss 0.572 kl 0.000 anchor 0.000 0.104s/rec
101
+ ep0 step 950/2818 loss 0.577 kl 0.000 anchor 0.000 0.104s/rec
102
+ ep0 step 960/2818 loss 0.653 kl 0.000 anchor 0.000 0.104s/rec
103
+ ep0 step 970/2818 loss 0.523 kl 0.000 anchor 0.000 0.104s/rec
104
+ ep0 step 980/2818 loss 0.748 kl 0.000 anchor 0.000 0.103s/rec
105
+ ep0 step 990/2818 loss 0.678 kl 0.000 anchor 0.000 0.103s/rec
106
+ ep0 step 1000/2818 loss 0.794 kl 0.000 anchor 0.000 0.103s/rec
107
+ ep0 step 1010/2818 loss 0.581 kl 0.000 anchor 0.000 0.103s/rec
108
+ ep0 step 1020/2818 loss 0.646 kl 0.000 anchor 0.000 0.104s/rec
109
+ ep0 step 1030/2818 loss 0.592 kl 0.000 anchor 0.000 0.104s/rec
110
+ ep0 step 1040/2818 loss 0.697 kl 0.000 anchor 0.000 0.104s/rec
111
+ ep0 step 1050/2818 loss 0.662 kl 0.000 anchor 0.000 0.104s/rec
112
+ ep0 step 1060/2818 loss 0.718 kl 0.000 anchor 0.000 0.104s/rec
113
+ ep0 step 1070/2818 loss 0.772 kl 0.000 anchor 0.000 0.104s/rec
114
+ ep0 step 1080/2818 loss 0.641 kl 0.000 anchor 0.000 0.104s/rec
115
+ ep0 step 1090/2818 loss 0.690 kl 0.000 anchor 0.000 0.104s/rec
116
+ ep0 step 1100/2818 loss 1.022 kl 0.000 anchor 0.000 0.104s/rec
117
+ ep0 step 1110/2818 loss 0.503 kl 0.000 anchor 0.000 0.104s/rec
118
+ ep0 step 1120/2818 loss 0.693 kl 0.000 anchor 0.000 0.104s/rec
119
+ ep0 step 1130/2818 loss 0.672 kl 0.000 anchor 0.000 0.104s/rec
120
+ ep0 step 1140/2818 loss 0.560 kl 0.000 anchor 0.000 0.104s/rec
121
+ ep0 step 1150/2818 loss 0.505 kl 0.000 anchor 0.000 0.104s/rec
122
+ ep0 step 1160/2818 loss 0.736 kl 0.000 anchor 0.000 0.104s/rec
123
+ ep0 step 1170/2818 loss 0.564 kl 0.000 anchor 0.000 0.104s/rec
124
+ ep0 step 1180/2818 loss 0.535 kl 0.000 anchor 0.000 0.104s/rec
125
+ ep0 step 1190/2818 loss 0.595 kl 0.000 anchor 0.000 0.104s/rec
126
+ ep0 step 1200/2818 loss 0.623 kl 0.000 anchor 0.000 0.104s/rec
127
+ ep0 step 1210/2818 loss 0.761 kl 0.000 anchor 0.000 0.104s/rec
128
+ ep0 step 1220/2818 loss 0.683 kl 0.000 anchor 0.000 0.104s/rec
129
+ ep0 step 1230/2818 loss 0.660 kl 0.000 anchor 0.000 0.104s/rec
130
+ ep0 step 1240/2818 loss 0.769 kl 0.000 anchor 0.000 0.104s/rec
131
+ ep0 step 1250/2818 loss 0.551 kl 0.000 anchor 0.000 0.104s/rec
132
+ ep0 step 1260/2818 loss 0.667 kl 0.000 anchor 0.000 0.104s/rec
133
+ ep0 step 1270/2818 loss 0.588 kl 0.000 anchor 0.000 0.104s/rec
134
+ ep0 step 1280/2818 loss 0.639 kl 0.000 anchor 0.000 0.104s/rec
135
+ ep0 step 1290/2818 loss 0.615 kl 0.000 anchor 0.000 0.104s/rec
136
+ ep0 step 1300/2818 loss 0.652 kl 0.000 anchor 0.000 0.104s/rec
137
+ ep0 step 1310/2818 loss 0.547 kl 0.000 anchor 0.000 0.104s/rec
138
+ ep0 step 1320/2818 loss 0.647 kl 0.000 anchor 0.000 0.104s/rec
139
+ ep0 step 1330/2818 loss 0.565 kl 0.000 anchor 0.000 0.104s/rec
140
+ ep0 step 1340/2818 loss 0.551 kl 0.000 anchor 0.000 0.104s/rec
141
+ ep0 step 1350/2818 loss 0.817 kl 0.000 anchor 0.000 0.104s/rec
142
+ ep0 step 1360/2818 loss 0.554 kl 0.000 anchor 0.000 0.104s/rec
143
+ ep0 step 1370/2818 loss 0.648 kl 0.000 anchor 0.000 0.104s/rec
144
+ ep0 step 1380/2818 loss 0.577 kl 0.000 anchor 0.000 0.104s/rec
145
+ ep0 step 1390/2818 loss 0.622 kl 0.000 anchor 0.000 0.104s/rec
146
+ ep0 step 1400/2818 loss 0.427 kl 0.000 anchor 0.000 0.104s/rec
147
+ ep0 step 1410/2818 loss 0.639 kl 0.000 anchor 0.000 0.104s/rec
148
+ ep0 step 1420/2818 loss 0.591 kl 0.000 anchor 0.000 0.104s/rec
149
+ ep0 step 1430/2818 loss 0.556 kl 0.000 anchor 0.000 0.104s/rec
150
+ ep0 step 1440/2818 loss 0.767 kl 0.000 anchor 0.000 0.104s/rec
151
+ ep0 step 1450/2818 loss 0.620 kl 0.000 anchor 0.000 0.104s/rec
152
+ ep0 step 1460/2818 loss 0.527 kl 0.000 anchor 0.000 0.104s/rec
153
+ ep0 step 1470/2818 loss 0.504 kl 0.000 anchor 0.000 0.104s/rec
154
+ ep0 step 1480/2818 loss 0.512 kl 0.000 anchor 0.000 0.104s/rec
155
+ ep0 step 1490/2818 loss 0.503 kl 0.000 anchor 0.000 0.104s/rec
156
+ ep0 step 1500/2818 loss 0.482 kl 0.000 anchor 0.000 0.105s/rec
157
+ ep0 step 1510/2818 loss 0.621 kl 0.000 anchor 0.000 0.105s/rec
158
+ ep0 step 1520/2818 loss 0.546 kl 0.000 anchor 0.000 0.105s/rec
159
+ ep0 step 1530/2818 loss 0.446 kl 0.000 anchor 0.000 0.105s/rec
160
+ ep0 step 1540/2818 loss 0.570 kl 0.000 anchor 0.000 0.105s/rec
161
+ ep0 step 1550/2818 loss 0.580 kl 0.000 anchor 0.000 0.105s/rec
162
+ ep0 step 1560/2818 loss 0.655 kl 0.000 anchor 0.000 0.105s/rec
163
+ ep0 step 1570/2818 loss 0.582 kl 0.000 anchor 0.000 0.105s/rec
164
+ ep0 step 1580/2818 loss 0.642 kl 0.000 anchor 0.000 0.105s/rec
165
+ ep0 step 1590/2818 loss 0.603 kl 0.000 anchor 0.000 0.105s/rec
166
+ ep0 step 1600/2818 loss 0.601 kl 0.000 anchor 0.000 0.105s/rec
167
+ ep0 step 1610/2818 loss 0.666 kl 0.000 anchor 0.000 0.105s/rec
168
+ ep0 step 1620/2818 loss 0.470 kl 0.000 anchor 0.000 0.105s/rec
169
+ ep0 step 1630/2818 loss 0.610 kl 0.000 anchor 0.000 0.105s/rec
170
+ ep0 step 1640/2818 loss 0.567 kl 0.000 anchor 0.000 0.105s/rec
171
+ ep0 step 1650/2818 loss 0.507 kl 0.000 anchor 0.000 0.105s/rec
172
+ ep0 step 1660/2818 loss 0.426 kl 0.000 anchor 0.000 0.104s/rec
173
+ ep0 step 1670/2818 loss 0.541 kl 0.000 anchor 0.000 0.104s/rec
174
+ ep0 step 1680/2818 loss 0.582 kl 0.000 anchor 0.000 0.104s/rec
175
+ ep0 step 1690/2818 loss 0.408 kl 0.000 anchor 0.000 0.104s/rec
176
+ ep0 step 1700/2818 loss 0.497 kl 0.000 anchor 0.000 0.104s/rec
177
+ ep0 step 1710/2818 loss 0.580 kl 0.000 anchor 0.000 0.104s/rec
178
+ ep0 step 1720/2818 loss 0.559 kl 0.000 anchor 0.000 0.104s/rec
179
+ ep0 step 1730/2818 loss 0.771 kl 0.000 anchor 0.000 0.104s/rec
180
+ ep0 step 1740/2818 loss 0.624 kl 0.000 anchor 0.000 0.104s/rec
181
+ ep0 step 1750/2818 loss 0.478 kl 0.000 anchor 0.000 0.104s/rec
182
+ ep0 step 1760/2818 loss 0.685 kl 0.000 anchor 0.000 0.104s/rec
183
+ ep0 step 1770/2818 loss 0.524 kl 0.000 anchor 0.000 0.104s/rec
184
+ ep0 step 1780/2818 loss 0.572 kl 0.000 anchor 0.000 0.104s/rec
185
+ ep0 step 1790/2818 loss 0.514 kl 0.000 anchor 0.000 0.104s/rec
186
+ ep0 step 1800/2818 loss 0.477 kl 0.000 anchor 0.000 0.104s/rec
187
+ ep0 step 1810/2818 loss 0.389 kl 0.000 anchor 0.000 0.104s/rec
188
+ ep0 step 1820/2818 loss 0.462 kl 0.000 anchor 0.000 0.104s/rec
189
+ ep0 step 1830/2818 loss 0.602 kl 0.000 anchor 0.000 0.104s/rec
190
+ ep0 step 1840/2818 loss 0.374 kl 0.000 anchor 0.000 0.104s/rec
191
+ ep0 step 1850/2818 loss 0.500 kl 0.000 anchor 0.000 0.104s/rec
192
+ ep0 step 1860/2818 loss 0.408 kl 0.000 anchor 0.000 0.104s/rec
193
+ ep0 step 1870/2818 loss 0.705 kl 0.000 anchor 0.000 0.104s/rec
194
+ ep0 step 1880/2818 loss 0.491 kl 0.000 anchor 0.000 0.104s/rec
195
+ ep0 step 1890/2818 loss 0.548 kl 0.000 anchor 0.000 0.104s/rec
196
+ ep0 step 1900/2818 loss 0.357 kl 0.000 anchor 0.000 0.104s/rec
197
+ ep0 step 1910/2818 loss 0.755 kl 0.000 anchor 0.000 0.104s/rec
198
+ ep0 step 1920/2818 loss 0.381 kl 0.000 anchor 0.000 0.104s/rec
199
+ ep0 step 1930/2818 loss 0.593 kl 0.000 anchor 0.000 0.104s/rec
200
+ ep0 step 1940/2818 loss 0.575 kl 0.000 anchor 0.000 0.104s/rec
201
+ ep0 step 1950/2818 loss 0.553 kl 0.000 anchor 0.000 0.104s/rec
202
+ ep0 step 1960/2818 loss 0.513 kl 0.000 anchor 0.000 0.104s/rec
203
+ ep0 step 1970/2818 loss 0.645 kl 0.000 anchor 0.000 0.104s/rec
204
+ ep0 step 1980/2818 loss 0.416 kl 0.000 anchor 0.000 0.104s/rec
205
+ ep0 step 1990/2818 loss 0.511 kl 0.000 anchor 0.000 0.104s/rec
206
+ ep0 step 2000/2818 loss 0.470 kl 0.000 anchor 0.000 0.104s/rec
207
+ ep0 step 2010/2818 loss 0.501 kl 0.000 anchor 0.000 0.104s/rec
208
+ ep0 step 2020/2818 loss 0.676 kl 0.000 anchor 0.000 0.104s/rec
209
+ ep0 step 2030/2818 loss 0.462 kl 0.000 anchor 0.000 0.104s/rec
210
+ ep0 step 2040/2818 loss 0.479 kl 0.000 anchor 0.000 0.104s/rec
211
+ ep0 step 2050/2818 loss 0.521 kl 0.000 anchor 0.000 0.104s/rec
212
+ ep0 step 2060/2818 loss 0.643 kl 0.000 anchor 0.000 0.104s/rec
213
+ ep0 step 2070/2818 loss 0.516 kl 0.000 anchor 0.000 0.104s/rec
214
+ ep0 step 2080/2818 loss 0.531 kl 0.000 anchor 0.000 0.104s/rec
215
+ ep0 step 2090/2818 loss 0.644 kl 0.000 anchor 0.000 0.104s/rec
216
+ ep0 step 2100/2818 loss 0.773 kl 0.000 anchor 0.000 0.104s/rec
217
+ ep0 step 2110/2818 loss 0.558 kl 0.000 anchor 0.000 0.104s/rec
218
+ ep0 step 2120/2818 loss 0.587 kl 0.000 anchor 0.000 0.104s/rec
219
+ ep0 step 2130/2818 loss 0.679 kl 0.000 anchor 0.000 0.104s/rec
220
+ ep0 step 2140/2818 loss 0.559 kl 0.000 anchor 0.000 0.104s/rec
221
+ ep0 step 2150/2818 loss 0.526 kl 0.000 anchor 0.000 0.104s/rec
222
+ ep0 step 2160/2818 loss 0.468 kl 0.000 anchor 0.000 0.104s/rec
223
+ ep0 step 2170/2818 loss 0.327 kl 0.000 anchor 0.000 0.104s/rec
224
+ ep0 step 2180/2818 loss 0.686 kl 0.000 anchor 0.000 0.104s/rec
225
+ ep0 step 2190/2818 loss 0.395 kl 0.000 anchor 0.000 0.104s/rec
226
+ ep0 step 2200/2818 loss 0.564 kl 0.000 anchor 0.000 0.104s/rec
227
+ ep0 step 2210/2818 loss 0.551 kl 0.000 anchor 0.000 0.104s/rec
228
+ ep0 step 2220/2818 loss 0.475 kl 0.000 anchor 0.000 0.104s/rec
229
+ ep0 step 2230/2818 loss 0.630 kl 0.000 anchor 0.000 0.104s/rec
230
+ ep0 step 2240/2818 loss 0.515 kl 0.000 anchor 0.000 0.104s/rec
231
+ ep0 step 2250/2818 loss 0.657 kl 0.000 anchor 0.000 0.104s/rec
232
+ ep0 step 2260/2818 loss 0.482 kl 0.000 anchor 0.000 0.104s/rec
233
+ ep0 step 2270/2818 loss 0.600 kl 0.000 anchor 0.000 0.104s/rec
234
+ ep0 step 2280/2818 loss 0.478 kl 0.000 anchor 0.000 0.104s/rec
235
+ ep0 step 2290/2818 loss 0.654 kl 0.000 anchor 0.000 0.104s/rec
236
+ ep0 step 2300/2818 loss 0.577 kl 0.000 anchor 0.000 0.104s/rec
237
+ ep0 step 2310/2818 loss 0.587 kl 0.000 anchor 0.000 0.104s/rec
238
+ ep0 step 2320/2818 loss 0.544 kl 0.000 anchor 0.000 0.104s/rec
239
+ ep0 step 2330/2818 loss 0.567 kl 0.000 anchor 0.000 0.103s/rec
240
+ ep0 step 2340/2818 loss 0.611 kl 0.000 anchor 0.000 0.103s/rec
241
+ ep0 step 2350/2818 loss 0.441 kl 0.000 anchor 0.000 0.103s/rec
242
+ ep0 step 2360/2818 loss 0.711 kl 0.000 anchor 0.000 0.104s/rec
243
+ ep0 step 2370/2818 loss 0.550 kl 0.000 anchor 0.000 0.104s/rec
244
+ ep0 step 2380/2818 loss 0.408 kl 0.000 anchor 0.000 0.104s/rec
245
+ ep0 step 2390/2818 loss 0.652 kl 0.000 anchor 0.000 0.104s/rec
246
+ ep0 step 2400/2818 loss 0.684 kl 0.000 anchor 0.000 0.104s/rec
247
+ ep0 step 2410/2818 loss 0.591 kl 0.000 anchor 0.000 0.104s/rec
248
+ ep0 step 2420/2818 loss 0.505 kl 0.000 anchor 0.000 0.104s/rec
249
+ ep0 step 2430/2818 loss 0.546 kl 0.000 anchor 0.000 0.104s/rec
250
+ ep0 step 2440/2818 loss 0.448 kl 0.000 anchor 0.000 0.104s/rec
251
+ ep0 step 2450/2818 loss 0.379 kl 0.000 anchor 0.000 0.104s/rec
252
+ ep0 step 2460/2818 loss 0.336 kl 0.000 anchor 0.000 0.104s/rec
253
+ ep0 step 2470/2818 loss 0.582 kl 0.000 anchor 0.000 0.104s/rec
254
+ ep0 step 2480/2818 loss 0.503 kl 0.000 anchor 0.000 0.104s/rec
255
+ ep0 step 2490/2818 loss 0.648 kl 0.000 anchor 0.000 0.104s/rec
256
+ ep0 step 2500/2818 loss 0.549 kl 0.000 anchor 0.000 0.104s/rec
257
+ ep0 step 2510/2818 loss 0.527 kl 0.000 anchor 0.000 0.104s/rec
258
+ ep0 step 2520/2818 loss 0.602 kl 0.000 anchor 0.000 0.104s/rec
259
+ ep0 step 2530/2818 loss 0.490 kl 0.000 anchor 0.000 0.104s/rec
260
+ ep0 step 2540/2818 loss 0.546 kl 0.000 anchor 0.000 0.104s/rec
261
+ ep0 step 2550/2818 loss 0.541 kl 0.000 anchor 0.000 0.104s/rec
262
+ ep0 step 2560/2818 loss 0.613 kl 0.000 anchor 0.000 0.104s/rec
263
+ ep0 step 2570/2818 loss 0.578 kl 0.000 anchor 0.000 0.104s/rec
264
+ ep0 step 2580/2818 loss 0.578 kl 0.000 anchor 0.000 0.104s/rec
265
+ ep0 step 2590/2818 loss 0.528 kl 0.000 anchor 0.000 0.104s/rec
266
+ ep0 step 2600/2818 loss 0.648 kl 0.000 anchor 0.000 0.104s/rec
267
+ ep0 step 2610/2818 loss 0.561 kl 0.000 anchor 0.000 0.104s/rec
268
+ ep0 step 2620/2818 loss 0.392 kl 0.000 anchor 0.000 0.103s/rec
269
+ ep0 step 2630/2818 loss 0.369 kl 0.000 anchor 0.000 0.103s/rec
270
+ ep0 step 2640/2818 loss 0.467 kl 0.000 anchor 0.000 0.103s/rec
271
+ ep0 step 2650/2818 loss 0.465 kl 0.000 anchor 0.000 0.104s/rec
272
+ ep0 step 2660/2818 loss 0.504 kl 0.000 anchor 0.000 0.104s/rec
273
+ ep0 step 2670/2818 loss 0.449 kl 0.000 anchor 0.000 0.104s/rec
274
+ ep0 step 2680/2818 loss 0.388 kl 0.000 anchor 0.000 0.104s/rec
275
+ ep0 step 2690/2818 loss 0.525 kl 0.000 anchor 0.000 0.104s/rec
276
+ ep0 step 2700/2818 loss 0.450 kl 0.000 anchor 0.000 0.104s/rec
277
+ ep0 step 2710/2818 loss 0.626 kl 0.000 anchor 0.000 0.104s/rec
278
+ ep0 step 2720/2818 loss 0.564 kl 0.000 anchor 0.000 0.104s/rec
279
+ ep0 step 2730/2818 loss 0.474 kl 0.000 anchor 0.000 0.103s/rec
280
+ ep0 step 2740/2818 loss 0.650 kl 0.000 anchor 0.000 0.104s/rec
281
+ ep0 step 2750/2818 loss 0.456 kl 0.000 anchor 0.000 0.104s/rec
282
+ ep0 step 2760/2818 loss 0.695 kl 0.000 anchor 0.000 0.104s/rec
283
+ ep0 step 2770/2818 loss 0.435 kl 0.000 anchor 0.000 0.104s/rec
284
+ ep0 step 2780/2818 loss 0.579 kl 0.000 anchor 0.000 0.104s/rec
285
+ ep0 step 2790/2818 loss 0.465 kl 0.000 anchor 0.000 0.104s/rec
286
+ ep0 step 2800/2818 loss 0.585 kl 0.000 anchor 0.000 0.104s/rec
287
+ ep0 step 2810/2818 loss 0.415 kl 0.000 anchor 0.000 0.103s/rec
288
+ saved /runs/r15-08b/00-trial-0/checkpoint
training_config.json CHANGED
@@ -3,26 +3,30 @@
3
  "base": "Qwen/Qwen3.5-0.8B-Base",
4
  "n_per_source": 1000,
5
  "epochs": 1,
6
- "lr": 4e-05,
7
  "head_lr": 0.0,
8
  "weight_decay": 0.01,
9
  "lora": 16,
10
- "accum": 1,
11
  "holdout": "",
12
  "perm_kl": 0.0,
13
  "perm_frac": 0.3,
14
  "ord_w": 0.0,
 
 
 
15
  "suite": "/root/evals/v7/decision-v7",
16
  "train_sources": "",
17
  "device": "cuda",
18
- "batch": 8,
19
  "dtype": "bf16",
20
  "weights_dtype": "fp32",
21
- "checkpointing": 0,
22
  "option_isolation": 0,
23
  "special_embeddings": 0,
24
  "head_dim": 256,
25
  "lora_targets": "all",
 
26
  "base_revision": "dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68",
27
  "p_none": 0.1,
28
  "p_none_distract": 0.12,
@@ -33,9 +37,10 @@
33
  "anchor": "",
34
  "anchor_w": 0.0,
35
  "anchor_sources": "",
36
- "out": "/runs/night2-08b-du2/00-trial-0/checkpoint",
37
- "data": "evals/night2/dates_unknowable.jsonl",
38
- "replay": 2000,
 
39
  "init_from": "jaredpalmer/kev-0.8b",
40
  "seed": 1
41
  },
@@ -43,9 +48,10 @@
43
  "base_revision": "dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68",
44
  "init_source": {
45
  "init_from": "jaredpalmer/kev-0.8b",
46
- "resolved": "/__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-0.8b/snapshots/c917edefdfd72b3e9ba71455584700acc70595f6",
47
- "adapter_sha256": "d9fa619fd3b0490122454c386bfa1b53c23850189d1b5c4346962162e2e39a64",
48
- "head_sha256": "8610dac1c30a64bbcea7715f20a258f4d084d7b862c10f96a1625c732049a32a"
 
49
  },
50
  "ordinal_objective": "ranked_probability_score",
51
  "holdout": []
 
3
  "base": "Qwen/Qwen3.5-0.8B-Base",
4
  "n_per_source": 1000,
5
  "epochs": 1,
6
+ "lr": 2e-05,
7
  "head_lr": 0.0,
8
  "weight_decay": 0.01,
9
  "lora": 16,
10
+ "accum": 2,
11
  "holdout": "",
12
  "perm_kl": 0.0,
13
  "perm_frac": 0.3,
14
  "ord_w": 0.0,
15
+ "label_smoothing": 0.0,
16
+ "brier_w": 0.0,
17
+ "focal_gamma": 0.0,
18
  "suite": "/root/evals/v7/decision-v7",
19
  "train_sources": "",
20
  "device": "cuda",
21
+ "batch": 4,
22
  "dtype": "bf16",
23
  "weights_dtype": "fp32",
24
+ "checkpointing": 1,
25
  "option_isolation": 0,
26
  "special_embeddings": 0,
27
  "head_dim": 256,
28
  "lora_targets": "all",
29
+ "lora_placement": "full",
30
  "base_revision": "dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68",
31
  "p_none": 0.1,
32
  "p_none_distract": 0.12,
 
37
  "anchor": "",
38
  "anchor_w": 0.0,
39
  "anchor_sources": "",
40
+ "out": "/runs/r15-08b/00-trial-0/checkpoint",
41
+ "data": "evals/round15/joint/train.jsonl",
42
+ "max_state": 7552,
43
+ "replay": 6000,
44
  "init_from": "jaredpalmer/kev-0.8b",
45
  "seed": 1
46
  },
 
48
  "base_revision": "dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68",
49
  "init_source": {
50
  "init_from": "jaredpalmer/kev-0.8b",
51
+ "resolved": "/__modal/volumes/vo-kEMu8BkBAIrorAQI6V8f2D/hub/models--jaredpalmer--kev-0.8b/snapshots/54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8",
52
+ "adapter_sha256": "c81d5716f0af7622d8d2b97013c333d48263ca01113f9a7cf4526e96f6ac0b26",
53
+ "head_sha256": "39f4343ccccc65e583bbfff0de0e11bfedb849fcfaaf94b50ac4f2b73bc79c65",
54
+ "adapter_tensors": 372
55
  },
56
  "ordinal_objective": "ranked_probability_score",
57
  "holdout": []
training_metrics.json CHANGED
@@ -1,14 +1,14 @@
1
  {
2
- "wall_seconds": 313.94338822364807,
3
- "records_seen": 3937,
4
- "requested_records": 3425,
5
  "truncated_records": 0,
6
  "rejected_records": 0,
7
- "optimizer_steps": 429,
8
- "forward_tokens": 640184,
9
- "peak_device_bytes": 43896273920,
10
  "device": "cuda",
11
  "dtype": "bf16",
12
- "batch": 8,
13
- "peak_rss_bytes": 10316816384
14
  }
 
1
  {
2
+ "wall_seconds": 3136.3209748268127,
3
+ "records_seen": 30329,
4
+ "requested_records": 22539,
5
  "truncated_records": 0,
6
  "rejected_records": 0,
7
+ "optimizer_steps": 2818,
8
+ "forward_tokens": 15531716,
9
+ "peak_device_bytes": 23937857536,
10
  "device": "cuda",
11
  "dtype": "bf16",
12
+ "batch": 4,
13
+ "peak_rss_bytes": 10316873728
14
  }