GoatHerder commited on
Commit
5f2a11f
·
verified ·
1 Parent(s): 086613e

Upload 49 files

Browse files
README.md CHANGED
@@ -2,327 +2,187 @@
2
  license: apache-2.0
3
  language:
4
  - en
5
- base_model:
6
- - convaiinnovations/laya
7
- base_model_relation: adapter
8
  pipeline_tag: text-classification
 
 
 
 
9
  tags:
10
  - laya
11
- - typed-decisions
12
- - system-one
13
- - classification
14
  - adapter
15
- - linear-interface
16
- - representation-alignment
17
- datasets:
18
- - LocalLLaMA/typed-decisions
19
- metrics:
20
- - accuracy
21
- - brier_score
22
- library_name: transformers
23
  ---
24
- # Ariadne-Laya-TD
25
 
26
- **Ariadne** is a lightweight representation-alignment adapter for the Laya base model.
27
 
28
- Ariadne keeps the underlying Laya encoder frozen and trains only a single identity-initialized affine transformation immediately before the encoder stack:
29
 
30
- $$
31
- h' = Wh + b,\qquad W_0 = I,\qquad b_0 = 0
32
- $$
33
 
34
- The interface contains just **1,049,600 trainable parameters**.
35
 
36
- The experiment asks a simple question:
37
 
38
- > **Instead of adapting the representation to the task, can we adapt the task to the representation?**
39
 
40
- Early results suggest that much of the discrimination required for typed decisions is already accessible in the frozen Laya base representation.
41
- Ariadne learns a small transformation that allows the existing model to use that structure without modifying the encoder itself.
42
-
43
- This makes specialisms extremely lightweight: a deployment can keep one frozen Laya base resident and switch between task-specific Ariadne interfaces containing only 1.05M parameters—roughly 2 MB in BF16—rather than loading a separate ~421M-parameter model for each specialism.
44
 
45
  ## Architecture
46
 
47
- Ariadne adds one trainable affine transformation before the frozen Laya encoder:
48
-
49
- ```text
50
- input
51
- ↓
52
- Laya embeddings
53
- ↓
54
- Ariadne linear interface
55
- ↓
56
- frozen Laya encoder
57
- ↓
58
- existing decision pathway
59
- ```
60
-
61
- The interface is deliberately minimal:
62
-
63
- - one 1024 × 1024 linear transformation
64
- - one 1024-dimensional bias
65
- - exact identity initialization
66
- - **1,049,600 trainable parameters**
67
- - no LoRA matrices
68
- - no additional transformer blocks
69
- - no modification of the Laya encoder during interface training
70
-
71
- At initialization, Ariadne is exactly equivalent to the unmodified base model.
72
-
73
- Training therefore learns only how to transform the representation presented to the existing frozen network.
74
-
75
- ## Base model
76
-
77
- Ariadne-Laya-TD is an adapter for:
78
-
79
- `convaiinnovations/laya`
80
-
81
- The underlying Laya base model remains frozen during Ariadne training.
82
-
83
- The principal comparison model is the official Laya typed-decisions specialist:
84
 
85
- `convaiinnovations/laya-typed-decisions`
86
 
87
- That specialist is obtained through conventional full-model specialization and serves as the reference for how much typed-decisions capability can be recovered without modifying the base encoder.
88
 
89
- ## Typed-decisions result
90
-
91
- The primary experiment trains Ariadne on the `LocalLLaMA/typed-decisions` training data while keeping the entire Laya base model frozen.
92
-
93
- Only the 1.05M-parameter interface is updated.
94
 
95
- ### Preliminary deterministic results
96
 
97
- The deterministic multi-seed sweep is currently being completed.
 
 
 
 
 
 
 
 
 
98
 
99
- Individual runs have already demonstrated typed-decisions performance at or above the official Laya specialist's approximately **76.7%** reference accuracy while leaving the underlying encoder frozen.
100
 
101
- One deterministic run reached:
102
 
103
- | Model | Trainable parameters | Typed-decisions accuracy |
104
  |---|---:|---:|
105
- | Official Laya typed-decisions specialist | ~421M during specialization | 76.70% |
106
- | **Ariadne on frozen Laya base** | **1,049,600** | **76.95%** |
107
-
108
- This is currently a **single-run result**, not the final multi-seed mean.
109
-
110
- Ariadne training has shown significant convergence sensitivity between runs. The final five-seed deterministic statistics will replace this preliminary result before the benchmark is treated as stable.
111
-
112
- ## What this result means
113
-
114
- The goal is not simply to reduce the number of trainable parameters.
115
-
116
- The frozen-base experiment tests whether task-specific discrimination requires rewriting the internal representation of the model at all.
117
-
118
- The working hypothesis is:
119
 
120
- > **The base model already contains much of the discriminative structure required by typed decisions. Ariadne learns how to align the task with that existing structure rather than modifying hundreds of millions of encoder parameters to align the representation with the task.**
121
 
122
- A successful interface therefore does not teach the frozen encoder new contextual computation.
123
 
124
- Instead, it changes the coordinates in which the existing computation receives its input.
 
 
 
 
 
 
125
 
126
- This can be viewed as **representation alignment** rather than conventional representation adaptation.
127
 
128
- ## Post-specialization experiment
129
 
130
- A second experiment applies the same interface concept to the official Laya typed-decisions specialist.
131
-
132
- In this experiment:
133
-
134
- - the specialist itself is frozen
135
- - only the Ariadne interface is trained
136
- - five seeds are evaluated under the same benchmark bundle
137
-
138
- Results:
139
-
140
- | Method | Trainable parameters | Typed-decisions accuracy |
141
  |---|---:|---:|
142
- | Untouched Laya specialist | — | 76.70% |
143
- | **Specialist + Ariadne linear interface** | **1,049,600** | **78.10 ± 0.31%** |
144
- | Residual-interface ablation | 1,049,600 | 78.24 ± 0.75% |
145
- | LayerNorm + Linear ablation | ~1.05M | 77.72 ± 0.53% |
146
- | Continued specialist full fine-tune | 421,029,889 | 78.42 ± 0.75% |
147
-
148
- Values after ± are sample standard deviations across five seeds.
 
149
 
150
- The plain linear interface was the most stable adaptation method tested.
151
 
152
- Continued full-model fine-tuning produced a mean accuracy only **0.32 percentage points higher** than the linear interface while updating roughly **400× more parameters**.
153
 
154
- With only five seeds, this small difference does not establish a reliable ordering between the two methods.
155
-
156
- ## Cross-task retention after specialist alignment
157
-
158
- The post-specialization Ariadne interface was also evaluated on seven additional English benchmarks.
159
-
160
- | Benchmark | Untouched specialist | Specialist + Ariadne |
161
- |---|---:|---:|
162
- | Typed decisions | 76.70% | **78.10 ± 0.31%** |
163
- | AG News | 94.67% | 94.17 ± 0.37% |
164
- | BoolQ | 82.00% | 81.10 ± 0.90% |
165
- | DAIR Emotion | 58.17% | 57.40 ± 0.65% |
166
- | Prompt injections | 67.24% | 67.41 ± 1.87% |
167
- | SST-5 | 48.00% | 49.33 ± 1.67% |
168
- | MASSIVE intent EN | 81.00% | 79.40 ± 1.23% |
169
- | XNLI EN | 86.33% | 88.33 ± 0.53% |
170
 
171
- Across the seven non-target tasks, the mean accuracy change relative to the untouched specialist is approximately **−0.04 percentage points**.
172
-
173
- This does not imply that every capability is preserved identically. Individual tasks move in both directions.
174
-
175
- It does indicate that the gain on typed decisions is not accompanied by a broad average collapse across this evaluation bundle.
 
176
 
177
  ## Training
178
 
179
- ### Frozen-base Ariadne
180
-
181
- The released Ariadne checkpoint is trained from the official Laya base model.
182
-
183
- During training:
184
-
185
- - all Laya parameters remain frozen
186
- - the frozen encoder is run in evaluation mode
187
- - only the Ariadne interface is trainable
188
- - initialization is exact identity
189
- - training uses soft typed-decision targets
190
- - checkpoint selection uses validation soft cross-entropy
191
-
192
- Current training experiments use deterministic execution settings to better separate seed effects from GPU numerical nondeterminism.
193
-
194
- The final training configuration and five-seed statistics will be added once the current deterministic sweep is complete.
195
-
196
- ### Post-specialist Ariadne experiment
197
-
198
- The completed post-specialist sweep used:
199
 
200
- - trainable parameters: 1,049,600
201
- - effective batch size: 32
202
- - optimizer: AdamW
203
- - learning rate: `3e-4`
204
- - weight decay: `0.01`
205
- - gradient clipping: `1.0`
206
- - objective: soft-target decision cross-entropy
207
- - seeds: 0, 1, 2, 3, 4
208
- - checkpoint selection: lowest validation soft cross-entropy
209
 
210
- ## Dataset
211
-
212
- Ariadne is trained on:
213
-
214
- `LocalLLaMA/typed-decisions`
215
-
216
- The primary benchmark contains:
217
 
218
- - **400 test cases**
219
- - **2,000 individual decisions**
220
 
221
- Training and validation use data derived from the public training split.
222
 
223
- The official test split is excluded from training and checkpoint selection.
224
 
225
- The test set was inspected during development, so results should be interpreted as research benchmark results rather than performance on a completely untouched external test set.
226
 
227
- ## Evaluation
228
 
229
- Accuracy counts invalid or missing decisions as incorrect.
230
 
231
- Probability-based metrics are calculated over valid decisions.
232
 
233
- The evaluation harness tracks:
234
 
235
- - hard accuracy
236
- - soft accuracy
237
- - soft cross-entropy
238
- - hard Brier score
239
- - soft Brier score
240
- - expected calibration error (ECE)
241
 
242
- ECE uses the predicted option probability rather than a separate provider confidence field.
243
 
244
- Brier scores sum over classes.
 
245
 
246
- Evaluation bundle SHA-256:
247
-
248
- ```text
249
- 10fb671245cdac1ee09e9c4d2b9b09829fbf1abd06e0954fddee0b9ec8519854
 
 
 
 
250
  ```
251
 
252
- ## Calibration
253
-
254
- Hard decision accuracy and probability calibration should be treated separately.
255
-
256
- The upstream Laya checkpoint may emit warnings for stored temperature values outside the runtime-supported calibration range. Laya clamps affected values and warns that confidence for those entries should be treated as uncalibrated.
257
-
258
- Ariadne's improvements in hard accuracy should therefore **not** be interpreted as evidence of improved calibration.
259
 
260
- Soft cross-entropy, Brier scores and ECE are reported separately where available.
261
 
262
- ## Why "Ariadne"?
263
 
264
- Ariadne is named after the thread through the labyrinth.
265
-
266
- The metaphor reflects the central idea of the method: the useful representational structure may already exist inside the model; the interface learns a path through that structure rather than reconstructing it.
267
-
268
- In practical terms:
269
-
270
- > **Conventional fine-tuning moves the representation toward the task. Ariadne attempts to move the task toward the representation.**
271
-
272
- ## Intended use
273
-
274
- Ariadne is intended for research and experimentation involving:
275
-
276
- - representation alignment
277
- - frozen encoders
278
- - parameter-efficient adaptation
279
- - System-1 and selector models
280
- - typed decision systems
281
- - reuse of pretrained representations
282
- - comparison with conventional full fine-tuning
283
- - comparison with PEFT methods such as LoRA and adapters
284
-
285
- The approach may be particularly useful when many specialist tasks need to share a large common frozen encoder.
286
-
287
- ## Limitations
288
-
289
- - The current release focuses on English typed decisions.
290
- - The primary benchmark consists of synthetic decision workflows.
291
- - Frozen-base training currently shows significant convergence sensitivity.
292
- - Individual high-performing runs should not be confused with multi-seed average performance.
293
- - The deterministic frozen-base five-seed sweep is still being completed.
294
- - The typed-decisions test set was inspected during development.
295
- - Cross-task results should not be assumed to generalize to arbitrary domains.
296
- - Hard accuracy does not imply good probability calibration.
297
- - Ariadne depends on the upstream Laya model and is not a standalone foundation model.
298
- - The model should not be used for safety-critical decisions without independent task-specific validation.
299
-
300
- ## Reproducibility
301
-
302
- The experimental artifacts include:
303
-
304
- - deterministic and non-deterministic training histories
305
- - per-seed validation metrics
306
- - per-task predictions
307
- - checkpoint hashes
308
- - frozen-parameter verification
309
- - machine-readable benchmark reports
310
- - per-seed/per-task CSV results
311
-
312
- The test suite verifies that interface-only experiments do not alter the frozen upstream model parameters.
313
-
314
- The current audited experimental suite includes **43 passing tests**.
315
-
316
- ## Relationship to Laya
317
 
318
- Ariadne is independently developed and is not an official Convai Innovations model.
319
 
320
- The upstream Laya models and code are released under Apache-2.0.
 
 
 
 
 
 
321
 
322
- Ariadne's interface weights and accompanying code are also released under Apache-2.0.
 
 
 
323
 
324
- ## Citation
325
 
326
- A technical report describing Ariadne and the representation-alignment experiments is in preparation.
327
 
328
- Until then, please cite this Hugging Face repository and the upstream Laya project when using Ariadne.
 
2
  license: apache-2.0
3
  language:
4
  - en
 
 
 
5
  pipeline_tag: text-classification
6
+ base_model: convaiinnovations/laya
7
+ base_model_relation: adapter
8
+ datasets:
9
+ - LocalLLaMA/typed-decisions
10
  tags:
11
  - laya
12
+ - modernbert
 
 
13
  - adapter
14
+ - typed-decisions
15
+ - custom-code
 
 
 
 
 
 
16
  ---
 
17
 
18
+ # Ariadne-Laya-TD — stable training with a small input adapter
19
 
20
+ **A small typed-decision adapter for frozen Laya, with a more stable training recipe.**
21
 
22
+ A 1,049,600-parameter affine input adapter trained on typed decisions while every parameter of base Laya stays frozen. It requires the pinned base model; the small adapter file is not a standalone model.
 
 
23
 
24
+ The included checkpoint scores **75.45%** on the local Typed Decisions test, compared with **36.25%** for the matched base. It is seed 4, epoch 8, selected by validation loss. This release uses LR 0.0001.
25
 
26
+ Across five deterministic runs, typed-decision accuracy was **75.28% ± 0.93 percentage points** (sample SD). The matched native base scored **36.25%**. The table below reports all eight tasks and every seed.
27
 
28
+ **Lower learning rate improved training stability.** Changing LR from 3e-4 to 1e-4 raised mean typed-decision accuracy from 70.23% to 75.28% and reduced the across-seed standard deviation from 6.68 to 0.93 points. Sample variance fell by 98.1% in this five-seed comparison. The complete learning-rate comparison is below.
29
 
30
+ **Consistency across seeds.** Typed-decision accuracy ranged from 74.25% to 76.55%. AG News stayed between 93.67% and 95.17%, and MASSIVE between 75.00% and 79.00%. These are measurements of this recipe on five seeds and the fixed benchmark bundle.
 
 
 
31
 
32
  ## Architecture
33
 
34
+ The map is `x′ = Wx + b`, initialized with `W = I`, `b = 0`, immediately after Laya's original token embedding and normalization. Original embeddings, normalization, all 28 ModernBERT blocks, and decision heads remain frozen and in evaluation mode. Gradients pass through the frozen stack to update only the map. There is no added LLM and no extra BERT forward pass.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
35
 
36
+ The combined model has 422,343,427 parameters; 1,049,600 were trainable (about 0.25%). Each FP32 adapter checkpoint is approximately 4.2 MB. The original model and tokenizer are downloaded separately.
37
 
38
+ Base: [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya/tree/55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851), revision `55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851`. Source weights were hashed before training and at checkpoint saves; all pretrained tensors remained unchanged.
39
 
40
+ ## Evaluation
 
 
 
 
41
 
42
+ All scores below are local measurements on identical frozen cases. Accuracy is the highest-probability label accuracy, including ordinal questions; failed decisions count as incorrect. Values are percentages; variance is sample variance in percentage-points squared. Five seeds are used for the mean and SD. These are not the upstream model card's historical benchmark numbers.
43
 
44
+ | Benchmark | Decisions | Native base | Adapter mean ± SD | Variance | Selected checkpoint |
45
+ |---|---:|---:|---:|---:|---:|
46
+ | Typed decisions | 2000 | 36.25 | 75.28 ± 0.93 | 0.8620 | 75.45 |
47
+ | AG News | 600 | 94.83 | 94.47 ± 0.58 | 0.3389 | 95.17 |
48
+ | BoolQ | 600 | 83.00 | 82.10 ± 0.89 | 0.7861 | 83.17 |
49
+ | DAIR Emotion | 600 | 57.33 | 57.20 ± 1.00 | 1.0056 | 57.83 |
50
+ | Prompt injections | 116 | 68.97 | 74.48 ± 3.93 | 15.4578 | 74.14 |
51
+ | SST-5 | 600 | 37.00 | 41.00 ± 2.46 | 6.0278 | 41.00 |
52
+ | MASSIVE intent EN | 300 | 78.67 | 76.80 ± 1.56 | 2.4222 | 77.00 |
53
+ | XNLI EN | 300 | 86.00 | 87.27 ± 1.04 | 1.0778 | 86.33 |
54
 
55
+ ### What changed from the first release
56
 
57
+ The earlier adapter used LR 3e-4. This version uses 1e-4 with the same seeds, data, initialization, batch, stopping rule and deterministic runtime. The table compares all five runs of each recipe. Different stopping epochs are a consequence of the shared validation rule.
58
 
59
+ | Benchmark | Earlier LR 3e-4 mean ± SD | This LR 1e-4 mean ± SD |
60
  |---|---:|---:|
61
+ | Typed decisions | 70.23 ± 6.68 | 75.28 ± 0.93 |
62
+ | AG News | 54.57 ± 35.43 | 94.47 ± 0.58 |
63
+ | BoolQ | 63.27 ± 16.02 | 82.10 ± 0.89 |
64
+ | DAIR Emotion | 30.33 ± 24.82 | 57.20 ± 1.00 |
65
+ | Prompt injections | 57.76 ± 15.54 | 74.48 ± 3.93 |
66
+ | SST-5 | 30.53 ± 13.61 | 41.00 ± 2.46 |
67
+ | MASSIVE intent EN | 32.53 ± 39.58 | 76.80 ± 1.56 |
68
+ | XNLI EN | 55.53 ± 29.33 | 87.27 ± 1.04 |
 
 
 
 
 
 
69
 
70
+ The earlier published checkpoint was seed 1, epoch 14. The new validation-selected checkpoint improves AG News from 93.17% to 95.17%, BoolQ from 79.83% to 83.17%, and MASSIVE from 74.33% to 77.00%. Its typed-decision score is 75.45%, compared with 76.95% for the earlier checkpoint. The five-seed means above measure the improvement in the training recipe.
71
 
72
+ ### Individual seeds
73
 
74
+ | Seed | Typed | AG News | BoolQ | Emotion | Injection | SST-5 | MASSIVE | XNLI | Best epoch | Stop epoch |
75
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
76
+ | 0 | 74.25 | 94.17 | 80.83 | 58.50 | 69.83 | 38.17 | 79.00 | 89.00 | 11 | 14 |
77
+ | 1 | 75.65 | 94.83 | 82.67 | 57.17 | 78.45 | 44.83 | 75.67 | 87.00 | 9 | 12 |
78
+ | 2 | 74.50 | 94.50 | 81.83 | 56.00 | 71.55 | 39.83 | 75.00 | 87.33 | 13 | 16 |
79
+ | 3 | 76.55 | 93.67 | 82.00 | 56.50 | 78.45 | 41.17 | 77.33 | 86.67 | 8 | 11 |
80
+ | 4 | 75.45 | 95.17 | 83.17 | 57.83 | 74.14 | 41.00 | 77.00 | 86.33 | 8 | 11 |
81
 
82
+ ### Selected checkpoint versus base: paired uncertainty
83
 
84
+ These descriptive 95% intervals resample cases 10,000 times, keeping each typed case's five decisions together. They describe benchmark-case uncertainty for the selected model, separately from the across-seed SD above. They have no multiple-comparison correction and were not used to select the checkpoint.
85
 
86
+ | Benchmark | Change from base (pp) | Paired 95% interval (pp) |
 
 
 
 
 
 
 
 
 
 
87
  |---|---:|---:|
88
+ | Typed decisions | +39.20 | [+35.90, +42.50] |
89
+ | AG News | +0.33 | [-0.50, +1.33] |
90
+ | BoolQ | +0.17 | [-1.50, +1.83] |
91
+ | DAIR Emotion | +0.50 | [-1.50, +2.67] |
92
+ | Prompt injections | +5.17 | [+1.72, +9.48] |
93
+ | SST-5 | +4.00 | [+1.00, +6.83] |
94
+ | MASSIVE intent EN | -1.67 | [-4.33, +1.00] |
95
+ | XNLI EN | +0.33 | [-1.67, +2.33] |
96
 
97
+ ### Checkpoint selection
98
 
99
+ The default package contains **seed 4, epoch 8**, selected by the lowest validation soft cross-entropy across the five seeds (0.854020). This selection rule was recorded before the sweep. Test performance was not used to choose the checkpoint. Only the selected adapter weights are included; all five seeds’ results are reported.
100
 
101
+ ### Benchmark scope
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
102
 
103
+ - Typed Decisions: 400 test cases / 2,000 decisions from four synthetic workflows. Training and validation case IDs are disjoint from test IDs. The test task itself is the adaptation task.
104
+ - AG News, BoolQ, Emotion and SST-5: fixed prefixes of 600 examples each. Prompt injections: all 116 test examples. MASSIVE and XNLI: first 300 English test examples each.
105
+ - MASSIVE uses the gold intent plus 19 seeded distractors (20 options), not all intents. XNLI uses three labels.
106
+ - These seven retention tasks were excluded from adapter training; some were present in the base model's training. These are held-out examples for this adaptation, not proof of wholly unseen-task generalization.
107
+ - The test suite was inspected during earlier exploratory experiments. This recipe was selected adaptively after earlier results; seed 0 began as the lower-LR pilot and was retained for the five-seed confirmation. Checkpoint selection used validation only, but the research process is not a blind test. Earlier runs with changing stopping rules are excluded from the main table.
108
+ - No downstream deployment, multilingual retention, or broad safety claim is established by these small English benchmark subsets.
109
 
110
  ## Training
111
 
112
+ Training data: [LocalLLaMA/typed-decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions), revision `f7a2487edd7a043a5441a5e9ccc7fe5ddbd9ebe8`. A case-level split with seed 42 yields 1,080 training cases (5,400 decisions) and 120 validation cases (600 decisions). Questions from one case never cross splits.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
113
 
114
+ - Seeds 0, 1, 2, 3, 4; RTX PRO 6000 Blackwell only, sequential runs.
115
+ - Every map starts at the same identity transformation and native dropout stays off. Seeds control training-example shuffling; this comparison measures sensitivity to data order with deterministic computation.
116
+ - AdamW, LR 0.0001, weight decay 0.01, gradient clipping at norm 1; effective batch 32 (microbatch 8 × accumulation 4). No learning-rate warmup or scheduler.
117
+ - Soft-target cross-entropy on raw logits. The pretrained stack stays in eval mode; its dropout is off. The affine map is computed in FP32; the remaining CUDA forward uses BF16 autocast.
118
+ - Minimum 10 completed epochs; stop after three consecutive epochs without a strictly lower validation loss. Patience counts from epoch 1, with stopping gated by the minimum. No fixed maximum. Keep the best trained checkpoint; epoch zero is a separate baseline.
119
+ - Input/head token budgets: 1024/256, shared by native baseline and adapter. These differ from the original base config's defaults, so this card reports matched-budget local baselines.
 
 
 
120
 
121
+ ## Reproducibility
 
 
 
 
 
 
122
 
123
+ Strict deterministic algorithms were enabled with errors on unsupported operations, cuDNN benchmarking and cuDNN SDPA were disabled, and float32 matmul precision was set to highest. Python, NumPy and PyTorch seeds were set. Processes used `PYTHONHASHSEED=0` and `CUBLAS_WORKSPACE_CONFIG=:4096:8`.
 
124
 
125
+ Two independent two-epoch seed-0 verification runs (338 optimizer updates each) matched losses, validation metrics, and checkpoint tensors bit for bit. The reported seed-0 run matched that two-epoch metric prefix. This verifies the checked environment and prefix; it does not promise bitwise agreement across GPUs, libraries or operating systems. See `environment.json` and `evidence/replay-check.json`.
126
 
127
+ Before training, the identity adapter matched every native-base answer object on all 5,116 benchmark decisions. All five final runs were independently rescored and checked for unchanged pretrained weights. Export verification is recorded separately in `export-verification.json` after testing the portable package.
128
 
129
+ ## Confidence and temperature warning
130
 
131
+ The pinned base checkpoint ships `choice:11+` temperature 0.1005828. Laya 0.3.20 clamps it to 0.5 and emits a warning on load. Both baseline and adapters use that same behavior. Only the 300 MASSIVE examples use this bucket in this suite. These temperatures are not used in the training loss.
132
 
133
+ Temperature scaling preserves class order mathematically but changes probability sharpness and calibration metrics. Calibration was not refitted after adaptation. Confidence calibration should be evaluated for the intended application; the base model's calibration claims have not been established for this adapter.
134
 
135
+ ## Load locally
136
 
137
+ This is a custom adapter. Use the bundled loader: `laya.load(repo_id)` and `AutoModel.from_pretrained(repo_id)` do not install this input transformation.
138
 
139
+ Download the repository files with the Hugging Face UI or `hf download GoatHerder/Ariadne-Laya-TD --local-dir ariadne-laya-td`, then change into that directory. The learning rate and selected seed are recorded in `training_protocol.json` and `adapter_config.json`.
 
 
 
 
 
140
 
141
+ Install a PyTorch build for your GPU and the versions in `requirements.txt`. The measured environment used PyTorch 2.13.0+cu130; package versions and CUDA/cuDNN details are in `environment.json`. Run from this downloaded repository directory:
142
 
143
+ ```python
144
+ from load_adapter import load
145
 
146
+ model = load(device="cuda:0") # defaults to the validation-selected seed
147
+ answer = model.predict("The customer says the delivery arrived damaged.", {
148
+ "route": {"type": "choice", "instructions": "Choose the support queue.",
149
+ "criteria": {"delivery": "Delivery and damaged items",
150
+ "billing": "Payments and invoices"}}
151
+ })
152
+ print(answer)
153
+ model.close()
154
  ```
155
 
156
+ Only the validation-selected seed is included in this compact package. Use `device='cpu'` for CPU inference. The loader downloads the pinned base, verifies base configuration/tokenizer hashes and model tensors, checks adapter weights, then freezes all parameters for inference. `local_files_only=True` uses cached weights; `base_path` can point to an existing copy of the pinned base.
 
 
 
 
 
 
157
 
158
+ ## Reproduce preparation and one training run
159
 
160
+ The bundled `ariadne_bench` source contains the actual tokenizer/prompt preparation, training loop, adapters and scorer. Dataset files are downloaded from their pinned sources; evaluation data is not redistributed.
161
 
162
+ ```bash
163
+ python -m ariadne_bench.experiments.prepare_alignment --out data/alignment
164
+ python -m ariadne_bench prepare --profile laya-core \
165
+ --suites typed-decisions,ag-news,emotion,boolq,sst5,prompt-injections,massive-intent.en,xnli.en \
166
+ --out data/heldout-eight
167
+ ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
168
 
169
+ The expected case-file hashes and immutable dataset revisions are in `training_protocol.json` and `benchmarks/sources.lock.json`. `make_training_config.py` resolves the pinned base into a local training config:
170
 
171
+ ```bash
172
+ python make_training_config.py --seed 0 --device cuda:0 --out train-config.json
173
+ PYTHONHASHSEED=0 CUBLAS_WORKSPACE_CONFIG=:4096:8 USE_TF=0 OMP_NUM_THREADS=4 \
174
+ python -m ariadne_bench.experiments.align --config train-config.json \
175
+ --train data/alignment/train --validation data/alignment/validation \
176
+ --out training-seed-0 --epochs 0 --min-epochs 10 --early-stopping-patience 3 \
177
+ --batch-size 8 --accumulation 4 --lr 0.0001 --keep-best-only
178
 
179
+ PYTHONHASHSEED=0 CUBLAS_WORKSPACE_CONFIG=:4096:8 USE_TF=0 OMP_NUM_THREADS=4 \
180
+ python -m ariadne_bench run --data data/heldout-eight \
181
+ --config training-seed-0/benchmark-config.json --out evaluation-seed-0 --warmup 5
182
+ ```
183
 
184
+ ## License and attribution
185
 
186
+ Apache-2.0. The [base-model card](https://huggingface.co/convaiinnovations/laya/blob/55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851/README.md) and [Typed Decisions dataset card](https://huggingface.co/datasets/LocalLLaMA/typed-decisions/blob/f7a2487edd7a043a5441a5e9ccc7fe5ddbd9ebe8/README.md) list Apache-2.0. Base weights are downloaded from their original repository. Evaluation datasets retain their own licenses. Prompt preparation credits and upstream notices are in `THIRD_PARTY.md` and `licenses/`.
187
 
188
+ This is an independent adaptation of Laya. The upstream authors did not produce these adapter weights or experimental results.
USAGE.md CHANGED
@@ -1,48 +1,3 @@
1
- # Load the Laya linear adapter
2
 
3
- This package contains the validation-selected **seed 1, epoch 14** checkpoint from the deterministic frozen-base experiment. It trained 1,049,600 affine parameters at LR 3e-4 while all base Laya weights remained frozen. Its 4.2 MB weight file requires the pinned base model, which the loader downloads separately.
4
-
5
- The ongoing lower-learning-rate experiment is separate. These are the already verified weights, not a claim that the training stability investigation is complete.
6
-
7
- ## Upload
8
-
9
- Extract the ZIP and upload the files and folders **at the root of your existing Hugging Face model repository**. There is deliberately no README.md in this bundle, so it can coexist with your existing model card. The ZIP is a transport archive; upload its extracted contents for normal use.
10
-
11
- This is a custom input adapter. Loading requires the bundled loader; bare `laya.load(repo_id)` and `AutoModel.from_pretrained(repo_id)` do not install the extra affine layer.
12
-
13
- ## Load a downloaded copy
14
-
15
- Install an appropriate PyTorch build and the versions in `requirements.txt`. The measured GPU environment is recorded in `environment.json`. From the downloaded repository directory:
16
-
17
- ```python
18
- from load_adapter import load
19
-
20
- model = load(device="cuda:0")
21
- answer = model.predict("The delivery arrived damaged.", {
22
- "route": {
23
- "type": "choice",
24
- "instructions": "Choose the customer-support queue.",
25
- "criteria": {"delivery": "Delivery and damaged items", "billing": "Payments and invoices"},
26
- }
27
- })
28
- print(answer)
29
- model.close()
30
- ```
31
-
32
- After you upload, first download your repository with `huggingface_hub.snapshot_download("YOUR_ACCOUNT/YOUR_REPOSITORY")`, then run the example from that downloaded directory. Use `device="cpu"` for CPU inference. `base_path` can point to the pinned base snapshot already on disk; `local_files_only=True` uses cached base files.
33
-
34
- The loader verifies the adapter checksum, base configuration/tokenizer hashes and pretrained tensor hash. Strict determinism is enabled by default and requires the CUBLAS workspace environment to be set before CUDA is initialized. A fresh process handles this automatically. Numerical agreement across other hardware or software is not guaranteed.
35
-
36
- ## Measured performance and limits
37
-
38
- Selected checkpoint accuracy: typed decisions **76.95%**, AG News **93.17%**, BoolQ **79.83%**, Emotion **57.33%**, prompt injections **71.55%**, SST-5 **42.00%**, MASSIVE EN **74.33%**, XNLI EN **87.67%**. Native base typed accuracy was 36.25% under the matched protocol.
39
-
40
- Across all five training seeds, typed accuracy was **70.23% ± 6.68 pp** (sample SD). Three seeds had severe retention losses. The selected model also lost 3.17 pp on BoolQ and 4.33 pp on MASSIVE versus base. Full results and paired uncertainty are in `metrics.json` and `evidence/candidate-uncertainty.json`. Other seeds' scores are included for transparency; this bundle contains only seed 1's weights.
41
-
42
- The benchmark uses 2,000 typed decisions plus fixed subsets of seven other tasks, totaling 5,116 decisions. MASSIVE uses 20 candidate intents. The same test subsets were inspected in earlier experiments. These results describe an experimental task adapter, not established broad task improvement.
43
-
44
- The base checkpoint's choice:11+ temperature is clamped from 0.1005828 to 0.5 by Laya 0.3.20. This expected load warning concerns confidence calibration; raw-logit training and class ordering are unaffected. Calibration was not refitted after adaptation.
45
-
46
- All 5,116 answer objects matched the original selected checkpoint when the portable export was tested from outside the workspace. `export-verification.json` records that check. `upload-manifest.json` verifies that this bundle keeps the same weights and inference code, with the other seeds removed from its loading configuration.
47
-
48
- License: Apache-2.0. Base model: `convaiinnovations/laya`, pinned to revision `55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851`. Attribution is in `THIRD_PARTY.md` and `licenses/`.
 
1
+ # Using this adapter
2
 
3
+ This package uses LR 0.0001, seed 4, epoch 8. See [the model card](README.md#load-locally) for the loading example, all five seeds' results, requirements and limitations. Load with the bundled `load_adapter.py`; the adapter weights require the pinned base Laya.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
adapter.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e5ec5e1d494509f63960b6949af15be8390442271805ba45d796a1911b5c239a
3
  size 4198616
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081
3
  size 4198616
adapter_config.json CHANGED
@@ -4,7 +4,7 @@
4
  "position": "after native embeddings, before ModernBERT block 0",
5
  "base_model": "convaiinnovations/laya",
6
  "base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
7
- "candidate_seed": 1,
8
  "max_len": 1024,
9
  "head_max_len": 256,
10
  "trainable_parameters_during_training": 1049600,
@@ -16,10 +16,10 @@
16
  "tokenizer/tokenizer_config.json": "50044de60daaa73df97d262e15a40d4faf0160e7d742df64b377877a1320dd12"
17
  },
18
  "seeds": {
19
- "1": {
20
  "weights_file": "adapter.safetensors",
21
- "weights_sha256": "e5ec5e1d494509f63960b6949af15be8390442271805ba45d796a1911b5c239a",
22
- "selected_epoch": 14
23
  }
24
  }
25
  }
 
4
  "position": "after native embeddings, before ModernBERT block 0",
5
  "base_model": "convaiinnovations/laya",
6
  "base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
7
+ "candidate_seed": 4,
8
  "max_len": 1024,
9
  "head_max_len": 256,
10
  "trainable_parameters_during_training": 1049600,
 
16
  "tokenizer/tokenizer_config.json": "50044de60daaa73df97d262e15a40d4faf0160e7d742df64b377877a1320dd12"
17
  },
18
  "seeds": {
19
+ "4": {
20
  "weights_file": "adapter.safetensors",
21
+ "weights_sha256": "f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081",
22
+ "selected_epoch": 8
23
  }
24
  }
25
  }
evidence/adapter-equivalence.json CHANGED
@@ -35,5 +35,10 @@
35
  "forward_exact": true,
36
  "exact_match": true
37
  }
38
- ]
 
 
 
 
 
39
  }
 
35
  "forward_exact": true,
36
  "exact_match": true
37
  }
38
+ ],
39
+ "reuse_provenance": {
40
+ "experiment": "E0i",
41
+ "source_sha256": "c77faa691bd273d363dbb196155545b4f72697b9d0f4c80451ffb06b0090d200",
42
+ "reason": "unchanged adapter computation and strict runtime; this forward/backward diagnostic has no optimizer step or learning-rate dependence"
43
+ }
44
  }
evidence/audit.json CHANGED
@@ -5,52 +5,52 @@
5
  "runs": [
6
  {
7
  "seed": 0,
8
- "checkpoint_sha256": "bcb37755de54782bc4d1f79615539f48940aedb2b733dadc1033fba2f2304cf8",
9
  "checkpoint_bytes": 4198616,
10
  "all_pretrained_tensors_unchanged": true,
11
  "independent_scoring_matches": true,
12
  "min_epochs": 10,
13
- "stop_epoch": 23,
14
  "minimum_epoch_rule_verified": true
15
  },
16
  {
17
  "seed": 1,
18
- "checkpoint_sha256": "e5ec5e1d494509f63960b6949af15be8390442271805ba45d796a1911b5c239a",
19
  "checkpoint_bytes": 4198616,
20
  "all_pretrained_tensors_unchanged": true,
21
  "independent_scoring_matches": true,
22
  "min_epochs": 10,
23
- "stop_epoch": 17,
24
  "minimum_epoch_rule_verified": true
25
  },
26
  {
27
  "seed": 2,
28
- "checkpoint_sha256": "61391c3869c1f4a8449444db6a9e3b660692eefa25e3c149cee0d6af5daf264c",
29
  "checkpoint_bytes": 4198616,
30
  "all_pretrained_tensors_unchanged": true,
31
  "independent_scoring_matches": true,
32
  "min_epochs": 10,
33
- "stop_epoch": 18,
34
  "minimum_epoch_rule_verified": true
35
  },
36
  {
37
  "seed": 3,
38
- "checkpoint_sha256": "26ed2443b8ea146d3fab98e94ab55356557b9ef10ac5b8ba656b1a37cf4f0f64",
39
  "checkpoint_bytes": 4198616,
40
  "all_pretrained_tensors_unchanged": true,
41
  "independent_scoring_matches": true,
42
  "min_epochs": 10,
43
- "stop_epoch": 27,
44
  "minimum_epoch_rule_verified": true
45
  },
46
  {
47
  "seed": 4,
48
- "checkpoint_sha256": "603a83ca4996e32b4753250ea6b07856e7cc5196b41e635cdad35ee4270eca38",
49
  "checkpoint_bytes": 4198616,
50
  "all_pretrained_tensors_unchanged": true,
51
  "independent_scoring_matches": true,
52
  "min_epochs": 10,
53
- "stop_epoch": 10,
54
  "minimum_epoch_rule_verified": true
55
  }
56
  ],
 
5
  "runs": [
6
  {
7
  "seed": 0,
8
+ "checkpoint_sha256": "4794ce39ed5fe621ef8deb7e40900590951f7382e9c7e15543f8d0ef4b41df6b",
9
  "checkpoint_bytes": 4198616,
10
  "all_pretrained_tensors_unchanged": true,
11
  "independent_scoring_matches": true,
12
  "min_epochs": 10,
13
+ "stop_epoch": 14,
14
  "minimum_epoch_rule_verified": true
15
  },
16
  {
17
  "seed": 1,
18
+ "checkpoint_sha256": "0441f00f59c24e9b6efe2e01e71fefc31ca5e4b0da70ed2a5a2051a385782ddc",
19
  "checkpoint_bytes": 4198616,
20
  "all_pretrained_tensors_unchanged": true,
21
  "independent_scoring_matches": true,
22
  "min_epochs": 10,
23
+ "stop_epoch": 12,
24
  "minimum_epoch_rule_verified": true
25
  },
26
  {
27
  "seed": 2,
28
+ "checkpoint_sha256": "1a4cebb05017866daa8f339671599ac21ff3125fcbfac8605845813ef540c007",
29
  "checkpoint_bytes": 4198616,
30
  "all_pretrained_tensors_unchanged": true,
31
  "independent_scoring_matches": true,
32
  "min_epochs": 10,
33
+ "stop_epoch": 16,
34
  "minimum_epoch_rule_verified": true
35
  },
36
  {
37
  "seed": 3,
38
+ "checkpoint_sha256": "1ce28c37f42b247b9bdb51d126bc189161f33d8ff28990b922a3041e5a3ff911",
39
  "checkpoint_bytes": 4198616,
40
  "all_pretrained_tensors_unchanged": true,
41
  "independent_scoring_matches": true,
42
  "min_epochs": 10,
43
+ "stop_epoch": 11,
44
  "minimum_epoch_rule_verified": true
45
  },
46
  {
47
  "seed": 4,
48
+ "checkpoint_sha256": "f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081",
49
  "checkpoint_bytes": 4198616,
50
  "all_pretrained_tensors_unchanged": true,
51
  "independent_scoring_matches": true,
52
  "min_epochs": 10,
53
+ "stop_epoch": 11,
54
  "minimum_epoch_rule_verified": true
55
  }
56
  ],
evidence/candidate-uncertainty.json CHANGED
@@ -1,6 +1,6 @@
1
  {
2
- "candidate_seed": 1,
3
- "selected_epoch": 14,
4
  "benchmark_cases_sha256": "10fb671245cdac1ee09e9c4d2b9b09829fbf1abd06e0954fddee0b9ec8519854",
5
  "resamples": 10000,
6
  "bootstrap_seed_per_suite": 42,
@@ -10,66 +10,66 @@
10
  "suites": {
11
  "typed-decisions": {
12
  "cases": 400,
13
- "accuracy_difference_pp": 40.7,
14
  "case_bootstrap_95_interval_pp": [
15
- 37.3,
16
- 44.101249999999986
17
  ]
18
  },
19
  "ag-news": {
20
  "cases": 600,
21
- "accuracy_difference_pp": -1.6666666666666667,
22
  "case_bootstrap_95_interval_pp": [
23
- -3.5,
24
- 0.16666666666666666
25
  ]
26
  },
27
  "boolq": {
28
  "cases": 600,
29
- "accuracy_difference_pp": -3.1666666666666665,
30
  "case_bootstrap_95_interval_pp": [
31
- -5.833333333333333,
32
- -0.6666666666666666
33
  ]
34
  },
35
  "emotion": {
36
  "cases": 600,
37
- "accuracy_difference_pp": 0.0,
38
  "case_bootstrap_95_interval_pp": [
39
- -2.5,
40
  2.6666666666666665
41
  ]
42
  },
43
  "prompt-injections": {
44
  "cases": 116,
45
- "accuracy_difference_pp": 2.586206896551724,
46
  "case_bootstrap_95_interval_pp": [
47
- -5.172413793103448,
48
- 10.344827586206897
49
  ]
50
  },
51
  "sst5": {
52
  "cases": 600,
53
- "accuracy_difference_pp": 5.0,
54
  "case_bootstrap_95_interval_pp": [
55
- 1.1666666666666667,
56
- 9.0
57
  ]
58
  },
59
  "massive-intent.en": {
60
  "cases": 300,
61
- "accuracy_difference_pp": -4.333333333333333,
62
  "case_bootstrap_95_interval_pp": [
63
- -8.333333333333334,
64
- -0.3333333333333333
65
  ]
66
  },
67
  "xnli.en": {
68
  "cases": 300,
69
- "accuracy_difference_pp": 1.6666666666666667,
70
  "case_bootstrap_95_interval_pp": [
71
- -1.0,
72
- 4.666666666666667
73
  ]
74
  }
75
  }
 
1
  {
2
+ "candidate_seed": 4,
3
+ "selected_epoch": 8,
4
  "benchmark_cases_sha256": "10fb671245cdac1ee09e9c4d2b9b09829fbf1abd06e0954fddee0b9ec8519854",
5
  "resamples": 10000,
6
  "bootstrap_seed_per_suite": 42,
 
10
  "suites": {
11
  "typed-decisions": {
12
  "cases": 400,
13
+ "accuracy_difference_pp": 39.2,
14
  "case_bootstrap_95_interval_pp": [
15
+ 35.9,
16
+ 42.5
17
  ]
18
  },
19
  "ag-news": {
20
  "cases": 600,
21
+ "accuracy_difference_pp": 0.3333333333333333,
22
  "case_bootstrap_95_interval_pp": [
23
+ -0.5,
24
+ 1.3333333333333333
25
  ]
26
  },
27
  "boolq": {
28
  "cases": 600,
29
+ "accuracy_difference_pp": 0.16666666666666666,
30
  "case_bootstrap_95_interval_pp": [
31
+ -1.5,
32
+ 1.8333333333333333
33
  ]
34
  },
35
  "emotion": {
36
  "cases": 600,
37
+ "accuracy_difference_pp": 0.5,
38
  "case_bootstrap_95_interval_pp": [
39
+ -1.5,
40
  2.6666666666666665
41
  ]
42
  },
43
  "prompt-injections": {
44
  "cases": 116,
45
+ "accuracy_difference_pp": 5.172413793103448,
46
  "case_bootstrap_95_interval_pp": [
47
+ 1.7241379310344827,
48
+ 9.482758620689655
49
  ]
50
  },
51
  "sst5": {
52
  "cases": 600,
53
+ "accuracy_difference_pp": 4.0,
54
  "case_bootstrap_95_interval_pp": [
55
+ 1.0,
56
+ 6.833333333333333
57
  ]
58
  },
59
  "massive-intent.en": {
60
  "cases": 300,
61
+ "accuracy_difference_pp": -1.6666666666666667,
62
  "case_bootstrap_95_interval_pp": [
63
+ -4.333333333333333,
64
+ 1.0
65
  ]
66
  },
67
  "xnli.en": {
68
  "cases": 300,
69
+ "accuracy_difference_pp": 0.3333333333333333,
70
  "case_bootstrap_95_interval_pp": [
71
+ -1.6666666666666667,
72
+ 2.3333333333333335
73
  ]
74
  }
75
  }
evidence/lr-comparison.json ADDED
@@ -0,0 +1,315 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "lr_3e4": {
3
+ "typed-decisions": {
4
+ "n": 5,
5
+ "mean_percent": 70.23,
6
+ "sample_variance_pp_squared": 44.56824999999997,
7
+ "sample_std_pp": 6.675945625901995
8
+ },
9
+ "ag-news": {
10
+ "n": 5,
11
+ "mean_percent": 54.56666666666666,
12
+ "sample_variance_pp_squared": 1255.536111111111,
13
+ "sample_std_pp": 35.433544997799906
14
+ },
15
+ "boolq": {
16
+ "n": 5,
17
+ "mean_percent": 63.266666666666666,
18
+ "sample_variance_pp_squared": 256.79999999999995,
19
+ "sample_std_pp": 16.0249804992081
20
+ },
21
+ "emotion": {
22
+ "n": 5,
23
+ "mean_percent": 30.333333333333332,
24
+ "sample_variance_pp_squared": 616.1666666666666,
25
+ "sample_std_pp": 24.82270466058577
26
+ },
27
+ "prompt-injections": {
28
+ "n": 5,
29
+ "mean_percent": 57.758620689655174,
30
+ "sample_variance_pp_squared": 241.5279429250891,
31
+ "sample_std_pp": 15.541169290793055
32
+ },
33
+ "sst5": {
34
+ "n": 5,
35
+ "mean_percent": 30.53333333333333,
36
+ "sample_variance_pp_squared": 185.14444444444447,
37
+ "sample_std_pp": 13.606779356057938
38
+ },
39
+ "massive-intent.en": {
40
+ "n": 5,
41
+ "mean_percent": 32.53333333333333,
42
+ "sample_variance_pp_squared": 1566.1999999999998,
43
+ "sample_std_pp": 39.57524478761944
44
+ },
45
+ "xnli.en": {
46
+ "n": 5,
47
+ "mean_percent": 55.53333333333334,
48
+ "sample_variance_pp_squared": 860.5333333333334,
49
+ "sample_std_pp": 29.33484844571953
50
+ }
51
+ },
52
+ "lr_1e4": {
53
+ "typed-decisions": {
54
+ "n": 5,
55
+ "mean_percent": 75.28,
56
+ "sample_variance_pp_squared": 0.8619999999999957,
57
+ "sample_std_pp": 0.9284395510748105
58
+ },
59
+ "ag-news": {
60
+ "n": 5,
61
+ "mean_percent": 94.46666666666667,
62
+ "sample_variance_pp_squared": 0.3388888888888897,
63
+ "sample_std_pp": 0.5821416398857667
64
+ },
65
+ "boolq": {
66
+ "n": 5,
67
+ "mean_percent": 82.10000000000001,
68
+ "sample_variance_pp_squared": 0.7861111111111168,
69
+ "sample_std_pp": 0.8866290718846956
70
+ },
71
+ "emotion": {
72
+ "n": 5,
73
+ "mean_percent": 57.2,
74
+ "sample_variance_pp_squared": 1.0055555555555546,
75
+ "sample_std_pp": 1.0027739304327543
76
+ },
77
+ "prompt-injections": {
78
+ "n": 5,
79
+ "mean_percent": 74.48275862068965,
80
+ "sample_variance_pp_squared": 15.457788347205712,
81
+ "sample_std_pp": 3.93163939689358
82
+ },
83
+ "sst5": {
84
+ "n": 5,
85
+ "mean_percent": 41.0,
86
+ "sample_variance_pp_squared": 6.027777777777775,
87
+ "sample_std_pp": 2.4551533104427055
88
+ },
89
+ "massive-intent.en": {
90
+ "n": 5,
91
+ "mean_percent": 76.8,
92
+ "sample_variance_pp_squared": 2.422222222222218,
93
+ "sample_std_pp": 1.556349003990499
94
+ },
95
+ "xnli.en": {
96
+ "n": 5,
97
+ "mean_percent": 87.26666666666667,
98
+ "sample_variance_pp_squared": 1.0777777777777784,
99
+ "sample_std_pp": 1.038160766826496
100
+ }
101
+ },
102
+ "paired_seed_accuracy_differences_pp": {
103
+ "0": {
104
+ "typed-decisions": 13.250000000000007,
105
+ "ag-news": 62.83333333333333,
106
+ "emotion": 38.99999999999999,
107
+ "boolq": 24.16666666666667,
108
+ "sst5": 23.166666666666664,
109
+ "prompt-injections": 16.37931034482759,
110
+ "massive-intent.en": 76.66666666666667,
111
+ "xnli.en": 54.666666666666664
112
+ },
113
+ "1": {
114
+ "typed-decisions": -1.3000000000000012,
115
+ "ag-news": 1.6666666666666718,
116
+ "emotion": -0.16666666666667052,
117
+ "boolq": 2.833333333333332,
118
+ "sst5": 2.833333333333332,
119
+ "prompt-injections": 6.896551724137923,
120
+ "massive-intent.en": 1.333333333333342,
121
+ "xnli.en": -0.666666666666671
122
+ },
123
+ "2": {
124
+ "typed-decisions": 8.399999999999997,
125
+ "ag-news": 68.16666666666666,
126
+ "emotion": 51.16666666666667,
127
+ "boolq": 32.0,
128
+ "sst5": 17.166666666666668,
129
+ "prompt-injections": 37.06896551724138,
130
+ "massive-intent.en": 68.66666666666667,
131
+ "xnli.en": 53.0
132
+ },
133
+ "3": {
134
+ "typed-decisions": 5.099999999999993,
135
+ "ag-news": 65.16666666666666,
136
+ "emotion": 42.99999999999999,
137
+ "boolq": 33.166666666666664,
138
+ "sst5": 15.500000000000004,
139
+ "prompt-injections": 21.55172413793103,
140
+ "massive-intent.en": 75.0,
141
+ "xnli.en": 53.0
142
+ },
143
+ "4": {
144
+ "typed-decisions": -0.20000000000000018,
145
+ "ag-news": 1.6666666666666607,
146
+ "emotion": 1.333333333333342,
147
+ "boolq": 2.0000000000000018,
148
+ "sst5": -6.333333333333336,
149
+ "prompt-injections": 1.7241379310344862,
150
+ "massive-intent.en": -0.33333333333332993,
151
+ "xnli.en": -1.333333333333342
152
+ }
153
+ },
154
+ "previous_release": {
155
+ "seed": 1,
156
+ "epoch": 14,
157
+ "interface_lr": 0.0003,
158
+ "suites": {
159
+ "typed-decisions": {
160
+ "attempted": 2000,
161
+ "valid": 2000,
162
+ "failed": 0,
163
+ "coverage": 1.0,
164
+ "accuracy_all": 0.7695,
165
+ "accuracy_valid": 0.7695,
166
+ "ece_top_label": 0.2142237525419539,
167
+ "mean_confidence": 0.555276247458046,
168
+ "brier_hard": 0.3977842163576973,
169
+ "brier_hard_n": 2000,
170
+ "nll_hard": 0.7027635604099817,
171
+ "nll_hard_n": 2000,
172
+ "zero_probability_gold": 0.0,
173
+ "zero_probability_gold_n": 2000,
174
+ "soft_accuracy": 0.47103605907759827,
175
+ "soft_accuracy_n": 2000,
176
+ "brier_soft": 0.06255403710877039,
177
+ "brier_soft_n": 2000,
178
+ "kl_gold_to_prediction": 0.11643450461088396,
179
+ "kl_gold_to_prediction_n": 2000,
180
+ "total_variation": 0.17146599324503523,
181
+ "total_variation_n": 2000,
182
+ "score_mae": 0.2299944493828981,
183
+ "score_mae_n": 800,
184
+ "within_one_level": 0.98875,
185
+ "within_one_level_n": 800,
186
+ "macro_f1": 0.6487623885802034
187
+ },
188
+ "ag-news": {
189
+ "attempted": 600,
190
+ "valid": 600,
191
+ "failed": 0,
192
+ "coverage": 1.0,
193
+ "accuracy_all": 0.9316666666666666,
194
+ "accuracy_valid": 0.9316666666666666,
195
+ "ece_top_label": 0.03387674147756575,
196
+ "mean_confidence": 0.904708562141314,
197
+ "brier_hard": 0.10911433540423408,
198
+ "brier_hard_n": 600,
199
+ "nll_hard": 0.21361121678114303,
200
+ "nll_hard_n": 600,
201
+ "zero_probability_gold": 0.0,
202
+ "zero_probability_gold_n": 600,
203
+ "macro_f1": 0.9283987145646934
204
+ },
205
+ "emotion": {
206
+ "attempted": 600,
207
+ "valid": 600,
208
+ "failed": 0,
209
+ "coverage": 1.0,
210
+ "accuracy_all": 0.5733333333333334,
211
+ "accuracy_valid": 0.5733333333333334,
212
+ "ece_top_label": 0.3169210441851814,
213
+ "mean_confidence": 0.8880587177380197,
214
+ "brier_hard": 0.7264584486728032,
215
+ "brier_hard_n": 600,
216
+ "nll_hard": 2.146154603203373,
217
+ "nll_hard_n": 600,
218
+ "zero_probability_gold": 0.008333333333333333,
219
+ "zero_probability_gold_n": 600,
220
+ "macro_f1": 0.4862286665693279
221
+ },
222
+ "boolq": {
223
+ "attempted": 600,
224
+ "valid": 600,
225
+ "failed": 0,
226
+ "coverage": 1.0,
227
+ "accuracy_all": 0.7983333333333333,
228
+ "accuracy_valid": 0.7983333333333333,
229
+ "ece_top_label": 0.09755433333333334,
230
+ "mean_confidence": 0.8900093333333333,
231
+ "brier_hard": 0.29822958146666667,
232
+ "brier_hard_n": 600,
233
+ "nll_hard": 0.4711253960780479,
234
+ "nll_hard_n": 600,
235
+ "zero_probability_gold": 0.0,
236
+ "zero_probability_gold_n": 600,
237
+ "macro_f1": 0.7813983878883868
238
+ },
239
+ "sst5": {
240
+ "attempted": 600,
241
+ "valid": 600,
242
+ "failed": 0,
243
+ "coverage": 1.0,
244
+ "accuracy_all": 0.42,
245
+ "accuracy_valid": 0.42,
246
+ "ece_top_label": 0.1643168667044709,
247
+ "mean_confidence": 0.5761336745094923,
248
+ "brier_hard": 0.722245716690302,
249
+ "brier_hard_n": 600,
250
+ "nll_hard": 1.412054753482227,
251
+ "nll_hard_n": 600,
252
+ "zero_probability_gold": 0.0,
253
+ "zero_probability_gold_n": 600,
254
+ "score_mae": 0.7360310433394476,
255
+ "score_mae_n": 600,
256
+ "within_one_level": 0.7466666666666667,
257
+ "within_one_level_n": 600,
258
+ "macro_f1": 0.41387920942262574
259
+ },
260
+ "prompt-injections": {
261
+ "attempted": 116,
262
+ "valid": 116,
263
+ "failed": 0,
264
+ "coverage": 1.0,
265
+ "accuracy_all": 0.7155172413793104,
266
+ "accuracy_valid": 0.7155172413793104,
267
+ "ece_top_label": 0.20010517241379305,
268
+ "mean_confidence": 0.9105,
269
+ "brier_hard": 0.43610344172413795,
270
+ "brier_hard_n": 116,
271
+ "nll_hard": 1.2355666268201304,
272
+ "nll_hard_n": 116,
273
+ "zero_probability_gold": 0.008620689655172414,
274
+ "zero_probability_gold_n": 116,
275
+ "macro_f1": 0.7038756091900673
276
+ },
277
+ "massive-intent.en": {
278
+ "attempted": 300,
279
+ "valid": 300,
280
+ "failed": 0,
281
+ "coverage": 1.0,
282
+ "accuracy_all": 0.7433333333333333,
283
+ "accuracy_valid": 0.7433333333333333,
284
+ "ece_top_label": 0.1906540683935104,
285
+ "mean_confidence": 0.9339874017268438,
286
+ "brier_hard": 0.43706654758878133,
287
+ "brier_hard_n": 300,
288
+ "nll_hard": 2.7801978001907903,
289
+ "nll_hard_n": 300,
290
+ "zero_probability_gold": 0.07,
291
+ "zero_probability_gold_n": 300,
292
+ "macro_f1": 0.44346732036749054
293
+ },
294
+ "xnli.en": {
295
+ "attempted": 300,
296
+ "valid": 300,
297
+ "failed": 0,
298
+ "coverage": 1.0,
299
+ "accuracy_all": 0.8766666666666667,
300
+ "accuracy_valid": 0.8766666666666667,
301
+ "ece_top_label": 0.055158283453726184,
302
+ "mean_confidence": 0.913468962679153,
303
+ "brier_hard": 0.19180470102420566,
304
+ "brier_hard_n": 300,
305
+ "nll_hard": 0.3708980570966955,
306
+ "nll_hard_n": 300,
307
+ "zero_probability_gold": 0.0,
308
+ "zero_probability_gold_n": 300,
309
+ "macro_f1": 0.8772028178860477
310
+ }
311
+ },
312
+ "repository_revision": "eaea15162eac93240e5d70f263920eea3093e6c2"
313
+ },
314
+ "scope": "adaptive follow-up on a repeatedly inspected benchmark bundle"
315
+ }
evidence/replay-check.json CHANGED
@@ -1,15 +1,17 @@
1
  {
2
  "passed": true,
3
  "seed": 0,
 
4
  "separate_processes": true,
5
  "epochs_per_process": 2,
6
  "optimizer_updates_per_process": 338,
7
  "initial_validation_exact": true,
8
  "epoch_losses_and_metrics_exact": true,
9
  "selected_checkpoint_tensors_bitwise_equal": true,
 
10
  "checkpoint_sha256": [
11
- "3a30ce9881420cebc61829e52b41a9e669a114b4dffe0f8f836831572bd70331",
12
- "3a30ce9881420cebc61829e52b41a9e669a114b4dffe0f8f836831572bd70331"
13
  ],
14
  "reproducibility": {
15
  "deterministic_algorithms": true,
@@ -28,5 +30,6 @@
28
  "cudnn_version": 92000,
29
  "scope": "fixed hardware/runtime; cross-platform bitwise agreement is not promised"
30
  },
31
- "limit": "replay verified for this data, device, runtime, and two-epoch prefix"
 
32
  }
 
1
  {
2
  "passed": true,
3
  "seed": 0,
4
+ "interface_lr": 0.0001,
5
  "separate_processes": true,
6
  "epochs_per_process": 2,
7
  "optimizer_updates_per_process": 338,
8
  "initial_validation_exact": true,
9
  "epoch_losses_and_metrics_exact": true,
10
  "selected_checkpoint_tensors_bitwise_equal": true,
11
+ "reported_seed_0_two_epoch_metrics_exact": true,
12
  "checkpoint_sha256": [
13
+ "3c52574d070140a55de8929922302c193d4493f76f35a502edf419ee5005881e",
14
+ "3c52574d070140a55de8929922302c193d4493f76f35a502edf419ee5005881e"
15
  ],
16
  "reproducibility": {
17
  "deterministic_algorithms": true,
 
30
  "cudnn_version": 92000,
31
  "scope": "fixed hardware/runtime; cross-platform bitwise agreement is not promised"
32
  },
33
+ "timing": "verification after the completed five-seed sweep",
34
+ "limit": "same runtime, hardware, data and two-epoch prefix; not a full-trajectory or cross-platform replay"
35
  }
evidence/seed-0-history.json CHANGED
@@ -9,399 +9,246 @@
9
  "epochs": [
10
  {
11
  "epoch": 1,
12
- "train_soft_cross_entropy": 1.1371864740936843,
13
  "validation": {
14
- "soft_cross_entropy": 1.195652896563212,
15
- "accuracy": 0.3433333333333333,
16
- "brier_soft": 0.2543176457285881,
17
  "decisions": 600
18
  },
19
  "early_stopping": {
20
  "patience": 3,
21
  "min_epochs": 10,
22
  "epochs_without_improvement": 0,
23
- "best_validation_loss": 1.195652896563212,
24
  "improved": true
25
  }
26
  },
27
  {
28
  "epoch": 2,
29
- "train_soft_cross_entropy": 1.1867035720966481,
30
  "validation": {
31
- "soft_cross_entropy": 1.1503166087468466,
32
- "accuracy": 0.43666666666666665,
33
- "brier_soft": 0.2273359453678131,
34
  "decisions": 600
35
  },
36
  "early_stopping": {
37
  "patience": 3,
38
  "min_epochs": 10,
39
  "epochs_without_improvement": 0,
40
- "best_validation_loss": 1.1503166087468466,
41
  "improved": true
42
  }
43
  },
44
  {
45
  "epoch": 3,
46
- "train_soft_cross_entropy": 1.1479326088340194,
47
  "validation": {
48
- "soft_cross_entropy": 1.1359616955121359,
49
- "accuracy": 0.4583333333333333,
50
- "brier_soft": 0.21933125893274943,
51
  "decisions": 600
52
  },
53
  "early_stopping": {
54
  "patience": 3,
55
  "min_epochs": 10,
56
  "epochs_without_improvement": 0,
57
- "best_validation_loss": 1.1359616955121359,
58
  "improved": true
59
  }
60
  },
61
  {
62
  "epoch": 4,
63
- "train_soft_cross_entropy": 1.17192115077266,
64
  "validation": {
65
- "soft_cross_entropy": 1.1432366434733072,
66
- "accuracy": 0.4533333333333333,
67
- "brier_soft": 0.22181885927915573,
68
  "decisions": 600
69
  },
70
  "early_stopping": {
71
  "patience": 3,
72
  "min_epochs": 10,
73
  "epochs_without_improvement": 1,
74
- "best_validation_loss": 1.1359616955121359,
75
  "improved": false
76
  }
77
  },
78
  {
79
  "epoch": 5,
80
- "train_soft_cross_entropy": 1.1521123818114951,
81
  "validation": {
82
- "soft_cross_entropy": 1.1532048479715984,
83
- "accuracy": 0.42,
84
- "brier_soft": 0.22427557816108068,
85
  "decisions": 600
86
  },
87
  "early_stopping": {
88
  "patience": 3,
89
  "min_epochs": 10,
90
- "epochs_without_improvement": 2,
91
- "best_validation_loss": 1.1359616955121359,
92
- "improved": false
93
  }
94
  },
95
  {
96
  "epoch": 6,
97
- "train_soft_cross_entropy": 1.139600450904281,
98
  "validation": {
99
- "soft_cross_entropy": 1.1171702599525453,
100
- "accuracy": 0.465,
101
- "brier_soft": 0.21267332146565118,
102
  "decisions": 600
103
  },
104
  "early_stopping": {
105
  "patience": 3,
106
  "min_epochs": 10,
107
  "epochs_without_improvement": 0,
108
- "best_validation_loss": 1.1171702599525453,
109
  "improved": true
110
  }
111
  },
112
  {
113
  "epoch": 7,
114
- "train_soft_cross_entropy": 1.1084874327094467,
115
  "validation": {
116
- "soft_cross_entropy": 1.101010274887085,
117
- "accuracy": 0.48333333333333334,
118
- "brier_soft": 0.20341493268807728,
119
  "decisions": 600
120
  },
121
  "early_stopping": {
122
  "patience": 3,
123
  "min_epochs": 10,
124
  "epochs_without_improvement": 0,
125
- "best_validation_loss": 1.101010274887085,
126
  "improved": true
127
  }
128
  },
129
  {
130
  "epoch": 8,
131
- "train_soft_cross_entropy": 1.0972168815577472,
132
- "validation": {
133
- "soft_cross_entropy": 1.0927943936983744,
134
- "accuracy": 0.4866666666666667,
135
- "brier_soft": 0.20212342927853266,
136
- "decisions": 600
137
- },
138
- "early_stopping": {
139
- "patience": 3,
140
- "min_epochs": 10,
141
- "epochs_without_improvement": 0,
142
- "best_validation_loss": 1.0927943936983744,
143
- "improved": true
144
- }
145
- },
146
- {
147
- "epoch": 9,
148
- "train_soft_cross_entropy": 1.0959268622045164,
149
  "validation": {
150
- "soft_cross_entropy": 1.0886747964223227,
151
- "accuracy": 0.485,
152
- "brier_soft": 0.1981955200433731,
153
- "decisions": 600
154
- },
155
- "early_stopping": {
156
- "patience": 3,
157
- "min_epochs": 10,
158
- "epochs_without_improvement": 0,
159
- "best_validation_loss": 1.0886747964223227,
160
- "improved": true
161
- }
162
- },
163
- {
164
- "epoch": 10,
165
- "train_soft_cross_entropy": 1.0879168816849036,
166
- "validation": {
167
- "soft_cross_entropy": 1.0776302735010783,
168
- "accuracy": 0.5133333333333333,
169
- "brier_soft": 0.19176260660092037,
170
- "decisions": 600
171
- },
172
- "early_stopping": {
173
- "patience": 3,
174
- "min_epochs": 10,
175
- "epochs_without_improvement": 0,
176
- "best_validation_loss": 1.0776302735010783,
177
- "improved": true
178
- }
179
- },
180
- {
181
- "epoch": 11,
182
- "train_soft_cross_entropy": 1.0794475747037817,
183
- "validation": {
184
- "soft_cross_entropy": 1.0989113728205362,
185
- "accuracy": 0.49666666666666665,
186
- "brier_soft": 0.20668872609734534,
187
  "decisions": 600
188
  },
189
  "early_stopping": {
190
  "patience": 3,
191
  "min_epochs": 10,
192
  "epochs_without_improvement": 1,
193
- "best_validation_loss": 1.0776302735010783,
194
  "improved": false
195
  }
196
  },
197
  {
198
- "epoch": 12,
199
- "train_soft_cross_entropy": 1.0683085640271506,
200
- "validation": {
201
- "soft_cross_entropy": 1.0758668931325277,
202
- "accuracy": 0.5016666666666667,
203
- "brier_soft": 0.19317058285077413,
204
- "decisions": 600
205
- },
206
- "early_stopping": {
207
- "patience": 3,
208
- "min_epochs": 10,
209
- "epochs_without_improvement": 0,
210
- "best_validation_loss": 1.0758668931325277,
211
- "improved": true
212
- }
213
- },
214
- {
215
- "epoch": 13,
216
- "train_soft_cross_entropy": 1.052984880871243,
217
- "validation": {
218
- "soft_cross_entropy": 1.0600487383206685,
219
- "accuracy": 0.5316666666666666,
220
- "brier_soft": 0.18257314254840215,
221
- "decisions": 600
222
- },
223
- "early_stopping": {
224
- "patience": 3,
225
- "min_epochs": 10,
226
- "epochs_without_improvement": 0,
227
- "best_validation_loss": 1.0600487383206685,
228
- "improved": true
229
- }
230
- },
231
- {
232
- "epoch": 14,
233
- "train_soft_cross_entropy": 1.0401919462062694,
234
- "validation": {
235
- "soft_cross_entropy": 1.0525298221906025,
236
- "accuracy": 0.5366666666666666,
237
- "brier_soft": 0.1811353324353695,
238
- "decisions": 600
239
- },
240
- "early_stopping": {
241
- "patience": 3,
242
- "min_epochs": 10,
243
- "epochs_without_improvement": 0,
244
- "best_validation_loss": 1.0525298221906025,
245
- "improved": true
246
- }
247
- },
248
- {
249
- "epoch": 15,
250
- "train_soft_cross_entropy": 1.0262903751267327,
251
- "validation": {
252
- "soft_cross_entropy": 1.0581993921597799,
253
- "accuracy": 0.5233333333333333,
254
- "brier_soft": 0.18550985043247542,
255
- "decisions": 600
256
- },
257
- "early_stopping": {
258
- "patience": 3,
259
- "min_epochs": 10,
260
- "epochs_without_improvement": 1,
261
- "best_validation_loss": 1.0525298221906025,
262
- "improved": false
263
- }
264
- },
265
- {
266
- "epoch": 16,
267
- "train_soft_cross_entropy": 1.0127459985238534,
268
- "validation": {
269
- "soft_cross_entropy": 1.0545752588907877,
270
- "accuracy": 0.54,
271
- "brier_soft": 0.18113683501879374,
272
- "decisions": 600
273
- },
274
- "early_stopping": {
275
- "patience": 3,
276
- "min_epochs": 10,
277
- "epochs_without_improvement": 2,
278
- "best_validation_loss": 1.0525298221906025,
279
- "improved": false
280
- }
281
- },
282
- {
283
- "epoch": 17,
284
- "train_soft_cross_entropy": 1.0096984642523306,
285
- "validation": {
286
- "soft_cross_entropy": 1.0468252456188203,
287
- "accuracy": 0.5483333333333333,
288
- "brier_soft": 0.17509076982736588,
289
- "decisions": 600
290
- },
291
- "early_stopping": {
292
- "patience": 3,
293
- "min_epochs": 10,
294
- "epochs_without_improvement": 0,
295
- "best_validation_loss": 1.0468252456188203,
296
- "improved": true
297
- }
298
- },
299
- {
300
- "epoch": 18,
301
- "train_soft_cross_entropy": 1.0022709012914588,
302
  "validation": {
303
- "soft_cross_entropy": 1.0376164364814757,
304
- "accuracy": 0.5766666666666667,
305
- "brier_soft": 0.17226109830041728,
306
  "decisions": 600
307
  },
308
  "early_stopping": {
309
  "patience": 3,
310
  "min_epochs": 10,
311
  "epochs_without_improvement": 0,
312
- "best_validation_loss": 1.0376164364814757,
313
  "improved": true
314
  }
315
  },
316
  {
317
- "epoch": 19,
318
- "train_soft_cross_entropy": 0.9934269437083492,
319
  "validation": {
320
- "soft_cross_entropy": 1.0526774628957112,
321
- "accuracy": 0.5516666666666666,
322
- "brier_soft": 0.17908830145994822,
323
  "decisions": 600
324
  },
325
  "early_stopping": {
326
  "patience": 3,
327
  "min_epochs": 10,
328
  "epochs_without_improvement": 1,
329
- "best_validation_loss": 1.0376164364814757,
330
  "improved": false
331
  }
332
  },
333
  {
334
- "epoch": 20,
335
- "train_soft_cross_entropy": 0.9915926403911025,
336
  "validation": {
337
- "soft_cross_entropy": 1.0292588464419048,
338
- "accuracy": 0.57,
339
- "brier_soft": 0.16918850486477216,
340
  "decisions": 600
341
  },
342
  "early_stopping": {
343
  "patience": 3,
344
  "min_epochs": 10,
345
  "epochs_without_improvement": 0,
346
- "best_validation_loss": 1.0292588464419048,
347
  "improved": true
348
  }
349
  },
350
  {
351
- "epoch": 21,
352
- "train_soft_cross_entropy": 0.9833765982698511,
353
  "validation": {
354
- "soft_cross_entropy": 1.0493251442909242,
355
- "accuracy": 0.555,
356
- "brier_soft": 0.17917159788310527,
357
  "decisions": 600
358
  },
359
  "early_stopping": {
360
  "patience": 3,
361
  "min_epochs": 10,
362
  "epochs_without_improvement": 1,
363
- "best_validation_loss": 1.0292588464419048,
364
  "improved": false
365
  }
366
  },
367
  {
368
- "epoch": 22,
369
- "train_soft_cross_entropy": 0.9786927858988445,
370
  "validation": {
371
- "soft_cross_entropy": 1.0412022503217062,
372
- "accuracy": 0.555,
373
- "brier_soft": 0.17517984464764594,
374
  "decisions": 600
375
  },
376
  "early_stopping": {
377
  "patience": 3,
378
  "min_epochs": 10,
379
  "epochs_without_improvement": 2,
380
- "best_validation_loss": 1.0292588464419048,
381
  "improved": false
382
  }
383
  },
384
  {
385
- "epoch": 23,
386
- "train_soft_cross_entropy": 0.9695542351404826,
387
  "validation": {
388
- "soft_cross_entropy": 1.0376375365257262,
389
- "accuracy": 0.5583333333333333,
390
- "brier_soft": 0.17371407074232897,
391
  "decisions": 600
392
  },
393
  "early_stopping": {
394
  "patience": 3,
395
  "min_epochs": 10,
396
  "epochs_without_improvement": 3,
397
- "best_validation_loss": 1.0292588464419048,
398
  "improved": false
399
  }
400
  }
401
  ],
402
  "stopping": {
403
  "reason": "early_stopping",
404
- "epochs_completed": 23,
405
  "patience": 3,
406
  "min_epochs": 10,
407
  "epochs_without_improvement": 3,
 
9
  "epochs": [
10
  {
11
  "epoch": 1,
12
+ "train_soft_cross_entropy": 1.087307652897305,
13
  "validation": {
14
+ "soft_cross_entropy": 1.0358176565170287,
15
+ "accuracy": 0.5516666666666666,
16
+ "brier_soft": 0.17500559210777283,
17
  "decisions": 600
18
  },
19
  "early_stopping": {
20
  "patience": 3,
21
  "min_epochs": 10,
22
  "epochs_without_improvement": 0,
23
+ "best_validation_loss": 1.0358176565170287,
24
  "improved": true
25
  }
26
  },
27
  {
28
  "epoch": 2,
29
+ "train_soft_cross_entropy": 0.9998172007666694,
30
  "validation": {
31
+ "soft_cross_entropy": 1.013402551015218,
32
+ "accuracy": 0.6066666666666667,
33
+ "brier_soft": 0.16418362177908422,
34
  "decisions": 600
35
  },
36
  "early_stopping": {
37
  "patience": 3,
38
  "min_epochs": 10,
39
  "epochs_without_improvement": 0,
40
+ "best_validation_loss": 1.013402551015218,
41
  "improved": true
42
  }
43
  },
44
  {
45
  "epoch": 3,
46
+ "train_soft_cross_entropy": 0.961990624445456,
47
  "validation": {
48
+ "soft_cross_entropy": 0.9978651634852092,
49
+ "accuracy": 0.6116666666666667,
50
+ "brier_soft": 0.15542398323615392,
51
  "decisions": 600
52
  },
53
  "early_stopping": {
54
  "patience": 3,
55
  "min_epochs": 10,
56
  "epochs_without_improvement": 0,
57
+ "best_validation_loss": 0.9978651634852092,
58
  "improved": true
59
  }
60
  },
61
  {
62
  "epoch": 4,
63
+ "train_soft_cross_entropy": 0.9616141191235295,
64
  "validation": {
65
+ "soft_cross_entropy": 1.1574800237019858,
66
+ "accuracy": 0.48,
67
+ "brier_soft": 0.23033222297827402,
68
  "decisions": 600
69
  },
70
  "early_stopping": {
71
  "patience": 3,
72
  "min_epochs": 10,
73
  "epochs_without_improvement": 1,
74
+ "best_validation_loss": 0.9978651634852092,
75
  "improved": false
76
  }
77
  },
78
  {
79
  "epoch": 5,
80
+ "train_soft_cross_entropy": 0.9278180778468097,
81
  "validation": {
82
+ "soft_cross_entropy": 0.9557415008544922,
83
+ "accuracy": 0.6833333333333333,
84
+ "brier_soft": 0.12863569288204113,
85
  "decisions": 600
86
  },
87
  "early_stopping": {
88
  "patience": 3,
89
  "min_epochs": 10,
90
+ "epochs_without_improvement": 0,
91
+ "best_validation_loss": 0.9557415008544922,
92
+ "improved": true
93
  }
94
  },
95
  {
96
  "epoch": 6,
97
+ "train_soft_cross_entropy": 0.8944051215383741,
98
  "validation": {
99
+ "soft_cross_entropy": 0.9157323372364045,
100
+ "accuracy": 0.6916666666666667,
101
+ "brier_soft": 0.1038101188838482,
102
  "decisions": 600
103
  },
104
  "early_stopping": {
105
  "patience": 3,
106
  "min_epochs": 10,
107
  "epochs_without_improvement": 0,
108
+ "best_validation_loss": 0.9157323372364045,
109
  "improved": true
110
  }
111
  },
112
  {
113
  "epoch": 7,
114
+ "train_soft_cross_entropy": 0.858607970961818,
115
  "validation": {
116
+ "soft_cross_entropy": 0.9070962381362915,
117
+ "accuracy": 0.7016666666666667,
118
+ "brier_soft": 0.09971169379850228,
119
  "decisions": 600
120
  },
121
  "early_stopping": {
122
  "patience": 3,
123
  "min_epochs": 10,
124
  "epochs_without_improvement": 0,
125
+ "best_validation_loss": 0.9070962381362915,
126
  "improved": true
127
  }
128
  },
129
  {
130
  "epoch": 8,
131
+ "train_soft_cross_entropy": 0.8490360851199539,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
132
  "validation": {
133
+ "soft_cross_entropy": 0.9083957560857137,
134
+ "accuracy": 0.7116666666666667,
135
+ "brier_soft": 0.10123874122897784,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
136
  "decisions": 600
137
  },
138
  "early_stopping": {
139
  "patience": 3,
140
  "min_epochs": 10,
141
  "epochs_without_improvement": 1,
142
+ "best_validation_loss": 0.9070962381362915,
143
  "improved": false
144
  }
145
  },
146
  {
147
+ "epoch": 9,
148
+ "train_soft_cross_entropy": 0.8271938294393045,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
149
  "validation": {
150
+ "soft_cross_entropy": 0.8859062685569128,
151
+ "accuracy": 0.7233333333333334,
152
+ "brier_soft": 0.08866231886049111,
153
  "decisions": 600
154
  },
155
  "early_stopping": {
156
  "patience": 3,
157
  "min_epochs": 10,
158
  "epochs_without_improvement": 0,
159
+ "best_validation_loss": 0.8859062685569128,
160
  "improved": true
161
  }
162
  },
163
  {
164
+ "epoch": 10,
165
+ "train_soft_cross_entropy": 0.8117916690861737,
166
  "validation": {
167
+ "soft_cross_entropy": 0.8980106647809346,
168
+ "accuracy": 0.72,
169
+ "brier_soft": 0.09496299955993891,
170
  "decisions": 600
171
  },
172
  "early_stopping": {
173
  "patience": 3,
174
  "min_epochs": 10,
175
  "epochs_without_improvement": 1,
176
+ "best_validation_loss": 0.8859062685569128,
177
  "improved": false
178
  }
179
  },
180
  {
181
+ "epoch": 11,
182
+ "train_soft_cross_entropy": 0.8015254776124601,
183
  "validation": {
184
+ "soft_cross_entropy": 0.8771124110619227,
185
+ "accuracy": 0.7116666666666667,
186
+ "brier_soft": 0.0828241604194045,
187
  "decisions": 600
188
  },
189
  "early_stopping": {
190
  "patience": 3,
191
  "min_epochs": 10,
192
  "epochs_without_improvement": 0,
193
+ "best_validation_loss": 0.8771124110619227,
194
  "improved": true
195
  }
196
  },
197
  {
198
+ "epoch": 12,
199
+ "train_soft_cross_entropy": 0.7930714415620874,
200
  "validation": {
201
+ "soft_cross_entropy": 0.8903550871213277,
202
+ "accuracy": 0.7233333333333334,
203
+ "brier_soft": 0.09005908486122886,
204
  "decisions": 600
205
  },
206
  "early_stopping": {
207
  "patience": 3,
208
  "min_epochs": 10,
209
  "epochs_without_improvement": 1,
210
+ "best_validation_loss": 0.8771124110619227,
211
  "improved": false
212
  }
213
  },
214
  {
215
+ "epoch": 13,
216
+ "train_soft_cross_entropy": 0.7864397717405248,
217
  "validation": {
218
+ "soft_cross_entropy": 0.8776129017273585,
219
+ "accuracy": 0.7283333333333334,
220
+ "brier_soft": 0.08296463950847587,
221
  "decisions": 600
222
  },
223
  "early_stopping": {
224
  "patience": 3,
225
  "min_epochs": 10,
226
  "epochs_without_improvement": 2,
227
+ "best_validation_loss": 0.8771124110619227,
228
  "improved": false
229
  }
230
  },
231
  {
232
+ "epoch": 14,
233
+ "train_soft_cross_entropy": 0.7832702258339634,
234
  "validation": {
235
+ "soft_cross_entropy": 0.8776572781801224,
236
+ "accuracy": 0.7466666666666667,
237
+ "brier_soft": 0.08209236452666421,
238
  "decisions": 600
239
  },
240
  "early_stopping": {
241
  "patience": 3,
242
  "min_epochs": 10,
243
  "epochs_without_improvement": 3,
244
+ "best_validation_loss": 0.8771124110619227,
245
  "improved": false
246
  }
247
  }
248
  ],
249
  "stopping": {
250
  "reason": "early_stopping",
251
+ "epochs_completed": 14,
252
  "patience": 3,
253
  "min_epochs": 10,
254
  "epochs_without_improvement": 3,
evidence/seed-1-history.json CHANGED
@@ -9,297 +9,212 @@
9
  "epochs": [
10
  {
11
  "epoch": 1,
12
- "train_soft_cross_entropy": 1.0002122403074194,
13
  "validation": {
14
- "soft_cross_entropy": 0.9555757478872935,
15
- "accuracy": 0.665,
16
- "brier_soft": 0.13395084381103517,
17
  "decisions": 600
18
  },
19
  "early_stopping": {
20
  "patience": 3,
21
  "min_epochs": 10,
22
  "epochs_without_improvement": 0,
23
- "best_validation_loss": 0.9555757478872935,
24
  "improved": true
25
  }
26
  },
27
  {
28
  "epoch": 2,
29
- "train_soft_cross_entropy": 0.895550852616628,
30
  "validation": {
31
- "soft_cross_entropy": 0.8875669745604197,
32
- "accuracy": 0.7216666666666667,
33
- "brier_soft": 0.08987226173281669,
34
  "decisions": 600
35
  },
36
  "early_stopping": {
37
  "patience": 3,
38
  "min_epochs": 10,
39
  "epochs_without_improvement": 0,
40
- "best_validation_loss": 0.8875669745604197,
41
  "improved": true
42
  }
43
  },
44
  {
45
  "epoch": 3,
46
- "train_soft_cross_entropy": 0.8399523215382187,
47
  "validation": {
48
- "soft_cross_entropy": 0.8776892550786336,
49
- "accuracy": 0.725,
50
- "brier_soft": 0.0829127719004949,
51
  "decisions": 600
52
  },
53
  "early_stopping": {
54
  "patience": 3,
55
  "min_epochs": 10,
56
  "epochs_without_improvement": 0,
57
- "best_validation_loss": 0.8776892550786336,
58
  "improved": true
59
  }
60
  },
61
  {
62
  "epoch": 4,
63
- "train_soft_cross_entropy": 0.8150597869908368,
64
  "validation": {
65
- "soft_cross_entropy": 0.8523939752578735,
66
- "accuracy": 0.7666666666666667,
67
- "brier_soft": 0.07098765720923741,
68
  "decisions": 600
69
  },
70
  "early_stopping": {
71
  "patience": 3,
72
  "min_epochs": 10,
73
  "epochs_without_improvement": 0,
74
- "best_validation_loss": 0.8523939752578735,
75
  "improved": true
76
  }
77
  },
78
  {
79
  "epoch": 5,
80
- "train_soft_cross_entropy": 0.7998251422246297,
81
  "validation": {
82
- "soft_cross_entropy": 0.853373521566391,
83
- "accuracy": 0.745,
84
- "brier_soft": 0.07141184127579132,
85
  "decisions": 600
86
  },
87
  "early_stopping": {
88
  "patience": 3,
89
  "min_epochs": 10,
90
- "epochs_without_improvement": 1,
91
- "best_validation_loss": 0.8523939752578735,
92
- "improved": false
93
  }
94
  },
95
  {
96
  "epoch": 6,
97
- "train_soft_cross_entropy": 0.7924317064991704,
98
  "validation": {
99
- "soft_cross_entropy": 0.8612222145001094,
100
- "accuracy": 0.7516666666666667,
101
- "brier_soft": 0.07590873730679353,
102
  "decisions": 600
103
  },
104
  "early_stopping": {
105
  "patience": 3,
106
  "min_epochs": 10,
107
- "epochs_without_improvement": 2,
108
- "best_validation_loss": 0.8523939752578735,
109
  "improved": false
110
  }
111
  },
112
  {
113
  "epoch": 7,
114
- "train_soft_cross_entropy": 0.7880261039733887,
115
  "validation": {
116
- "soft_cross_entropy": 0.8462035346031189,
117
  "accuracy": 0.7466666666666667,
118
- "brier_soft": 0.06817000946650903,
119
  "decisions": 600
120
  },
121
  "early_stopping": {
122
  "patience": 3,
123
  "min_epochs": 10,
124
- "epochs_without_improvement": 0,
125
- "best_validation_loss": 0.8462035346031189,
126
- "improved": true
127
  }
128
  },
129
  {
130
  "epoch": 8,
131
- "train_soft_cross_entropy": 0.7816576555923179,
132
  "validation": {
133
- "soft_cross_entropy": 0.8440792632102966,
134
- "accuracy": 0.7633333333333333,
135
- "brier_soft": 0.06713677939027547,
136
  "decisions": 600
137
  },
138
  "early_stopping": {
139
  "patience": 3,
140
  "min_epochs": 10,
141
- "epochs_without_improvement": 0,
142
- "best_validation_loss": 0.8440792632102966,
143
- "improved": true
144
  }
145
  },
146
  {
147
  "epoch": 9,
148
- "train_soft_cross_entropy": 0.7783506816404837,
149
  "validation": {
150
- "soft_cross_entropy": 0.8438087296485901,
151
- "accuracy": 0.7683333333333333,
152
- "brier_soft": 0.06672837336858113,
153
  "decisions": 600
154
  },
155
  "early_stopping": {
156
  "patience": 3,
157
  "min_epochs": 10,
158
  "epochs_without_improvement": 0,
159
- "best_validation_loss": 0.8438087296485901,
160
  "improved": true
161
  }
162
  },
163
  {
164
  "epoch": 10,
165
- "train_soft_cross_entropy": 0.7754799175262451,
166
  "validation": {
167
- "soft_cross_entropy": 0.8468067542711893,
168
- "accuracy": 0.755,
169
- "brier_soft": 0.06876023932515334,
170
  "decisions": 600
171
  },
172
  "early_stopping": {
173
  "patience": 3,
174
  "min_epochs": 10,
175
  "epochs_without_improvement": 1,
176
- "best_validation_loss": 0.8438087296485901,
177
  "improved": false
178
  }
179
  },
180
  {
181
  "epoch": 11,
182
- "train_soft_cross_entropy": 0.7734219932114637,
183
  "validation": {
184
- "soft_cross_entropy": 0.8521546524763107,
185
- "accuracy": 0.7783333333333333,
186
- "brier_soft": 0.07055049358556668,
187
  "decisions": 600
188
  },
189
  "early_stopping": {
190
  "patience": 3,
191
  "min_epochs": 10,
192
  "epochs_without_improvement": 2,
193
- "best_validation_loss": 0.8438087296485901,
194
  "improved": false
195
  }
196
  },
197
  {
198
  "epoch": 12,
199
- "train_soft_cross_entropy": 0.7719193545094243,
200
- "validation": {
201
- "soft_cross_entropy": 0.842797059615453,
202
- "accuracy": 0.7766666666666666,
203
- "brier_soft": 0.06609537469533583,
204
- "decisions": 600
205
- },
206
- "early_stopping": {
207
- "patience": 3,
208
- "min_epochs": 10,
209
- "epochs_without_improvement": 0,
210
- "best_validation_loss": 0.842797059615453,
211
- "improved": true
212
- }
213
- },
214
- {
215
- "epoch": 13,
216
- "train_soft_cross_entropy": 0.7722371352601934,
217
- "validation": {
218
- "soft_cross_entropy": 0.848033101161321,
219
- "accuracy": 0.7633333333333333,
220
- "brier_soft": 0.06775808438037832,
221
- "decisions": 600
222
- },
223
- "early_stopping": {
224
- "patience": 3,
225
- "min_epochs": 10,
226
- "epochs_without_improvement": 1,
227
- "best_validation_loss": 0.842797059615453,
228
- "improved": false
229
- }
230
- },
231
- {
232
- "epoch": 14,
233
- "train_soft_cross_entropy": 0.7689926949695305,
234
- "validation": {
235
- "soft_cross_entropy": 0.8377540612220764,
236
- "accuracy": 0.765,
237
- "brier_soft": 0.06276483290052662,
238
- "decisions": 600
239
- },
240
- "early_stopping": {
241
- "patience": 3,
242
- "min_epochs": 10,
243
- "epochs_without_improvement": 0,
244
- "best_validation_loss": 0.8377540612220764,
245
- "improved": true
246
- }
247
- },
248
- {
249
- "epoch": 15,
250
- "train_soft_cross_entropy": 0.7658979249000549,
251
- "validation": {
252
- "soft_cross_entropy": 0.8519912085930507,
253
- "accuracy": 0.765,
254
- "brier_soft": 0.07007166295622785,
255
- "decisions": 600
256
- },
257
- "early_stopping": {
258
- "patience": 3,
259
- "min_epochs": 10,
260
- "epochs_without_improvement": 1,
261
- "best_validation_loss": 0.8377540612220764,
262
- "improved": false
263
- }
264
- },
265
- {
266
- "epoch": 16,
267
- "train_soft_cross_entropy": 0.7651108237107594,
268
- "validation": {
269
- "soft_cross_entropy": 0.8408213845888773,
270
- "accuracy": 0.7633333333333333,
271
- "brier_soft": 0.06544813718336324,
272
- "decisions": 600
273
- },
274
- "early_stopping": {
275
- "patience": 3,
276
- "min_epochs": 10,
277
- "epochs_without_improvement": 2,
278
- "best_validation_loss": 0.8377540612220764,
279
- "improved": false
280
- }
281
- },
282
- {
283
- "epoch": 17,
284
- "train_soft_cross_entropy": 0.7817914198063038,
285
  "validation": {
286
- "soft_cross_entropy": 0.8468167595068614,
287
- "accuracy": 0.7683333333333333,
288
- "brier_soft": 0.06789324807934463,
289
  "decisions": 600
290
  },
291
  "early_stopping": {
292
  "patience": 3,
293
  "min_epochs": 10,
294
  "epochs_without_improvement": 3,
295
- "best_validation_loss": 0.8377540612220764,
296
  "improved": false
297
  }
298
  }
299
  ],
300
  "stopping": {
301
  "reason": "early_stopping",
302
- "epochs_completed": 17,
303
  "patience": 3,
304
  "min_epochs": 10,
305
  "epochs_without_improvement": 3,
 
9
  "epochs": [
10
  {
11
  "epoch": 1,
12
+ "train_soft_cross_entropy": 1.0518495603843971,
13
  "validation": {
14
+ "soft_cross_entropy": 0.9761401589711507,
15
+ "accuracy": 0.6666666666666666,
16
+ "brier_soft": 0.14186986642579238,
17
  "decisions": 600
18
  },
19
  "early_stopping": {
20
  "patience": 3,
21
  "min_epochs": 10,
22
  "epochs_without_improvement": 0,
23
+ "best_validation_loss": 0.9761401589711507,
24
  "improved": true
25
  }
26
  },
27
  {
28
  "epoch": 2,
29
+ "train_soft_cross_entropy": 0.912936895776678,
30
  "validation": {
31
+ "soft_cross_entropy": 0.9058315515518188,
32
+ "accuracy": 0.7133333333333334,
33
+ "brier_soft": 0.09862347106138865,
34
  "decisions": 600
35
  },
36
  "early_stopping": {
37
  "patience": 3,
38
  "min_epochs": 10,
39
  "epochs_without_improvement": 0,
40
+ "best_validation_loss": 0.9058315515518188,
41
  "improved": true
42
  }
43
  },
44
  {
45
  "epoch": 3,
46
+ "train_soft_cross_entropy": 0.8595328189708569,
47
  "validation": {
48
+ "soft_cross_entropy": 0.8869859429200491,
49
+ "accuracy": 0.7333333333333333,
50
+ "brier_soft": 0.08666441923628251,
51
  "decisions": 600
52
  },
53
  "early_stopping": {
54
  "patience": 3,
55
  "min_epochs": 10,
56
  "epochs_without_improvement": 0,
57
+ "best_validation_loss": 0.8869859429200491,
58
  "improved": true
59
  }
60
  },
61
  {
62
  "epoch": 4,
63
+ "train_soft_cross_entropy": 0.8311581687574033,
64
  "validation": {
65
+ "soft_cross_entropy": 0.875844070315361,
66
+ "accuracy": 0.7466666666666667,
67
+ "brier_soft": 0.08233745716512203,
68
  "decisions": 600
69
  },
70
  "early_stopping": {
71
  "patience": 3,
72
  "min_epochs": 10,
73
  "epochs_without_improvement": 0,
74
+ "best_validation_loss": 0.875844070315361,
75
  "improved": true
76
  }
77
  },
78
  {
79
  "epoch": 5,
80
+ "train_soft_cross_entropy": 0.812676587899526,
81
  "validation": {
82
+ "soft_cross_entropy": 0.8657626557350159,
83
+ "accuracy": 0.74,
84
+ "brier_soft": 0.07910436561020712,
85
  "decisions": 600
86
  },
87
  "early_stopping": {
88
  "patience": 3,
89
  "min_epochs": 10,
90
+ "epochs_without_improvement": 0,
91
+ "best_validation_loss": 0.8657626557350159,
92
+ "improved": true
93
  }
94
  },
95
  {
96
  "epoch": 6,
97
+ "train_soft_cross_entropy": 0.7997228077605919,
98
  "validation": {
99
+ "soft_cross_entropy": 0.8763724052906037,
100
+ "accuracy": 0.7433333333333333,
101
+ "brier_soft": 0.08383749471356472,
102
  "decisions": 600
103
  },
104
  "early_stopping": {
105
  "patience": 3,
106
  "min_epochs": 10,
107
+ "epochs_without_improvement": 1,
108
+ "best_validation_loss": 0.8657626557350159,
109
  "improved": false
110
  }
111
  },
112
  {
113
  "epoch": 7,
114
+ "train_soft_cross_entropy": 0.7919920165892,
115
  "validation": {
116
+ "soft_cross_entropy": 0.8720610602696737,
117
  "accuracy": 0.7466666666666667,
118
+ "brier_soft": 0.08182110945383708,
119
  "decisions": 600
120
  },
121
  "early_stopping": {
122
  "patience": 3,
123
  "min_epochs": 10,
124
+ "epochs_without_improvement": 2,
125
+ "best_validation_loss": 0.8657626557350159,
126
+ "improved": false
127
  }
128
  },
129
  {
130
  "epoch": 8,
131
+ "train_soft_cross_entropy": 0.7865607251061334,
132
  "validation": {
133
+ "soft_cross_entropy": 0.8708612742026647,
134
+ "accuracy": 0.735,
135
+ "brier_soft": 0.08140341289962331,
136
  "decisions": 600
137
  },
138
  "early_stopping": {
139
  "patience": 3,
140
  "min_epochs": 10,
141
+ "epochs_without_improvement": 3,
142
+ "best_validation_loss": 0.8657626557350159,
143
+ "improved": false
144
  }
145
  },
146
  {
147
  "epoch": 9,
148
+ "train_soft_cross_entropy": 0.781722169496395,
149
  "validation": {
150
+ "soft_cross_entropy": 0.8583925066391627,
151
+ "accuracy": 0.7466666666666667,
152
+ "brier_soft": 0.07437195796364297,
153
  "decisions": 600
154
  },
155
  "early_stopping": {
156
  "patience": 3,
157
  "min_epochs": 10,
158
  "epochs_without_improvement": 0,
159
+ "best_validation_loss": 0.8583925066391627,
160
  "improved": true
161
  }
162
  },
163
  {
164
  "epoch": 10,
165
+ "train_soft_cross_entropy": 0.7776996397530591,
166
  "validation": {
167
+ "soft_cross_entropy": 0.8674025861422221,
168
+ "accuracy": 0.7433333333333333,
169
+ "brier_soft": 0.07951540602991979,
170
  "decisions": 600
171
  },
172
  "early_stopping": {
173
  "patience": 3,
174
  "min_epochs": 10,
175
  "epochs_without_improvement": 1,
176
+ "best_validation_loss": 0.8583925066391627,
177
  "improved": false
178
  }
179
  },
180
  {
181
  "epoch": 11,
182
+ "train_soft_cross_entropy": 0.7751605097894315,
183
  "validation": {
184
+ "soft_cross_entropy": 0.86817025522391,
185
+ "accuracy": 0.7433333333333333,
186
+ "brier_soft": 0.07928484301393231,
187
  "decisions": 600
188
  },
189
  "early_stopping": {
190
  "patience": 3,
191
  "min_epochs": 10,
192
  "epochs_without_improvement": 2,
193
+ "best_validation_loss": 0.8583925066391627,
194
  "improved": false
195
  }
196
  },
197
  {
198
  "epoch": 12,
199
+ "train_soft_cross_entropy": 0.7721157057638521,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
200
  "validation": {
201
+ "soft_cross_entropy": 0.8638703493277232,
202
+ "accuracy": 0.7583333333333333,
203
+ "brier_soft": 0.0763223198801279,
204
  "decisions": 600
205
  },
206
  "early_stopping": {
207
  "patience": 3,
208
  "min_epochs": 10,
209
  "epochs_without_improvement": 3,
210
+ "best_validation_loss": 0.8583925066391627,
211
  "improved": false
212
  }
213
  }
214
  ],
215
  "stopping": {
216
  "reason": "early_stopping",
217
+ "epochs_completed": 12,
218
  "patience": 3,
219
  "min_epochs": 10,
220
  "epochs_without_improvement": 3,
evidence/seed-2-history.json CHANGED
@@ -9,314 +9,280 @@
9
  "epochs": [
10
  {
11
  "epoch": 1,
12
- "train_soft_cross_entropy": 1.1849676999339351,
13
  "validation": {
14
- "soft_cross_entropy": 1.1535689290364584,
15
- "accuracy": 0.4266666666666667,
16
- "brier_soft": 0.2362905572851499,
17
  "decisions": 600
18
  },
19
  "early_stopping": {
20
  "patience": 3,
21
  "min_epochs": 10,
22
  "epochs_without_improvement": 0,
23
- "best_validation_loss": 1.1535689290364584,
24
  "improved": true
25
  }
26
  },
27
  {
28
  "epoch": 2,
29
- "train_soft_cross_entropy": 1.1660638781830117,
30
  "validation": {
31
- "soft_cross_entropy": 1.1240008862813313,
32
- "accuracy": 0.44666666666666666,
33
- "brier_soft": 0.2163119477033615,
34
  "decisions": 600
35
  },
36
  "early_stopping": {
37
  "patience": 3,
38
  "min_epochs": 10,
39
  "epochs_without_improvement": 0,
40
- "best_validation_loss": 1.1240008862813313,
41
  "improved": true
42
  }
43
  },
44
  {
45
  "epoch": 3,
46
- "train_soft_cross_entropy": 1.1152002820262203,
47
  "validation": {
48
- "soft_cross_entropy": 1.111793461640676,
49
- "accuracy": 0.4816666666666667,
50
- "brier_soft": 0.21397013902664186,
51
  "decisions": 600
52
  },
53
  "early_stopping": {
54
  "patience": 3,
55
  "min_epochs": 10,
56
  "epochs_without_improvement": 0,
57
- "best_validation_loss": 1.111793461640676,
58
  "improved": true
59
  }
60
  },
61
  {
62
  "epoch": 4,
63
- "train_soft_cross_entropy": 1.0905909954177009,
64
  "validation": {
65
- "soft_cross_entropy": 1.0822320612271628,
66
- "accuracy": 0.49833333333333335,
67
- "brier_soft": 0.19861485362052916,
68
  "decisions": 600
69
  },
70
  "early_stopping": {
71
  "patience": 3,
72
  "min_epochs": 10,
73
  "epochs_without_improvement": 0,
74
- "best_validation_loss": 1.0822320612271628,
75
  "improved": true
76
  }
77
  },
78
  {
79
  "epoch": 5,
80
- "train_soft_cross_entropy": 1.0488388820047732,
81
  "validation": {
82
- "soft_cross_entropy": 1.0466791025797526,
83
- "accuracy": 0.5383333333333333,
84
- "brier_soft": 0.18209601615866025,
85
  "decisions": 600
86
  },
87
  "early_stopping": {
88
  "patience": 3,
89
  "min_epochs": 10,
90
- "epochs_without_improvement": 0,
91
- "best_validation_loss": 1.0466791025797526,
92
- "improved": true
93
  }
94
  },
95
  {
96
  "epoch": 6,
97
- "train_soft_cross_entropy": 1.0179926924352292,
98
  "validation": {
99
- "soft_cross_entropy": 1.0302869494756062,
100
- "accuracy": 0.5633333333333334,
101
- "brier_soft": 0.1717192947367827,
102
  "decisions": 600
103
  },
104
  "early_stopping": {
105
  "patience": 3,
106
  "min_epochs": 10,
107
  "epochs_without_improvement": 0,
108
- "best_validation_loss": 1.0302869494756062,
109
  "improved": true
110
  }
111
  },
112
  {
113
  "epoch": 7,
114
- "train_soft_cross_entropy": 1.0011521806540313,
115
  "validation": {
116
- "soft_cross_entropy": 1.0306965279579163,
117
- "accuracy": 0.5533333333333333,
118
- "brier_soft": 0.17337830337385338,
119
  "decisions": 600
120
  },
121
  "early_stopping": {
122
  "patience": 3,
123
  "min_epochs": 10,
124
- "epochs_without_improvement": 1,
125
- "best_validation_loss": 1.0302869494756062,
126
- "improved": false
127
  }
128
  },
129
  {
130
  "epoch": 8,
131
- "train_soft_cross_entropy": 0.9898686430189345,
132
  "validation": {
133
- "soft_cross_entropy": 1.0352703229586284,
134
- "accuracy": 0.5533333333333333,
135
- "brier_soft": 0.17590213686227799,
136
  "decisions": 600
137
  },
138
  "early_stopping": {
139
  "patience": 3,
140
  "min_epochs": 10,
141
- "epochs_without_improvement": 2,
142
- "best_validation_loss": 1.0302869494756062,
143
- "improved": false
144
  }
145
  },
146
  {
147
  "epoch": 9,
148
- "train_soft_cross_entropy": 0.9732429999775357,
149
  "validation": {
150
- "soft_cross_entropy": 1.0375931040445963,
151
- "accuracy": 0.56,
152
- "brier_soft": 0.1798627228786548,
153
  "decisions": 600
154
  },
155
  "early_stopping": {
156
  "patience": 3,
157
  "min_epochs": 10,
158
- "epochs_without_improvement": 3,
159
- "best_validation_loss": 1.0302869494756062,
160
  "improved": false
161
  }
162
  },
163
  {
164
  "epoch": 10,
165
- "train_soft_cross_entropy": 0.9616305458987201,
166
  "validation": {
167
- "soft_cross_entropy": 1.0139233843485513,
168
- "accuracy": 0.6116666666666667,
169
- "brier_soft": 0.16307140870640674,
170
  "decisions": 600
171
  },
172
  "early_stopping": {
173
  "patience": 3,
174
  "min_epochs": 10,
175
  "epochs_without_improvement": 0,
176
- "best_validation_loss": 1.0139233843485513,
177
  "improved": true
178
  }
179
  },
180
  {
181
  "epoch": 11,
182
- "train_soft_cross_entropy": 0.9440543747831274,
183
- "validation": {
184
- "soft_cross_entropy": 1.0138032054901123,
185
- "accuracy": 0.6166666666666667,
186
- "brier_soft": 0.165894419302543,
187
- "decisions": 600
188
- },
189
- "early_stopping": {
190
- "patience": 3,
191
- "min_epochs": 10,
192
- "epochs_without_improvement": 0,
193
- "best_validation_loss": 1.0138032054901123,
194
- "improved": true
195
- }
196
- },
197
- {
198
- "epoch": 12,
199
- "train_soft_cross_entropy": 0.9309376364284091,
200
  "validation": {
201
- "soft_cross_entropy": 0.9929217569033305,
202
- "accuracy": 0.6266666666666667,
203
- "brier_soft": 0.15395031906664372,
204
- "decisions": 600
205
- },
206
- "early_stopping": {
207
- "patience": 3,
208
- "min_epochs": 10,
209
- "epochs_without_improvement": 0,
210
- "best_validation_loss": 0.9929217569033305,
211
- "improved": true
212
- }
213
- },
214
- {
215
- "epoch": 13,
216
- "train_soft_cross_entropy": 0.9144250082969666,
217
- "validation": {
218
- "soft_cross_entropy": 0.9964303588867187,
219
- "accuracy": 0.6266666666666667,
220
- "brier_soft": 0.1560174826408426,
221
  "decisions": 600
222
  },
223
  "early_stopping": {
224
  "patience": 3,
225
  "min_epochs": 10,
226
  "epochs_without_improvement": 1,
227
- "best_validation_loss": 0.9929217569033305,
228
  "improved": false
229
  }
230
  },
231
  {
232
- "epoch": 14,
233
- "train_soft_cross_entropy": 0.9063218560925237,
234
  "validation": {
235
- "soft_cross_entropy": 1.005851571559906,
236
- "accuracy": 0.6133333333333333,
237
- "brier_soft": 0.1608497215807438,
238
  "decisions": 600
239
  },
240
  "early_stopping": {
241
  "patience": 3,
242
  "min_epochs": 10,
243
  "epochs_without_improvement": 2,
244
- "best_validation_loss": 0.9929217569033305,
245
  "improved": false
246
  }
247
  },
248
  {
249
- "epoch": 15,
250
- "train_soft_cross_entropy": 0.9014872776137458,
251
  "validation": {
252
- "soft_cross_entropy": 0.9867499772707621,
253
- "accuracy": 0.6316666666666667,
254
- "brier_soft": 0.14927835414807003,
255
  "decisions": 600
256
  },
257
  "early_stopping": {
258
  "patience": 3,
259
  "min_epochs": 10,
260
  "epochs_without_improvement": 0,
261
- "best_validation_loss": 0.9867499772707621,
262
  "improved": true
263
  }
264
  },
265
  {
266
- "epoch": 16,
267
- "train_soft_cross_entropy": 0.8920391595805133,
268
  "validation": {
269
- "soft_cross_entropy": 0.9926939264933268,
270
- "accuracy": 0.6283333333333333,
271
- "brier_soft": 0.15533705055713654,
272
  "decisions": 600
273
  },
274
  "early_stopping": {
275
  "patience": 3,
276
  "min_epochs": 10,
277
  "epochs_without_improvement": 1,
278
- "best_validation_loss": 0.9867499772707621,
279
  "improved": false
280
  }
281
  },
282
  {
283
- "epoch": 17,
284
- "train_soft_cross_entropy": 0.8795346831833875,
285
  "validation": {
286
- "soft_cross_entropy": 1.02633438428243,
287
- "accuracy": 0.6583333333333333,
288
- "brier_soft": 0.16615503872434298,
289
  "decisions": 600
290
  },
291
  "early_stopping": {
292
  "patience": 3,
293
  "min_epochs": 10,
294
  "epochs_without_improvement": 2,
295
- "best_validation_loss": 0.9867499772707621,
296
  "improved": false
297
  }
298
  },
299
  {
300
- "epoch": 18,
301
- "train_soft_cross_entropy": 0.879185232453876,
302
  "validation": {
303
- "soft_cross_entropy": 1.0007199597358705,
304
- "accuracy": 0.65,
305
- "brier_soft": 0.15312831775595745,
306
  "decisions": 600
307
  },
308
  "early_stopping": {
309
  "patience": 3,
310
  "min_epochs": 10,
311
  "epochs_without_improvement": 3,
312
- "best_validation_loss": 0.9867499772707621,
313
  "improved": false
314
  }
315
  }
316
  ],
317
  "stopping": {
318
  "reason": "early_stopping",
319
- "epochs_completed": 18,
320
  "patience": 3,
321
  "min_epochs": 10,
322
  "epochs_without_improvement": 3,
 
9
  "epochs": [
10
  {
11
  "epoch": 1,
12
+ "train_soft_cross_entropy": 1.0231783160456904,
13
  "validation": {
14
+ "soft_cross_entropy": 0.9674482258160909,
15
+ "accuracy": 0.6516666666666666,
16
+ "brier_soft": 0.139310811907053,
17
  "decisions": 600
18
  },
19
  "early_stopping": {
20
  "patience": 3,
21
  "min_epochs": 10,
22
  "epochs_without_improvement": 0,
23
+ "best_validation_loss": 0.9674482258160909,
24
  "improved": true
25
  }
26
  },
27
  {
28
  "epoch": 2,
29
+ "train_soft_cross_entropy": 0.8996694993972778,
30
  "validation": {
31
+ "soft_cross_entropy": 0.8974683928489685,
32
+ "accuracy": 0.75,
33
+ "brier_soft": 0.09312087532132864,
34
  "decisions": 600
35
  },
36
  "early_stopping": {
37
  "patience": 3,
38
  "min_epochs": 10,
39
  "epochs_without_improvement": 0,
40
+ "best_validation_loss": 0.8974683928489685,
41
  "improved": true
42
  }
43
  },
44
  {
45
  "epoch": 3,
46
+ "train_soft_cross_entropy": 0.8483260071277618,
47
  "validation": {
48
+ "soft_cross_entropy": 0.8809152638912201,
49
+ "accuracy": 0.7433333333333333,
50
+ "brier_soft": 0.08576706000914176,
51
  "decisions": 600
52
  },
53
  "early_stopping": {
54
  "patience": 3,
55
  "min_epochs": 10,
56
  "epochs_without_improvement": 0,
57
+ "best_validation_loss": 0.8809152638912201,
58
  "improved": true
59
  }
60
  },
61
  {
62
  "epoch": 4,
63
+ "train_soft_cross_entropy": 0.8239163917523843,
64
  "validation": {
65
+ "soft_cross_entropy": 0.8699432893594106,
66
+ "accuracy": 0.75,
67
+ "brier_soft": 0.07744032381723324,
68
  "decisions": 600
69
  },
70
  "early_stopping": {
71
  "patience": 3,
72
  "min_epochs": 10,
73
  "epochs_without_improvement": 0,
74
+ "best_validation_loss": 0.8699432893594106,
75
  "improved": true
76
  }
77
  },
78
  {
79
  "epoch": 5,
80
+ "train_soft_cross_entropy": 0.8087966615623898,
81
  "validation": {
82
+ "soft_cross_entropy": 0.8749503823121388,
83
+ "accuracy": 0.725,
84
+ "brier_soft": 0.08202173060427109,
85
  "decisions": 600
86
  },
87
  "early_stopping": {
88
  "patience": 3,
89
  "min_epochs": 10,
90
+ "epochs_without_improvement": 1,
91
+ "best_validation_loss": 0.8699432893594106,
92
+ "improved": false
93
  }
94
  },
95
  {
96
  "epoch": 6,
97
+ "train_soft_cross_entropy": 0.798634327782525,
98
  "validation": {
99
+ "soft_cross_entropy": 0.8646231130758921,
100
+ "accuracy": 0.7633333333333333,
101
+ "brier_soft": 0.07594866358985504,
102
  "decisions": 600
103
  },
104
  "early_stopping": {
105
  "patience": 3,
106
  "min_epochs": 10,
107
  "epochs_without_improvement": 0,
108
+ "best_validation_loss": 0.8646231130758921,
109
  "improved": true
110
  }
111
  },
112
  {
113
  "epoch": 7,
114
+ "train_soft_cross_entropy": 0.7900045536182545,
115
  "validation": {
116
+ "soft_cross_entropy": 0.8598153670628865,
117
+ "accuracy": 0.755,
118
+ "brier_soft": 0.07338624291742842,
119
  "decisions": 600
120
  },
121
  "early_stopping": {
122
  "patience": 3,
123
  "min_epochs": 10,
124
+ "epochs_without_improvement": 0,
125
+ "best_validation_loss": 0.8598153670628865,
126
+ "improved": true
127
  }
128
  },
129
  {
130
  "epoch": 8,
131
+ "train_soft_cross_entropy": 0.7848307268266325,
132
  "validation": {
133
+ "soft_cross_entropy": 0.8584532356262207,
134
+ "accuracy": 0.7483333333333333,
135
+ "brier_soft": 0.07417237816999356,
136
  "decisions": 600
137
  },
138
  "early_stopping": {
139
  "patience": 3,
140
  "min_epochs": 10,
141
+ "epochs_without_improvement": 0,
142
+ "best_validation_loss": 0.8584532356262207,
143
+ "improved": true
144
  }
145
  },
146
  {
147
  "epoch": 9,
148
+ "train_soft_cross_entropy": 0.7802449210043306,
149
  "validation": {
150
+ "soft_cross_entropy": 0.863755419254303,
151
+ "accuracy": 0.7433333333333333,
152
+ "brier_soft": 0.07594012685120105,
153
  "decisions": 600
154
  },
155
  "early_stopping": {
156
  "patience": 3,
157
  "min_epochs": 10,
158
+ "epochs_without_improvement": 1,
159
+ "best_validation_loss": 0.8584532356262207,
160
  "improved": false
161
  }
162
  },
163
  {
164
  "epoch": 10,
165
+ "train_soft_cross_entropy": 0.777841743098365,
166
  "validation": {
167
+ "soft_cross_entropy": 0.8546390326817831,
168
+ "accuracy": 0.7566666666666667,
169
+ "brier_soft": 0.07055099235226711,
170
  "decisions": 600
171
  },
172
  "early_stopping": {
173
  "patience": 3,
174
  "min_epochs": 10,
175
  "epochs_without_improvement": 0,
176
+ "best_validation_loss": 0.8546390326817831,
177
  "improved": true
178
  }
179
  },
180
  {
181
  "epoch": 11,
182
+ "train_soft_cross_entropy": 0.7738726913045954,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
183
  "validation": {
184
+ "soft_cross_entropy": 0.8676082201798757,
185
+ "accuracy": 0.7466666666666667,
186
+ "brier_soft": 0.07854201994836331,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
187
  "decisions": 600
188
  },
189
  "early_stopping": {
190
  "patience": 3,
191
  "min_epochs": 10,
192
  "epochs_without_improvement": 1,
193
+ "best_validation_loss": 0.8546390326817831,
194
  "improved": false
195
  }
196
  },
197
  {
198
+ "epoch": 12,
199
+ "train_soft_cross_entropy": 0.7718432058228387,
200
  "validation": {
201
+ "soft_cross_entropy": 0.8577940479914348,
202
+ "accuracy": 0.7366666666666667,
203
+ "brier_soft": 0.07410156331335505,
204
  "decisions": 600
205
  },
206
  "early_stopping": {
207
  "patience": 3,
208
  "min_epochs": 10,
209
  "epochs_without_improvement": 2,
210
+ "best_validation_loss": 0.8546390326817831,
211
  "improved": false
212
  }
213
  },
214
  {
215
+ "epoch": 13,
216
+ "train_soft_cross_entropy": 0.7713618384467231,
217
  "validation": {
218
+ "soft_cross_entropy": 0.8543428432941437,
219
+ "accuracy": 0.7583333333333333,
220
+ "brier_soft": 0.07187043125430743,
221
  "decisions": 600
222
  },
223
  "early_stopping": {
224
  "patience": 3,
225
  "min_epochs": 10,
226
  "epochs_without_improvement": 0,
227
+ "best_validation_loss": 0.8543428432941437,
228
  "improved": true
229
  }
230
  },
231
  {
232
+ "epoch": 14,
233
+ "train_soft_cross_entropy": 0.7684376273331819,
234
  "validation": {
235
+ "soft_cross_entropy": 0.8578398124376932,
236
+ "accuracy": 0.7483333333333333,
237
+ "brier_soft": 0.07348543658852577,
238
  "decisions": 600
239
  },
240
  "early_stopping": {
241
  "patience": 3,
242
  "min_epochs": 10,
243
  "epochs_without_improvement": 1,
244
+ "best_validation_loss": 0.8543428432941437,
245
  "improved": false
246
  }
247
  },
248
  {
249
+ "epoch": 15,
250
+ "train_soft_cross_entropy": 0.766730530526903,
251
  "validation": {
252
+ "soft_cross_entropy": 0.860800955692927,
253
+ "accuracy": 0.7383333333333333,
254
+ "brier_soft": 0.07522662562007705,
255
  "decisions": 600
256
  },
257
  "early_stopping": {
258
  "patience": 3,
259
  "min_epochs": 10,
260
  "epochs_without_improvement": 2,
261
+ "best_validation_loss": 0.8543428432941437,
262
  "improved": false
263
  }
264
  },
265
  {
266
+ "epoch": 16,
267
+ "train_soft_cross_entropy": 0.7659876039734593,
268
  "validation": {
269
+ "soft_cross_entropy": 0.8564839088916778,
270
+ "accuracy": 0.7533333333333333,
271
+ "brier_soft": 0.07288925250992179,
272
  "decisions": 600
273
  },
274
  "early_stopping": {
275
  "patience": 3,
276
  "min_epochs": 10,
277
  "epochs_without_improvement": 3,
278
+ "best_validation_loss": 0.8543428432941437,
279
  "improved": false
280
  }
281
  }
282
  ],
283
  "stopping": {
284
  "reason": "early_stopping",
285
+ "epochs_completed": 16,
286
  "patience": 3,
287
  "min_epochs": 10,
288
  "epochs_without_improvement": 3,
evidence/seed-3-history.json CHANGED
@@ -9,467 +9,195 @@
9
  "epochs": [
10
  {
11
  "epoch": 1,
12
- "train_soft_cross_entropy": 1.1980749452555621,
13
  "validation": {
14
- "soft_cross_entropy": 1.16634921391805,
15
- "accuracy": 0.4266666666666667,
16
- "brier_soft": 0.2427163557211558,
17
  "decisions": 600
18
  },
19
  "early_stopping": {
20
  "patience": 3,
21
  "min_epochs": 10,
22
  "epochs_without_improvement": 0,
23
- "best_validation_loss": 1.16634921391805,
24
  "improved": true
25
  }
26
  },
27
  {
28
  "epoch": 2,
29
- "train_soft_cross_entropy": 1.1646321240177862,
30
  "validation": {
31
- "soft_cross_entropy": 1.1217502697308859,
32
- "accuracy": 0.4533333333333333,
33
- "brier_soft": 0.2186164912581444,
34
  "decisions": 600
35
  },
36
  "early_stopping": {
37
  "patience": 3,
38
  "min_epochs": 10,
39
  "epochs_without_improvement": 0,
40
- "best_validation_loss": 1.1217502697308859,
41
  "improved": true
42
  }
43
  },
44
  {
45
  "epoch": 3,
46
- "train_soft_cross_entropy": 1.1056555556367944,
47
  "validation": {
48
- "soft_cross_entropy": 1.093972578048706,
49
- "accuracy": 0.505,
50
- "brier_soft": 0.199787005285422,
51
  "decisions": 600
52
  },
53
  "early_stopping": {
54
  "patience": 3,
55
  "min_epochs": 10,
56
  "epochs_without_improvement": 0,
57
- "best_validation_loss": 1.093972578048706,
58
  "improved": true
59
  }
60
  },
61
  {
62
  "epoch": 4,
63
- "train_soft_cross_entropy": 1.0671771099832323,
64
  "validation": {
65
- "soft_cross_entropy": 1.0549713468551636,
66
- "accuracy": 0.5516666666666666,
67
- "brier_soft": 0.18068428387244542,
68
  "decisions": 600
69
  },
70
  "early_stopping": {
71
  "patience": 3,
72
  "min_epochs": 10,
73
  "epochs_without_improvement": 0,
74
- "best_validation_loss": 1.0549713468551636,
75
  "improved": true
76
  }
77
  },
78
  {
79
  "epoch": 5,
80
- "train_soft_cross_entropy": 1.0494003568755257,
81
  "validation": {
82
- "soft_cross_entropy": 1.0567963059743246,
83
- "accuracy": 0.54,
84
- "brier_soft": 0.18257476242880027,
85
  "decisions": 600
86
  },
87
  "early_stopping": {
88
  "patience": 3,
89
  "min_epochs": 10,
90
  "epochs_without_improvement": 1,
91
- "best_validation_loss": 1.0549713468551636,
92
  "improved": false
93
  }
94
  },
95
  {
96
  "epoch": 6,
97
- "train_soft_cross_entropy": 1.038703513851872,
98
- "validation": {
99
- "soft_cross_entropy": 1.0705021890004476,
100
- "accuracy": 0.525,
101
- "brier_soft": 0.19120527582863966,
102
- "decisions": 600
103
- },
104
- "early_stopping": {
105
- "patience": 3,
106
- "min_epochs": 10,
107
- "epochs_without_improvement": 2,
108
- "best_validation_loss": 1.0549713468551636,
109
- "improved": false
110
- }
111
- },
112
- {
113
- "epoch": 7,
114
- "train_soft_cross_entropy": 1.0313762316880404,
115
- "validation": {
116
- "soft_cross_entropy": 1.0447989773750306,
117
- "accuracy": 0.5466666666666666,
118
- "brier_soft": 0.17501757830381393,
119
- "decisions": 600
120
- },
121
- "early_stopping": {
122
- "patience": 3,
123
- "min_epochs": 10,
124
- "epochs_without_improvement": 0,
125
- "best_validation_loss": 1.0447989773750306,
126
- "improved": true
127
- }
128
- },
129
- {
130
- "epoch": 8,
131
- "train_soft_cross_entropy": 1.017818791601393,
132
- "validation": {
133
- "soft_cross_entropy": 1.024161856174469,
134
- "accuracy": 0.5816666666666667,
135
- "brier_soft": 0.1630909529576699,
136
- "decisions": 600
137
- },
138
- "early_stopping": {
139
- "patience": 3,
140
- "min_epochs": 10,
141
- "epochs_without_improvement": 0,
142
- "best_validation_loss": 1.024161856174469,
143
- "improved": true
144
- }
145
- },
146
- {
147
- "epoch": 9,
148
- "train_soft_cross_entropy": 0.9899537578335514,
149
- "validation": {
150
- "soft_cross_entropy": 1.0265704361597696,
151
- "accuracy": 0.6083333333333333,
152
- "brier_soft": 0.16808087141563496,
153
- "decisions": 600
154
- },
155
- "early_stopping": {
156
- "patience": 3,
157
- "min_epochs": 10,
158
- "epochs_without_improvement": 1,
159
- "best_validation_loss": 1.024161856174469,
160
- "improved": false
161
- }
162
- },
163
- {
164
- "epoch": 10,
165
- "train_soft_cross_entropy": 0.9632394627288535,
166
- "validation": {
167
- "soft_cross_entropy": 1.0183856010437011,
168
- "accuracy": 0.5983333333333334,
169
- "brier_soft": 0.1608257148663203,
170
- "decisions": 600
171
- },
172
- "early_stopping": {
173
- "patience": 3,
174
- "min_epochs": 10,
175
- "epochs_without_improvement": 0,
176
- "best_validation_loss": 1.0183856010437011,
177
- "improved": true
178
- }
179
- },
180
- {
181
- "epoch": 11,
182
- "train_soft_cross_entropy": 0.9342622081438701,
183
- "validation": {
184
- "soft_cross_entropy": 1.0024159566561381,
185
- "accuracy": 0.6333333333333333,
186
- "brier_soft": 0.14996731283764045,
187
- "decisions": 600
188
- },
189
- "early_stopping": {
190
- "patience": 3,
191
- "min_epochs": 10,
192
- "epochs_without_improvement": 0,
193
- "best_validation_loss": 1.0024159566561381,
194
- "improved": true
195
- }
196
- },
197
- {
198
- "epoch": 12,
199
- "train_soft_cross_entropy": 0.9137524432606168,
200
- "validation": {
201
- "soft_cross_entropy": 0.9873138956228892,
202
- "accuracy": 0.6516666666666666,
203
- "brier_soft": 0.14277802929282188,
204
- "decisions": 600
205
- },
206
- "early_stopping": {
207
- "patience": 3,
208
- "min_epochs": 10,
209
- "epochs_without_improvement": 0,
210
- "best_validation_loss": 0.9873138956228892,
211
- "improved": true
212
- }
213
- },
214
- {
215
- "epoch": 13,
216
- "train_soft_cross_entropy": 0.8881368139496556,
217
- "validation": {
218
- "soft_cross_entropy": 0.9626709421475729,
219
- "accuracy": 0.6716666666666666,
220
- "brier_soft": 0.12283751085400581,
221
- "decisions": 600
222
- },
223
- "early_stopping": {
224
- "patience": 3,
225
- "min_epochs": 10,
226
- "epochs_without_improvement": 0,
227
- "best_validation_loss": 0.9626709421475729,
228
- "improved": true
229
- }
230
- },
231
- {
232
- "epoch": 14,
233
- "train_soft_cross_entropy": 0.865500467883216,
234
  "validation": {
235
- "soft_cross_entropy": 0.9580995849768321,
236
- "accuracy": 0.685,
237
- "brier_soft": 0.12383397127191226,
238
  "decisions": 600
239
  },
240
  "early_stopping": {
241
  "patience": 3,
242
  "min_epochs": 10,
243
  "epochs_without_improvement": 0,
244
- "best_validation_loss": 0.9580995849768321,
245
  "improved": true
246
  }
247
  },
248
  {
249
- "epoch": 15,
250
- "train_soft_cross_entropy": 0.8549449095461104,
251
- "validation": {
252
- "soft_cross_entropy": 0.9462178750832876,
253
- "accuracy": 0.6916666666666667,
254
- "brier_soft": 0.1187205430244406,
255
- "decisions": 600
256
- },
257
- "early_stopping": {
258
- "patience": 3,
259
- "min_epochs": 10,
260
- "epochs_without_improvement": 0,
261
- "best_validation_loss": 0.9462178750832876,
262
- "improved": true
263
- }
264
- },
265
- {
266
- "epoch": 16,
267
- "train_soft_cross_entropy": 0.8394836834183446,
268
- "validation": {
269
- "soft_cross_entropy": 0.9463119049866994,
270
- "accuracy": 0.6933333333333334,
271
- "brier_soft": 0.11875337022046248,
272
- "decisions": 600
273
- },
274
- "early_stopping": {
275
- "patience": 3,
276
- "min_epochs": 10,
277
- "epochs_without_improvement": 1,
278
- "best_validation_loss": 0.9462178750832876,
279
- "improved": false
280
- }
281
- },
282
- {
283
- "epoch": 17,
284
- "train_soft_cross_entropy": 0.8286281617040987,
285
- "validation": {
286
- "soft_cross_entropy": 0.941375896135966,
287
- "accuracy": 0.685,
288
- "brier_soft": 0.11455494280904531,
289
- "decisions": 600
290
- },
291
- "early_stopping": {
292
- "patience": 3,
293
- "min_epochs": 10,
294
- "epochs_without_improvement": 0,
295
- "best_validation_loss": 0.941375896135966,
296
- "improved": true
297
- }
298
- },
299
- {
300
- "epoch": 18,
301
- "train_soft_cross_entropy": 0.8222652976601212,
302
- "validation": {
303
- "soft_cross_entropy": 0.9331588689486185,
304
- "accuracy": 0.6866666666666666,
305
- "brier_soft": 0.11237604923546314,
306
- "decisions": 600
307
- },
308
- "early_stopping": {
309
- "patience": 3,
310
- "min_epochs": 10,
311
- "epochs_without_improvement": 0,
312
- "best_validation_loss": 0.9331588689486185,
313
- "improved": true
314
- }
315
- },
316
- {
317
- "epoch": 19,
318
- "train_soft_cross_entropy": 0.8160834203826056,
319
- "validation": {
320
- "soft_cross_entropy": 0.9387085942427317,
321
- "accuracy": 0.6916666666666667,
322
- "brier_soft": 0.11558652246991793,
323
- "decisions": 600
324
- },
325
- "early_stopping": {
326
- "patience": 3,
327
- "min_epochs": 10,
328
- "epochs_without_improvement": 1,
329
- "best_validation_loss": 0.9331588689486185,
330
- "improved": false
331
- }
332
- },
333
- {
334
- "epoch": 20,
335
- "train_soft_cross_entropy": 0.8087538785846146,
336
- "validation": {
337
- "soft_cross_entropy": 0.9408066284656524,
338
- "accuracy": 0.6783333333333333,
339
- "brier_soft": 0.11930533437679211,
340
- "decisions": 600
341
- },
342
- "early_stopping": {
343
- "patience": 3,
344
- "min_epochs": 10,
345
- "epochs_without_improvement": 2,
346
- "best_validation_loss": 0.9331588689486185,
347
- "improved": false
348
- }
349
- },
350
- {
351
- "epoch": 21,
352
- "train_soft_cross_entropy": 0.8048860243956248,
353
- "validation": {
354
- "soft_cross_entropy": 0.9236933688322703,
355
- "accuracy": 0.7016666666666667,
356
- "brier_soft": 0.10792215374608835,
357
- "decisions": 600
358
- },
359
- "early_stopping": {
360
- "patience": 3,
361
- "min_epochs": 10,
362
- "epochs_without_improvement": 0,
363
- "best_validation_loss": 0.9236933688322703,
364
- "improved": true
365
- }
366
- },
367
- {
368
- "epoch": 22,
369
- "train_soft_cross_entropy": 0.8001213801790167,
370
  "validation": {
371
- "soft_cross_entropy": 0.9347231006622314,
372
- "accuracy": 0.6983333333333334,
373
- "brier_soft": 0.11313647958139579,
374
  "decisions": 600
375
  },
376
  "early_stopping": {
377
  "patience": 3,
378
  "min_epochs": 10,
379
  "epochs_without_improvement": 1,
380
- "best_validation_loss": 0.9236933688322703,
381
  "improved": false
382
  }
383
  },
384
  {
385
- "epoch": 23,
386
- "train_soft_cross_entropy": 0.7940637088263476,
387
- "validation": {
388
- "soft_cross_entropy": 0.9299098292986552,
389
- "accuracy": 0.7033333333333334,
390
- "brier_soft": 0.11475709093113741,
391
- "decisions": 600
392
- },
393
- "early_stopping": {
394
- "patience": 3,
395
- "min_epochs": 10,
396
- "epochs_without_improvement": 2,
397
- "best_validation_loss": 0.9236933688322703,
398
- "improved": false
399
- }
400
- },
401
- {
402
- "epoch": 24,
403
- "train_soft_cross_entropy": 0.7884197377717054,
404
  "validation": {
405
- "soft_cross_entropy": 0.9105557099978129,
406
- "accuracy": 0.715,
407
- "brier_soft": 0.1013037375236551,
408
  "decisions": 600
409
  },
410
  "early_stopping": {
411
  "patience": 3,
412
  "min_epochs": 10,
413
  "epochs_without_improvement": 0,
414
- "best_validation_loss": 0.9105557099978129,
415
  "improved": true
416
  }
417
  },
418
  {
419
- "epoch": 25,
420
- "train_soft_cross_entropy": 0.784954049454795,
421
  "validation": {
422
- "soft_cross_entropy": 0.923290346066157,
423
- "accuracy": 0.71,
424
- "brier_soft": 0.10530036373684803,
425
  "decisions": 600
426
  },
427
  "early_stopping": {
428
  "patience": 3,
429
  "min_epochs": 10,
430
  "epochs_without_improvement": 1,
431
- "best_validation_loss": 0.9105557099978129,
432
  "improved": false
433
  }
434
  },
435
  {
436
- "epoch": 26,
437
- "train_soft_cross_entropy": 0.7828840752442677,
438
  "validation": {
439
- "soft_cross_entropy": 0.9299345835049947,
440
- "accuracy": 0.7066666666666667,
441
- "brier_soft": 0.11173659774164359,
442
  "decisions": 600
443
  },
444
  "early_stopping": {
445
  "patience": 3,
446
  "min_epochs": 10,
447
  "epochs_without_improvement": 2,
448
- "best_validation_loss": 0.9105557099978129,
449
  "improved": false
450
  }
451
  },
452
  {
453
- "epoch": 27,
454
- "train_soft_cross_entropy": 0.7790651953661883,
455
  "validation": {
456
- "soft_cross_entropy": 0.9137676254908244,
457
- "accuracy": 0.7216666666666667,
458
- "brier_soft": 0.10172271362195412,
459
  "decisions": 600
460
  },
461
  "early_stopping": {
462
  "patience": 3,
463
  "min_epochs": 10,
464
  "epochs_without_improvement": 3,
465
- "best_validation_loss": 0.9105557099978129,
466
  "improved": false
467
  }
468
  }
469
  ],
470
  "stopping": {
471
  "reason": "early_stopping",
472
- "epochs_completed": 27,
473
  "patience": 3,
474
  "min_epochs": 10,
475
  "epochs_without_improvement": 3,
 
9
  "epochs": [
10
  {
11
  "epoch": 1,
12
+ "train_soft_cross_entropy": 1.0557677305186237,
13
  "validation": {
14
+ "soft_cross_entropy": 0.990565903186798,
15
+ "accuracy": 0.6166666666666667,
16
+ "brier_soft": 0.15243564940989018,
17
  "decisions": 600
18
  },
19
  "early_stopping": {
20
  "patience": 3,
21
  "min_epochs": 10,
22
  "epochs_without_improvement": 0,
23
+ "best_validation_loss": 0.990565903186798,
24
  "improved": true
25
  }
26
  },
27
  {
28
  "epoch": 2,
29
+ "train_soft_cross_entropy": 0.9286279639491328,
30
  "validation": {
31
+ "soft_cross_entropy": 0.9022069962819418,
32
+ "accuracy": 0.71,
33
+ "brier_soft": 0.0947934572895368,
34
  "decisions": 600
35
  },
36
  "early_stopping": {
37
  "patience": 3,
38
  "min_epochs": 10,
39
  "epochs_without_improvement": 0,
40
+ "best_validation_loss": 0.9022069962819418,
41
  "improved": true
42
  }
43
  },
44
  {
45
  "epoch": 3,
46
+ "train_soft_cross_entropy": 0.8668451248716424,
47
  "validation": {
48
+ "soft_cross_entropy": 0.8912256610393524,
49
+ "accuracy": 0.725,
50
+ "brier_soft": 0.08988383966187637,
51
  "decisions": 600
52
  },
53
  "early_stopping": {
54
  "patience": 3,
55
  "min_epochs": 10,
56
  "epochs_without_improvement": 0,
57
+ "best_validation_loss": 0.8912256610393524,
58
  "improved": true
59
  }
60
  },
61
  {
62
  "epoch": 4,
63
+ "train_soft_cross_entropy": 0.8375598757796817,
64
  "validation": {
65
+ "soft_cross_entropy": 0.8716744474569956,
66
+ "accuracy": 0.7566666666666667,
67
+ "brier_soft": 0.07936330476154883,
68
  "decisions": 600
69
  },
70
  "early_stopping": {
71
  "patience": 3,
72
  "min_epochs": 10,
73
  "epochs_without_improvement": 0,
74
+ "best_validation_loss": 0.8716744474569956,
75
  "improved": true
76
  }
77
  },
78
  {
79
  "epoch": 5,
80
+ "train_soft_cross_entropy": 0.8176034839064986,
81
  "validation": {
82
+ "soft_cross_entropy": 0.8743925015131633,
83
+ "accuracy": 0.7333333333333333,
84
+ "brier_soft": 0.07994898026498655,
85
  "decisions": 600
86
  },
87
  "early_stopping": {
88
  "patience": 3,
89
  "min_epochs": 10,
90
  "epochs_without_improvement": 1,
91
+ "best_validation_loss": 0.8716744474569956,
92
  "improved": false
93
  }
94
  },
95
  {
96
  "epoch": 6,
97
+ "train_soft_cross_entropy": 0.8031262465318044,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
98
  "validation": {
99
+ "soft_cross_entropy": 0.8599367336432139,
100
+ "accuracy": 0.76,
101
+ "brier_soft": 0.07330287167181572,
102
  "decisions": 600
103
  },
104
  "early_stopping": {
105
  "patience": 3,
106
  "min_epochs": 10,
107
  "epochs_without_improvement": 0,
108
+ "best_validation_loss": 0.8599367336432139,
109
  "improved": true
110
  }
111
  },
112
  {
113
+ "epoch": 7,
114
+ "train_soft_cross_entropy": 0.7926116739820551,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
115
  "validation": {
116
+ "soft_cross_entropy": 0.8638877793153127,
117
+ "accuracy": 0.7566666666666667,
118
+ "brier_soft": 0.07329062350094319,
119
  "decisions": 600
120
  },
121
  "early_stopping": {
122
  "patience": 3,
123
  "min_epochs": 10,
124
  "epochs_without_improvement": 1,
125
+ "best_validation_loss": 0.8599367336432139,
126
  "improved": false
127
  }
128
  },
129
  {
130
+ "epoch": 8,
131
+ "train_soft_cross_entropy": 0.7867316756866596,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
132
  "validation": {
133
+ "soft_cross_entropy": 0.8545633804798126,
134
+ "accuracy": 0.77,
135
+ "brier_soft": 0.06962224062221746,
136
  "decisions": 600
137
  },
138
  "early_stopping": {
139
  "patience": 3,
140
  "min_epochs": 10,
141
  "epochs_without_improvement": 0,
142
+ "best_validation_loss": 0.8545633804798126,
143
  "improved": true
144
  }
145
  },
146
  {
147
+ "epoch": 9,
148
+ "train_soft_cross_entropy": 0.7820918090696688,
149
  "validation": {
150
+ "soft_cross_entropy": 0.8555870747566223,
151
+ "accuracy": 0.7683333333333333,
152
+ "brier_soft": 0.06890943528773884,
153
  "decisions": 600
154
  },
155
  "early_stopping": {
156
  "patience": 3,
157
  "min_epochs": 10,
158
  "epochs_without_improvement": 1,
159
+ "best_validation_loss": 0.8545633804798126,
160
  "improved": false
161
  }
162
  },
163
  {
164
+ "epoch": 10,
165
+ "train_soft_cross_entropy": 0.779588269745862,
166
  "validation": {
167
+ "soft_cross_entropy": 0.8600131020943323,
168
+ "accuracy": 0.76,
169
+ "brier_soft": 0.07248457937811811,
170
  "decisions": 600
171
  },
172
  "early_stopping": {
173
  "patience": 3,
174
  "min_epochs": 10,
175
  "epochs_without_improvement": 2,
176
+ "best_validation_loss": 0.8545633804798126,
177
  "improved": false
178
  }
179
  },
180
  {
181
+ "epoch": 11,
182
+ "train_soft_cross_entropy": 0.7761333607302772,
183
  "validation": {
184
+ "soft_cross_entropy": 0.8604028668006262,
185
+ "accuracy": 0.7616666666666667,
186
+ "brier_soft": 0.07338622493979831,
187
  "decisions": 600
188
  },
189
  "early_stopping": {
190
  "patience": 3,
191
  "min_epochs": 10,
192
  "epochs_without_improvement": 3,
193
+ "best_validation_loss": 0.8545633804798126,
194
  "improved": false
195
  }
196
  }
197
  ],
198
  "stopping": {
199
  "reason": "early_stopping",
200
+ "epochs_completed": 11,
201
  "patience": 3,
202
  "min_epochs": 10,
203
  "epochs_without_improvement": 3,
evidence/seed-4-history.json CHANGED
@@ -9,178 +9,195 @@
9
  "epochs": [
10
  {
11
  "epoch": 1,
12
- "train_soft_cross_entropy": 1.002756692568461,
13
  "validation": {
14
- "soft_cross_entropy": 0.9602992224693299,
15
- "accuracy": 0.6783333333333333,
16
- "brier_soft": 0.1296700432151556,
17
  "decisions": 600
18
  },
19
  "early_stopping": {
20
  "patience": 3,
21
  "min_epochs": 10,
22
  "epochs_without_improvement": 0,
23
- "best_validation_loss": 0.9602992224693299,
24
  "improved": true
25
  }
26
  },
27
  {
28
  "epoch": 2,
29
- "train_soft_cross_entropy": 0.8854288237624698,
30
  "validation": {
31
- "soft_cross_entropy": 0.8732180511951446,
32
- "accuracy": 0.7316666666666667,
33
- "brier_soft": 0.08001670623819034,
34
  "decisions": 600
35
  },
36
  "early_stopping": {
37
  "patience": 3,
38
  "min_epochs": 10,
39
  "epochs_without_improvement": 0,
40
- "best_validation_loss": 0.8732180511951446,
41
  "improved": true
42
  }
43
  },
44
  {
45
  "epoch": 3,
46
- "train_soft_cross_entropy": 0.8375967553809837,
47
  "validation": {
48
- "soft_cross_entropy": 0.8618760740756989,
49
- "accuracy": 0.7566666666666667,
50
- "brier_soft": 0.07448753335823616,
51
  "decisions": 600
52
  },
53
  "early_stopping": {
54
  "patience": 3,
55
  "min_epochs": 10,
56
  "epochs_without_improvement": 0,
57
- "best_validation_loss": 0.8618760740756989,
58
  "improved": true
59
  }
60
  },
61
  {
62
  "epoch": 4,
63
- "train_soft_cross_entropy": 0.8120107175244226,
64
  "validation": {
65
- "soft_cross_entropy": 0.8502214312553406,
66
- "accuracy": 0.7683333333333333,
67
- "brier_soft": 0.06676056392490864,
68
  "decisions": 600
69
  },
70
  "early_stopping": {
71
  "patience": 3,
72
  "min_epochs": 10,
73
  "epochs_without_improvement": 0,
74
- "best_validation_loss": 0.8502214312553406,
75
  "improved": true
76
  }
77
  },
78
  {
79
  "epoch": 5,
80
- "train_soft_cross_entropy": 0.7984476970301734,
81
  "validation": {
82
- "soft_cross_entropy": 0.8527251331011454,
83
- "accuracy": 0.765,
84
- "brier_soft": 0.07089023986210426,
85
  "decisions": 600
86
  },
87
  "early_stopping": {
88
  "patience": 3,
89
  "min_epochs": 10,
90
- "epochs_without_improvement": 1,
91
- "best_validation_loss": 0.8502214312553406,
92
- "improved": false
93
  }
94
  },
95
  {
96
  "epoch": 6,
97
- "train_soft_cross_entropy": 0.7902782445042221,
98
  "validation": {
99
- "soft_cross_entropy": 0.8496846203009287,
100
- "accuracy": 0.765,
101
- "brier_soft": 0.06887457605761786,
102
  "decisions": 600
103
  },
104
  "early_stopping": {
105
  "patience": 3,
106
  "min_epochs": 10,
107
- "epochs_without_improvement": 0,
108
- "best_validation_loss": 0.8496846203009287,
109
- "improved": true
110
  }
111
  },
112
  {
113
  "epoch": 7,
114
- "train_soft_cross_entropy": 0.7847388537283297,
115
  "validation": {
116
- "soft_cross_entropy": 0.8448993345101674,
117
- "accuracy": 0.7533333333333333,
118
- "brier_soft": 0.0668264798882107,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
119
  "decisions": 600
120
  },
121
  "early_stopping": {
122
  "patience": 3,
123
  "min_epochs": 10,
124
  "epochs_without_improvement": 0,
125
- "best_validation_loss": 0.8448993345101674,
126
  "improved": true
127
  }
128
  },
129
  {
130
- "epoch": 8,
131
- "train_soft_cross_entropy": 0.780297756018462,
132
  "validation": {
133
- "soft_cross_entropy": 0.8471182523171107,
134
- "accuracy": 0.7683333333333333,
135
- "brier_soft": 0.06819427679603299,
136
  "decisions": 600
137
  },
138
  "early_stopping": {
139
  "patience": 3,
140
  "min_epochs": 10,
141
  "epochs_without_improvement": 1,
142
- "best_validation_loss": 0.8448993345101674,
143
  "improved": false
144
  }
145
  },
146
  {
147
- "epoch": 9,
148
- "train_soft_cross_entropy": 0.7801684511590887,
149
  "validation": {
150
- "soft_cross_entropy": 0.8466147363185883,
151
- "accuracy": 0.775,
152
- "brier_soft": 0.06677873468647401,
153
  "decisions": 600
154
  },
155
  "early_stopping": {
156
  "patience": 3,
157
  "min_epochs": 10,
158
  "epochs_without_improvement": 2,
159
- "best_validation_loss": 0.8448993345101674,
160
  "improved": false
161
  }
162
  },
163
  {
164
- "epoch": 10,
165
- "train_soft_cross_entropy": 0.7775638208565888,
166
  "validation": {
167
- "soft_cross_entropy": 0.8455785755316416,
168
- "accuracy": 0.78,
169
- "brier_soft": 0.06686171248555184,
170
  "decisions": 600
171
  },
172
  "early_stopping": {
173
  "patience": 3,
174
  "min_epochs": 10,
175
  "epochs_without_improvement": 3,
176
- "best_validation_loss": 0.8448993345101674,
177
  "improved": false
178
  }
179
  }
180
  ],
181
  "stopping": {
182
  "reason": "early_stopping",
183
- "epochs_completed": 10,
184
  "patience": 3,
185
  "min_epochs": 10,
186
  "epochs_without_improvement": 3,
 
9
  "epochs": [
10
  {
11
  "epoch": 1,
12
+ "train_soft_cross_entropy": 1.022522153324551,
13
  "validation": {
14
+ "soft_cross_entropy": 0.9845864299933116,
15
+ "accuracy": 0.6683333333333333,
16
+ "brier_soft": 0.14738356913129488,
17
  "decisions": 600
18
  },
19
  "early_stopping": {
20
  "patience": 3,
21
  "min_epochs": 10,
22
  "epochs_without_improvement": 0,
23
+ "best_validation_loss": 0.9845864299933116,
24
  "improved": true
25
  }
26
  },
27
  {
28
  "epoch": 2,
29
+ "train_soft_cross_entropy": 0.9041571515136295,
30
  "validation": {
31
+ "soft_cross_entropy": 0.9089026947816213,
32
+ "accuracy": 0.7216666666666667,
33
+ "brier_soft": 0.10058322168886662,
34
  "decisions": 600
35
  },
36
  "early_stopping": {
37
  "patience": 3,
38
  "min_epochs": 10,
39
  "epochs_without_improvement": 0,
40
+ "best_validation_loss": 0.9089026947816213,
41
  "improved": true
42
  }
43
  },
44
  {
45
  "epoch": 3,
46
+ "train_soft_cross_entropy": 0.8495574627099214,
47
  "validation": {
48
+ "soft_cross_entropy": 0.8853352854649226,
49
+ "accuracy": 0.74,
50
+ "brier_soft": 0.08744065261135499,
51
  "decisions": 600
52
  },
53
  "early_stopping": {
54
  "patience": 3,
55
  "min_epochs": 10,
56
  "epochs_without_improvement": 0,
57
+ "best_validation_loss": 0.8853352854649226,
58
  "improved": true
59
  }
60
  },
61
  {
62
  "epoch": 4,
63
+ "train_soft_cross_entropy": 0.8263491551522856,
64
  "validation": {
65
+ "soft_cross_entropy": 0.8789768069982529,
66
+ "accuracy": 0.7416666666666667,
67
+ "brier_soft": 0.08262864720076323,
68
  "decisions": 600
69
  },
70
  "early_stopping": {
71
  "patience": 3,
72
  "min_epochs": 10,
73
  "epochs_without_improvement": 0,
74
+ "best_validation_loss": 0.8789768069982529,
75
  "improved": true
76
  }
77
  },
78
  {
79
  "epoch": 5,
80
+ "train_soft_cross_entropy": 0.8096302935812209,
81
  "validation": {
82
+ "soft_cross_entropy": 0.864648152589798,
83
+ "accuracy": 0.73,
84
+ "brier_soft": 0.07628292332092922,
85
  "decisions": 600
86
  },
87
  "early_stopping": {
88
  "patience": 3,
89
  "min_epochs": 10,
90
+ "epochs_without_improvement": 0,
91
+ "best_validation_loss": 0.864648152589798,
92
+ "improved": true
93
  }
94
  },
95
  {
96
  "epoch": 6,
97
+ "train_soft_cross_entropy": 0.7991735404950601,
98
  "validation": {
99
+ "soft_cross_entropy": 0.8743107922871908,
100
+ "accuracy": 0.74,
101
+ "brier_soft": 0.08211790287246307,
102
  "decisions": 600
103
  },
104
  "early_stopping": {
105
  "patience": 3,
106
  "min_epochs": 10,
107
+ "epochs_without_improvement": 1,
108
+ "best_validation_loss": 0.864648152589798,
109
+ "improved": false
110
  }
111
  },
112
  {
113
  "epoch": 7,
114
+ "train_soft_cross_entropy": 0.7918641125714337,
115
  "validation": {
116
+ "soft_cross_entropy": 0.86731658577919,
117
+ "accuracy": 0.7483333333333333,
118
+ "brier_soft": 0.07722779513647159,
119
+ "decisions": 600
120
+ },
121
+ "early_stopping": {
122
+ "patience": 3,
123
+ "min_epochs": 10,
124
+ "epochs_without_improvement": 2,
125
+ "best_validation_loss": 0.864648152589798,
126
+ "improved": false
127
+ }
128
+ },
129
+ {
130
+ "epoch": 8,
131
+ "train_soft_cross_entropy": 0.7861463651833711,
132
+ "validation": {
133
+ "soft_cross_entropy": 0.8540202794472377,
134
+ "accuracy": 0.76,
135
+ "brier_soft": 0.07019127607345581,
136
  "decisions": 600
137
  },
138
  "early_stopping": {
139
  "patience": 3,
140
  "min_epochs": 10,
141
  "epochs_without_improvement": 0,
142
+ "best_validation_loss": 0.8540202794472377,
143
  "improved": true
144
  }
145
  },
146
  {
147
+ "epoch": 9,
148
+ "train_soft_cross_entropy": 0.7806310887248428,
149
  "validation": {
150
+ "soft_cross_entropy": 0.8646407808860143,
151
+ "accuracy": 0.735,
152
+ "brier_soft": 0.07658187456429005,
153
  "decisions": 600
154
  },
155
  "early_stopping": {
156
  "patience": 3,
157
  "min_epochs": 10,
158
  "epochs_without_improvement": 1,
159
+ "best_validation_loss": 0.8540202794472377,
160
  "improved": false
161
  }
162
  },
163
  {
164
+ "epoch": 10,
165
+ "train_soft_cross_entropy": 0.7768384672094274,
166
  "validation": {
167
+ "soft_cross_entropy": 0.8618478431304296,
168
+ "accuracy": 0.73,
169
+ "brier_soft": 0.07576349244763454,
170
  "decisions": 600
171
  },
172
  "early_stopping": {
173
  "patience": 3,
174
  "min_epochs": 10,
175
  "epochs_without_improvement": 2,
176
+ "best_validation_loss": 0.8540202794472377,
177
  "improved": false
178
  }
179
  },
180
  {
181
+ "epoch": 11,
182
+ "train_soft_cross_entropy": 0.7740854684511821,
183
  "validation": {
184
+ "soft_cross_entropy": 0.8627219025293986,
185
+ "accuracy": 0.7283333333333334,
186
+ "brier_soft": 0.07741292332919936,
187
  "decisions": 600
188
  },
189
  "early_stopping": {
190
  "patience": 3,
191
  "min_epochs": 10,
192
  "epochs_without_improvement": 3,
193
+ "best_validation_loss": 0.8540202794472377,
194
  "improved": false
195
  }
196
  }
197
  ],
198
  "stopping": {
199
  "reason": "early_stopping",
200
+ "epochs_completed": 11,
201
  "patience": 3,
202
  "min_epochs": 10,
203
  "epochs_without_improvement": 3,
export-verification.json CHANGED
@@ -1,11 +1,11 @@
1
  {
2
  "passed": true,
3
- "candidate_seed": 1,
4
- "selected_epoch": 14,
5
  "cases": 3516,
6
  "decisions": 5116,
7
  "benchmark_cases_sha256": "10fb671245cdac1ee09e9c4d2b9b09829fbf1abd06e0954fddee0b9ec8519854",
8
- "checkpoint_sha256": "e5ec5e1d494509f63960b6949af15be8390442271805ba45d796a1911b5c239a",
9
  "all_answer_objects_exact": true,
10
  "all_suite_metrics_exact": true,
11
  "bundled_code_imported_from_outside_workspace": true,
@@ -35,8 +35,8 @@
35
  "portable_adapter": {
36
  "base_model": "convaiinnovations/laya",
37
  "base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
38
- "seed": 1,
39
- "selected_epoch": 14
40
  }
41
  }
42
  }
 
1
  {
2
  "passed": true,
3
+ "candidate_seed": 4,
4
+ "selected_epoch": 8,
5
  "cases": 3516,
6
  "decisions": 5116,
7
  "benchmark_cases_sha256": "10fb671245cdac1ee09e9c4d2b9b09829fbf1abd06e0954fddee0b9ec8519854",
8
+ "checkpoint_sha256": "f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081",
9
  "all_answer_objects_exact": true,
10
  "all_suite_metrics_exact": true,
11
  "bundled_code_imported_from_outside_workspace": true,
 
35
  "portable_adapter": {
36
  "base_model": "convaiinnovations/laya",
37
  "base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
38
+ "seed": 4,
39
+ "selected_epoch": 8
40
  }
41
  }
42
  }
make_training_config.py ADDED
@@ -0,0 +1,37 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Resolve the pinned base checkpoint and create a reproducible training config."""
2
+
3
+ import argparse
4
+ import json
5
+ from pathlib import Path
6
+ from huggingface_hub import snapshot_download
7
+
8
+ parser = argparse.ArgumentParser(description=__doc__)
9
+ parser.add_argument("--seed", type=int, choices=range(5), required=True)
10
+ parser.add_argument("--device", default="cuda:0")
11
+ parser.add_argument("--out", required=True)
12
+ args = parser.parse_args()
13
+ root = Path(__file__).resolve().parent
14
+ metadata = json.loads((root / "adapter_config.json").read_text())
15
+ base = snapshot_download(
16
+ metadata["base_model"],
17
+ revision=metadata["base_revision"],
18
+ allow_patterns=["model.safetensors", "rl_agent_config.json", "encoder/*", "tokenizer/*"],
19
+ )
20
+ config = {
21
+ "name": f"Deterministic frozen base + linear, seed {args.seed} (UNTRAINED)",
22
+ "adapter": "python",
23
+ "mode": "reference",
24
+ "factory": "ariadne_bench.frozen_input_interface:create",
25
+ "options": {
26
+ "laya_model": str(Path(base).resolve()),
27
+ "device": args.device,
28
+ "batch_size": 16,
29
+ "max_len": 1024,
30
+ "head_max_len": 256,
31
+ "seed": args.seed,
32
+ "training_scope": "bridge",
33
+ "add_interface": True,
34
+ "deterministic": True,
35
+ },
36
+ }
37
+ Path(args.out).write_text(json.dumps(config, indent=2) + "\n")
metrics.json CHANGED
@@ -12,8 +12,8 @@
12
  "runs": {
13
  "0": {
14
  "seed": 0,
15
- "selected_epoch": 20,
16
- "elapsed_s": 1038.9482859019772,
17
  "initial_validation": {
18
  "soft_cross_entropy": 1.544869564374288,
19
  "accuracy": 0.36833333333333335,
@@ -21,14 +21,14 @@
21
  "decisions": 600
22
  },
23
  "selected_validation": {
24
- "soft_cross_entropy": 1.0292588464419048,
25
- "accuracy": 0.57,
26
- "brier_soft": 0.16918850486477216,
27
  "decisions": 600
28
  },
29
  "stopping": {
30
  "reason": "early_stopping",
31
- "epochs_completed": 23,
32
  "patience": 3,
33
  "min_epochs": 10,
34
  "epochs_without_improvement": 3,
@@ -42,159 +42,159 @@
42
  "valid": 2000,
43
  "failed": 0,
44
  "coverage": 1.0,
45
- "accuracy_all": 0.61,
46
- "accuracy_valid": 0.61,
47
- "ece_top_label": 0.15303136435282014,
48
- "mean_confidence": 0.4574502381919959,
49
- "brier_hard": 0.549798004153015,
50
  "brier_hard_n": 2000,
51
- "nll_hard": 0.9521041848093966,
52
  "nll_hard_n": 2000,
53
  "zero_probability_gold": 0.0,
54
  "zero_probability_gold_n": 2000,
55
- "soft_accuracy": 0.3883819759829014,
56
  "soft_accuracy_n": 2000,
57
- "brier_soft": 0.15223049605383382,
58
  "brier_soft_n": 2000,
59
- "kl_gold_to_prediction": 0.2712886346630264,
60
  "kl_gold_to_prediction_n": 2000,
61
- "total_variation": 0.2797391333751494,
62
  "total_variation_n": 2000,
63
- "score_mae": 0.45095113683604626,
64
  "score_mae_n": 800,
65
- "within_one_level": 0.915,
66
  "within_one_level_n": 800,
67
- "macro_f1": 0.4150680787086972
68
  },
69
  "ag-news": {
70
  "attempted": 600,
71
  "valid": 600,
72
  "failed": 0,
73
  "coverage": 1.0,
74
- "accuracy_all": 0.31333333333333335,
75
- "accuracy_valid": 0.31333333333333335,
76
- "ece_top_label": 0.025934925132124264,
77
- "mean_confidence": 0.28915707486787573,
78
- "brier_hard": 0.7441339007069468,
79
  "brier_hard_n": 600,
80
- "nll_hard": 1.3750333163467814,
81
  "nll_hard_n": 600,
82
  "zero_probability_gold": 0.0,
83
  "zero_probability_gold_n": 600,
84
- "macro_f1": 0.29084523228646375
85
  },
86
  "emotion": {
87
  "attempted": 600,
88
  "valid": 600,
89
  "failed": 0,
90
  "coverage": 1.0,
91
- "accuracy_all": 0.195,
92
- "accuracy_valid": 0.195,
93
- "ece_top_label": 0.08537782470783518,
94
- "mean_confidence": 0.2764666886168944,
95
- "brier_hard": 0.8415061960524678,
96
  "brier_hard_n": 600,
97
- "nll_hard": 1.810555559185162,
98
  "nll_hard_n": 600,
99
- "zero_probability_gold": 0.0,
100
  "zero_probability_gold_n": 600,
101
- "macro_f1": 0.12883846057517262
102
  },
103
  "boolq": {
104
  "attempted": 600,
105
  "valid": 600,
106
  "failed": 0,
107
  "coverage": 1.0,
108
- "accuracy_all": 0.5666666666666667,
109
- "accuracy_valid": 0.5666666666666667,
110
- "ece_top_label": 0.08635800000000002,
111
- "mean_confidence": 0.5258943333333334,
112
- "brier_hard": 0.5020554209333333,
113
  "brier_hard_n": 600,
114
- "nll_hard": 0.6953013099942155,
115
  "nll_hard_n": 600,
116
  "zero_probability_gold": 0.0,
117
  "zero_probability_gold_n": 600,
118
- "macro_f1": 0.5074762578298646
119
  },
120
  "sst5": {
121
  "attempted": 600,
122
  "valid": 600,
123
  "failed": 0,
124
  "coverage": 1.0,
125
- "accuracy_all": 0.15,
126
- "accuracy_valid": 0.15,
127
- "ece_top_label": 0.11939667407180776,
128
- "mean_confidence": 0.2693966740718078,
129
- "brier_hard": 0.8207044411796734,
130
  "brier_hard_n": 600,
131
- "nll_hard": 1.6589117409852412,
132
  "nll_hard_n": 600,
133
  "zero_probability_gold": 0.0,
134
  "zero_probability_gold_n": 600,
135
- "score_mae": 1.2247300886774997,
136
  "score_mae_n": 600,
137
- "within_one_level": 0.42,
138
  "within_one_level_n": 600,
139
- "macro_f1": 0.10629896865332711
140
  },
141
  "prompt-injections": {
142
  "attempted": 116,
143
  "valid": 116,
144
  "failed": 0,
145
  "coverage": 1.0,
146
- "accuracy_all": 0.5344827586206896,
147
- "accuracy_valid": 0.5344827586206896,
148
- "ece_top_label": 0.02089741379310345,
149
- "mean_confidence": 0.5135853448275862,
150
- "brier_hard": 0.5003080198275862,
151
  "brier_hard_n": 116,
152
- "nll_hard": 0.6934547870034389,
153
  "nll_hard_n": 116,
154
- "zero_probability_gold": 0.0,
155
  "zero_probability_gold_n": 116,
156
- "macro_f1": 0.513664596273292
157
  },
158
  "massive-intent.en": {
159
  "attempted": 300,
160
  "valid": 300,
161
  "failed": 0,
162
  "coverage": 1.0,
163
- "accuracy_all": 0.023333333333333334,
164
- "accuracy_valid": 0.023333333333333334,
165
- "ece_top_label": 0.09797495336730408,
166
- "mean_confidence": 0.12130828670063742,
167
- "brier_hard": 0.9672041702729498,
168
  "brier_hard_n": 300,
169
- "nll_hard": 3.1313647377288074,
170
  "nll_hard_n": 300,
171
- "zero_probability_gold": 0.0,
172
  "zero_probability_gold_n": 300,
173
- "macro_f1": 0.011215849106652133
174
  },
175
  "xnli.en": {
176
  "attempted": 300,
177
  "valid": 300,
178
  "failed": 0,
179
  "coverage": 1.0,
180
- "accuracy_all": 0.3433333333333333,
181
- "accuracy_valid": 0.3433333333333333,
182
- "ece_top_label": 0.05212130996336308,
183
- "mean_confidence": 0.39545464329669644,
184
- "brier_hard": 0.6747449551934047,
185
  "brier_hard_n": 300,
186
- "nll_hard": 1.1111148411541347,
187
  "nll_hard_n": 300,
188
  "zero_probability_gold": 0.0,
189
  "zero_probability_gold_n": 300,
190
- "macro_f1": 0.24805362074756226
191
  }
192
  }
193
  },
194
  "1": {
195
  "seed": 1,
196
- "selected_epoch": 14,
197
- "elapsed_s": 766.9112328969641,
198
  "initial_validation": {
199
  "soft_cross_entropy": 1.544869564374288,
200
  "accuracy": 0.36833333333333335,
@@ -202,14 +202,14 @@
202
  "decisions": 600
203
  },
204
  "selected_validation": {
205
- "soft_cross_entropy": 0.8377540612220764,
206
- "accuracy": 0.765,
207
- "brier_soft": 0.06276483290052662,
208
  "decisions": 600
209
  },
210
  "stopping": {
211
  "reason": "early_stopping",
212
- "epochs_completed": 17,
213
  "patience": 3,
214
  "min_epochs": 10,
215
  "epochs_without_improvement": 3,
@@ -223,159 +223,159 @@
223
  "valid": 2000,
224
  "failed": 0,
225
  "coverage": 1.0,
226
- "accuracy_all": 0.7695,
227
- "accuracy_valid": 0.7695,
228
- "ece_top_label": 0.2142237525419539,
229
- "mean_confidence": 0.555276247458046,
230
- "brier_hard": 0.3977842163576973,
231
  "brier_hard_n": 2000,
232
- "nll_hard": 0.7027635604099817,
233
  "nll_hard_n": 2000,
234
  "zero_probability_gold": 0.0,
235
  "zero_probability_gold_n": 2000,
236
- "soft_accuracy": 0.47103605907759827,
237
  "soft_accuracy_n": 2000,
238
- "brier_soft": 0.06255403710877039,
239
  "brier_soft_n": 2000,
240
- "kl_gold_to_prediction": 0.11643450461088396,
241
  "kl_gold_to_prediction_n": 2000,
242
- "total_variation": 0.17146599324503523,
243
  "total_variation_n": 2000,
244
- "score_mae": 0.2299944493828981,
245
  "score_mae_n": 800,
246
- "within_one_level": 0.98875,
247
  "within_one_level_n": 800,
248
- "macro_f1": 0.6487623885802034
249
  },
250
  "ag-news": {
251
  "attempted": 600,
252
  "valid": 600,
253
  "failed": 0,
254
  "coverage": 1.0,
255
- "accuracy_all": 0.9316666666666666,
256
- "accuracy_valid": 0.9316666666666666,
257
- "ece_top_label": 0.03387674147756575,
258
- "mean_confidence": 0.904708562141314,
259
- "brier_hard": 0.10911433540423408,
260
  "brier_hard_n": 600,
261
- "nll_hard": 0.21361121678114303,
262
  "nll_hard_n": 600,
263
  "zero_probability_gold": 0.0,
264
  "zero_probability_gold_n": 600,
265
- "macro_f1": 0.9283987145646934
266
  },
267
  "emotion": {
268
  "attempted": 600,
269
  "valid": 600,
270
  "failed": 0,
271
  "coverage": 1.0,
272
- "accuracy_all": 0.5733333333333334,
273
- "accuracy_valid": 0.5733333333333334,
274
- "ece_top_label": 0.3169210441851814,
275
- "mean_confidence": 0.8880587177380197,
276
- "brier_hard": 0.7264584486728032,
277
  "brier_hard_n": 600,
278
- "nll_hard": 2.146154603203373,
279
  "nll_hard_n": 600,
280
- "zero_probability_gold": 0.008333333333333333,
281
  "zero_probability_gold_n": 600,
282
- "macro_f1": 0.4862286665693279
283
  },
284
  "boolq": {
285
  "attempted": 600,
286
  "valid": 600,
287
  "failed": 0,
288
  "coverage": 1.0,
289
- "accuracy_all": 0.7983333333333333,
290
- "accuracy_valid": 0.7983333333333333,
291
- "ece_top_label": 0.09755433333333334,
292
- "mean_confidence": 0.8900093333333333,
293
- "brier_hard": 0.29822958146666667,
294
  "brier_hard_n": 600,
295
- "nll_hard": 0.4711253960780479,
296
  "nll_hard_n": 600,
297
  "zero_probability_gold": 0.0,
298
  "zero_probability_gold_n": 600,
299
- "macro_f1": 0.7813983878883868
300
  },
301
  "sst5": {
302
  "attempted": 600,
303
  "valid": 600,
304
  "failed": 0,
305
  "coverage": 1.0,
306
- "accuracy_all": 0.42,
307
- "accuracy_valid": 0.42,
308
- "ece_top_label": 0.1643168667044709,
309
- "mean_confidence": 0.5761336745094923,
310
- "brier_hard": 0.722245716690302,
311
  "brier_hard_n": 600,
312
- "nll_hard": 1.412054753482227,
313
  "nll_hard_n": 600,
314
  "zero_probability_gold": 0.0,
315
  "zero_probability_gold_n": 600,
316
- "score_mae": 0.7360310433394476,
317
  "score_mae_n": 600,
318
- "within_one_level": 0.7466666666666667,
319
  "within_one_level_n": 600,
320
- "macro_f1": 0.41387920942262574
321
  },
322
  "prompt-injections": {
323
  "attempted": 116,
324
  "valid": 116,
325
  "failed": 0,
326
  "coverage": 1.0,
327
- "accuracy_all": 0.7155172413793104,
328
- "accuracy_valid": 0.7155172413793104,
329
- "ece_top_label": 0.20010517241379305,
330
- "mean_confidence": 0.9105,
331
- "brier_hard": 0.43610344172413795,
332
  "brier_hard_n": 116,
333
- "nll_hard": 1.2355666268201304,
334
  "nll_hard_n": 116,
335
- "zero_probability_gold": 0.008620689655172414,
336
  "zero_probability_gold_n": 116,
337
- "macro_f1": 0.7038756091900673
338
  },
339
  "massive-intent.en": {
340
  "attempted": 300,
341
  "valid": 300,
342
  "failed": 0,
343
  "coverage": 1.0,
344
- "accuracy_all": 0.7433333333333333,
345
- "accuracy_valid": 0.7433333333333333,
346
- "ece_top_label": 0.1906540683935104,
347
- "mean_confidence": 0.9339874017268438,
348
- "brier_hard": 0.43706654758878133,
349
  "brier_hard_n": 300,
350
- "nll_hard": 2.7801978001907903,
351
  "nll_hard_n": 300,
352
- "zero_probability_gold": 0.07,
353
  "zero_probability_gold_n": 300,
354
- "macro_f1": 0.44346732036749054
355
  },
356
  "xnli.en": {
357
  "attempted": 300,
358
  "valid": 300,
359
  "failed": 0,
360
  "coverage": 1.0,
361
- "accuracy_all": 0.8766666666666667,
362
- "accuracy_valid": 0.8766666666666667,
363
- "ece_top_label": 0.055158283453726184,
364
- "mean_confidence": 0.913468962679153,
365
- "brier_hard": 0.19180470102420566,
366
  "brier_hard_n": 300,
367
- "nll_hard": 0.3708980570966955,
368
  "nll_hard_n": 300,
369
  "zero_probability_gold": 0.0,
370
  "zero_probability_gold_n": 300,
371
- "macro_f1": 0.8772028178860477
372
  }
373
  }
374
  },
375
  "2": {
376
  "seed": 2,
377
- "selected_epoch": 15,
378
- "elapsed_s": 813.5451362769818,
379
  "initial_validation": {
380
  "soft_cross_entropy": 1.544869564374288,
381
  "accuracy": 0.36833333333333335,
@@ -383,14 +383,14 @@
383
  "decisions": 600
384
  },
385
  "selected_validation": {
386
- "soft_cross_entropy": 0.9867499772707621,
387
- "accuracy": 0.6316666666666667,
388
- "brier_soft": 0.14927835414807003,
389
  "decisions": 600
390
  },
391
  "stopping": {
392
  "reason": "early_stopping",
393
- "epochs_completed": 18,
394
  "patience": 3,
395
  "min_epochs": 10,
396
  "epochs_without_improvement": 3,
@@ -404,159 +404,159 @@
404
  "valid": 2000,
405
  "failed": 0,
406
  "coverage": 1.0,
407
- "accuracy_all": 0.661,
408
- "accuracy_valid": 0.661,
409
- "ece_top_label": 0.17250655060654457,
410
- "mean_confidence": 0.48883668313682976,
411
- "brier_hard": 0.5040496940863629,
412
  "brier_hard_n": 2000,
413
- "nll_hard": 0.8694405673021295,
414
  "nll_hard_n": 2000,
415
  "zero_probability_gold": 0.0,
416
  "zero_probability_gold_n": 2000,
417
- "soft_accuracy": 0.41427322765683283,
418
  "soft_accuracy_n": 2000,
419
- "brier_soft": 0.12329331823446005,
420
  "brier_soft_n": 2000,
421
- "kl_gold_to_prediction": 0.2167589625120186,
422
  "kl_gold_to_prediction_n": 2000,
423
- "total_variation": 0.24557641844905673,
424
  "total_variation_n": 2000,
425
- "score_mae": 0.38257377026959916,
426
  "score_mae_n": 800,
427
- "within_one_level": 0.92625,
428
  "within_one_level_n": 800,
429
- "macro_f1": 0.511489281675723
430
  },
431
  "ag-news": {
432
  "attempted": 600,
433
  "valid": 600,
434
  "failed": 0,
435
  "coverage": 1.0,
436
- "accuracy_all": 0.2633333333333333,
437
- "accuracy_valid": 0.2633333333333333,
438
- "ece_top_label": 0.0415880758215024,
439
- "mean_confidence": 0.30492140915483573,
440
- "brier_hard": 0.7533353201524136,
441
  "brier_hard_n": 600,
442
- "nll_hard": 1.392051963909145,
443
  "nll_hard_n": 600,
444
  "zero_probability_gold": 0.0,
445
  "zero_probability_gold_n": 600,
446
- "macro_f1": 0.2196261943166739
447
  },
448
  "emotion": {
449
  "attempted": 600,
450
  "valid": 600,
451
  "failed": 0,
452
  "coverage": 1.0,
453
- "accuracy_all": 0.04833333333333333,
454
- "accuracy_valid": 0.04833333333333333,
455
- "ece_top_label": 0.21796334391120661,
456
- "mean_confidence": 0.2662966772445399,
457
- "brier_hard": 0.9083949711964834,
458
  "brier_hard_n": 600,
459
- "nll_hard": 2.0132867760168445,
460
  "nll_hard_n": 600,
461
- "zero_probability_gold": 0.0,
462
  "zero_probability_gold_n": 600,
463
- "macro_f1": 0.03431938431938432
464
  },
465
  "boolq": {
466
  "attempted": 600,
467
  "valid": 600,
468
  "failed": 0,
469
  "coverage": 1.0,
470
- "accuracy_all": 0.49833333333333335,
471
- "accuracy_valid": 0.49833333333333335,
472
- "ece_top_label": 0.03133150000000007,
473
- "mean_confidence": 0.5290545,
474
- "brier_hard": 0.5040687543666666,
475
  "brier_hard_n": 600,
476
- "nll_hard": 0.6972355680599187,
477
  "nll_hard_n": 600,
478
  "zero_probability_gold": 0.0,
479
  "zero_probability_gold_n": 600,
480
- "macro_f1": 0.49597983919356775
481
  },
482
  "sst5": {
483
  "attempted": 600,
484
  "valid": 600,
485
  "failed": 0,
486
  "coverage": 1.0,
487
- "accuracy_all": 0.22666666666666666,
488
- "accuracy_valid": 0.22666666666666666,
489
- "ece_top_label": 0.08382145968032653,
490
- "mean_confidence": 0.30788872429147607,
491
- "brier_hard": 0.8111226172703578,
492
  "brier_hard_n": 600,
493
- "nll_hard": 1.6364255257258482,
494
  "nll_hard_n": 600,
495
  "zero_probability_gold": 0.0,
496
  "zero_probability_gold_n": 600,
497
- "score_mae": 1.223873285603664,
498
  "score_mae_n": 600,
499
- "within_one_level": 0.39,
500
  "within_one_level_n": 600,
501
- "macro_f1": 0.12552908285983264
502
  },
503
  "prompt-injections": {
504
  "attempted": 116,
505
  "valid": 116,
506
  "failed": 0,
507
  "coverage": 1.0,
508
- "accuracy_all": 0.3448275862068966,
509
- "accuracy_valid": 0.3448275862068966,
510
- "ece_top_label": 0.1914112068965517,
511
- "mean_confidence": 0.5362387931034482,
512
- "brier_hard": 0.5323824856896552,
513
  "brier_hard_n": 116,
514
- "nll_hard": 0.7258214322004152,
515
  "nll_hard_n": 116,
516
- "zero_probability_gold": 0.0,
517
  "zero_probability_gold_n": 116,
518
- "macro_f1": 0.275
519
  },
520
  "massive-intent.en": {
521
  "attempted": 300,
522
  "valid": 300,
523
  "failed": 0,
524
  "coverage": 1.0,
525
- "accuracy_all": 0.06333333333333334,
526
- "accuracy_valid": 0.06333333333333334,
527
- "ece_top_label": 0.03411377590323972,
528
- "mean_confidence": 0.09744710923657304,
529
- "brier_hard": 0.9608971275390831,
530
  "brier_hard_n": 300,
531
- "nll_hard": 3.135262799945936,
532
  "nll_hard_n": 300,
533
- "zero_probability_gold": 0.0,
534
  "zero_probability_gold_n": 300,
535
- "macro_f1": 0.02907709642455108
536
  },
537
  "xnli.en": {
538
  "attempted": 300,
539
  "valid": 300,
540
  "failed": 0,
541
  "coverage": 1.0,
542
- "accuracy_all": 0.3433333333333333,
543
- "accuracy_valid": 0.3433333333333333,
544
- "ece_top_label": 0.04294870734165375,
545
- "mean_confidence": 0.38628204067498706,
546
- "brier_hard": 0.6761361630010466,
547
  "brier_hard_n": 300,
548
- "nll_hard": 1.1128104653939794,
549
  "nll_hard_n": 300,
550
  "zero_probability_gold": 0.0,
551
  "zero_probability_gold_n": 300,
552
- "macro_f1": 0.3099922839506173
553
  }
554
  }
555
  },
556
  "3": {
557
  "seed": 3,
558
- "selected_epoch": 24,
559
- "elapsed_s": 1226.4128511130111,
560
  "initial_validation": {
561
  "soft_cross_entropy": 1.544869564374288,
562
  "accuracy": 0.36833333333333335,
@@ -564,14 +564,14 @@
564
  "decisions": 600
565
  },
566
  "selected_validation": {
567
- "soft_cross_entropy": 0.9105557099978129,
568
- "accuracy": 0.715,
569
- "brier_soft": 0.1013037375236551,
570
  "decisions": 600
571
  },
572
  "stopping": {
573
  "reason": "early_stopping",
574
- "epochs_completed": 27,
575
  "patience": 3,
576
  "min_epochs": 10,
577
  "epochs_without_improvement": 3,
@@ -585,159 +585,159 @@
585
  "valid": 2000,
586
  "failed": 0,
587
  "coverage": 1.0,
588
- "accuracy_all": 0.7145,
589
- "accuracy_valid": 0.7145,
590
- "ece_top_label": 0.18500231281217855,
591
- "mean_confidence": 0.5300184130804106,
592
- "brier_hard": 0.45082257101028994,
593
  "brier_hard_n": 2000,
594
- "nll_hard": 0.7891005479512444,
595
  "nll_hard_n": 2000,
596
  "zero_probability_gold": 0.0,
597
  "zero_probability_gold_n": 2000,
598
- "soft_accuracy": 0.4447570518885741,
599
  "soft_accuracy_n": 2000,
600
- "brier_soft": 0.09325671264191153,
601
  "brier_soft_n": 2000,
602
- "kl_gold_to_prediction": 0.17157487793178008,
603
  "kl_gold_to_prediction_n": 2000,
604
- "total_variation": 0.21097575386758916,
605
  "total_variation_n": 2000,
606
- "score_mae": 0.32394138130337413,
607
  "score_mae_n": 800,
608
- "within_one_level": 0.96375,
609
  "within_one_level_n": 800,
610
- "macro_f1": 0.5630278027844289
611
  },
612
  "ag-news": {
613
  "attempted": 600,
614
  "valid": 600,
615
  "failed": 0,
616
  "coverage": 1.0,
617
- "accuracy_all": 0.285,
618
- "accuracy_valid": 0.285,
619
- "ece_top_label": 0.03904648151058483,
620
- "mean_confidence": 0.3240464815105848,
621
- "brier_hard": 0.7625373960165811,
622
  "brier_hard_n": 600,
623
- "nll_hard": 1.4140716749662658,
624
  "nll_hard_n": 600,
625
  "zero_probability_gold": 0.0,
626
  "zero_probability_gold_n": 600,
627
- "macro_f1": 0.25772655273033623
628
  },
629
  "emotion": {
630
  "attempted": 600,
631
  "valid": 600,
632
  "failed": 0,
633
  "coverage": 1.0,
634
- "accuracy_all": 0.135,
635
- "accuracy_valid": 0.135,
636
- "ece_top_label": 0.31863940813484304,
637
- "mean_confidence": 0.45363940813484305,
638
- "brier_hard": 0.9777635802768799,
639
  "brier_hard_n": 600,
640
- "nll_hard": 2.137352366621794,
641
  "nll_hard_n": 600,
642
- "zero_probability_gold": 0.0,
643
  "zero_probability_gold_n": 600,
644
- "macro_f1": 0.06590661903472231
645
  },
646
  "boolq": {
647
  "attempted": 600,
648
  "valid": 600,
649
  "failed": 0,
650
  "coverage": 1.0,
651
- "accuracy_all": 0.48833333333333334,
652
- "accuracy_valid": 0.48833333333333334,
653
- "ece_top_label": 0.10787050000000001,
654
- "mean_confidence": 0.5912141666666667,
655
- "brier_hard": 0.5250116893,
656
  "brier_hard_n": 600,
657
- "nll_hard": 0.719837552287629,
658
  "nll_hard_n": 600,
659
  "zero_probability_gold": 0.0,
660
  "zero_probability_gold_n": 600,
661
- "macro_f1": 0.48696381173075903
662
  },
663
  "sst5": {
664
  "attempted": 600,
665
  "valid": 600,
666
  "failed": 0,
667
  "coverage": 1.0,
668
- "accuracy_all": 0.25666666666666665,
669
- "accuracy_valid": 0.25666666666666665,
670
- "ece_top_label": 0.05066152500332892,
671
- "mean_confidence": 0.3073281916699956,
672
- "brier_hard": 0.8215936069738192,
673
  "brier_hard_n": 600,
674
- "nll_hard": 1.6775697808435297,
675
  "nll_hard_n": 600,
676
  "zero_probability_gold": 0.0,
677
  "zero_probability_gold_n": 600,
678
- "score_mae": 1.245226132804566,
679
  "score_mae_n": 600,
680
- "within_one_level": 0.44333333333333336,
681
  "within_one_level_n": 600,
682
- "macro_f1": 0.13260195835186606
683
  },
684
  "prompt-injections": {
685
  "attempted": 116,
686
  "valid": 116,
687
  "failed": 0,
688
  "coverage": 1.0,
689
- "accuracy_all": 0.5689655172413793,
690
- "accuracy_valid": 0.5689655172413793,
691
- "ece_top_label": 0.07609396551724136,
692
- "mean_confidence": 0.6450594827586207,
693
- "brier_hard": 0.4881634129310345,
694
  "brier_hard_n": 116,
695
- "nll_hard": 0.6835747994579267,
696
  "nll_hard_n": 116,
697
- "zero_probability_gold": 0.0,
698
  "zero_probability_gold_n": 116,
699
- "macro_f1": 0.5461658841940532
700
  },
701
  "massive-intent.en": {
702
  "attempted": 300,
703
  "valid": 300,
704
  "failed": 0,
705
  "coverage": 1.0,
706
- "accuracy_all": 0.023333333333333334,
707
- "accuracy_valid": 0.023333333333333334,
708
- "ece_top_label": 0.14827139054632427,
709
- "mean_confidence": 0.1716047238796576,
710
- "brier_hard": 1.012268483556247,
711
  "brier_hard_n": 300,
712
- "nll_hard": 3.540188562309456,
713
  "nll_hard_n": 300,
714
- "zero_probability_gold": 0.0,
715
  "zero_probability_gold_n": 300,
716
- "macro_f1": 0.014601824457593688
717
  },
718
  "xnli.en": {
719
  "attempted": 300,
720
  "valid": 300,
721
  "failed": 0,
722
  "coverage": 1.0,
723
- "accuracy_all": 0.33666666666666667,
724
- "accuracy_valid": 0.33666666666666667,
725
- "ece_top_label": 0.21980800490067007,
726
- "mean_confidence": 0.5564746715673368,
727
- "brier_hard": 0.7709342547858443,
728
  "brier_hard_n": 300,
729
- "nll_hard": 1.2650663651832452,
730
  "nll_hard_n": 300,
731
  "zero_probability_gold": 0.0,
732
  "zero_probability_gold_n": 300,
733
- "macro_f1": 0.20238216957239646
734
  }
735
  }
736
  },
737
  "4": {
738
  "seed": 4,
739
- "selected_epoch": 7,
740
- "elapsed_s": 452.7137592760264,
741
  "initial_validation": {
742
  "soft_cross_entropy": 1.544869564374288,
743
  "accuracy": 0.36833333333333335,
@@ -745,14 +745,14 @@
745
  "decisions": 600
746
  },
747
  "selected_validation": {
748
- "soft_cross_entropy": 0.8448993345101674,
749
- "accuracy": 0.7533333333333333,
750
- "brier_soft": 0.0668264798882107,
751
  "decisions": 600
752
  },
753
  "stopping": {
754
  "reason": "early_stopping",
755
- "epochs_completed": 10,
756
  "patience": 3,
757
  "min_epochs": 10,
758
  "epochs_without_improvement": 3,
@@ -766,152 +766,152 @@
766
  "valid": 2000,
767
  "failed": 0,
768
  "coverage": 1.0,
769
- "accuracy_all": 0.7565,
770
- "accuracy_valid": 0.7565,
771
- "ece_top_label": 0.20311433485058888,
772
- "mean_confidence": 0.5533856651494111,
773
- "brier_hard": 0.4063906772103004,
774
  "brier_hard_n": 2000,
775
- "nll_hard": 0.7152837956058272,
776
  "nll_hard_n": 2000,
777
  "zero_probability_gold": 0.0,
778
  "zero_probability_gold_n": 2000,
779
- "soft_accuracy": 0.46913207673904656,
780
  "soft_accuracy_n": 2000,
781
- "brier_soft": 0.06497074467877281,
782
  "brier_soft_n": 2000,
783
- "kl_gold_to_prediction": 0.12038107186141371,
784
  "kl_gold_to_prediction_n": 2000,
785
- "total_variation": 0.17619801386776057,
786
  "total_variation_n": 2000,
787
- "score_mae": 0.2447599606823941,
788
  "score_mae_n": 800,
789
- "within_one_level": 0.99375,
790
  "within_one_level_n": 800,
791
- "macro_f1": 0.6349242079343602
792
  },
793
  "ag-news": {
794
  "attempted": 600,
795
  "valid": 600,
796
  "failed": 0,
797
  "coverage": 1.0,
798
- "accuracy_all": 0.935,
799
- "accuracy_valid": 0.935,
800
- "ece_top_label": 0.0277188639855619,
801
- "mean_confidence": 0.9106097216911906,
802
- "brier_hard": 0.10223891825148304,
803
  "brier_hard_n": 600,
804
- "nll_hard": 0.2046134786862912,
805
  "nll_hard_n": 600,
806
  "zero_probability_gold": 0.0,
807
  "zero_probability_gold_n": 600,
808
- "macro_f1": 0.9319475717225341
809
  },
810
  "emotion": {
811
  "attempted": 600,
812
  "valid": 600,
813
  "failed": 0,
814
  "coverage": 1.0,
815
- "accuracy_all": 0.565,
816
- "accuracy_valid": 0.565,
817
- "ece_top_label": 0.3127789762695219,
818
- "mean_confidence": 0.8750361088494466,
819
- "brier_hard": 0.7278820816911535,
820
  "brier_hard_n": 600,
821
- "nll_hard": 2.179428847187102,
822
  "nll_hard_n": 600,
823
- "zero_probability_gold": 0.01,
824
  "zero_probability_gold_n": 600,
825
- "macro_f1": 0.47063728412574746
826
  },
827
  "boolq": {
828
  "attempted": 600,
829
  "valid": 600,
830
  "failed": 0,
831
  "coverage": 1.0,
832
- "accuracy_all": 0.8116666666666666,
833
- "accuracy_valid": 0.8116666666666666,
834
- "ece_top_label": 0.08128733333333335,
835
- "mean_confidence": 0.8882866666666667,
836
- "brier_hard": 0.2800035268666667,
837
  "brier_hard_n": 600,
838
- "nll_hard": 0.4526075086973918,
839
  "nll_hard_n": 600,
840
- "zero_probability_gold": 0.0,
841
  "zero_probability_gold_n": 600,
842
- "macro_f1": 0.7989317880539385
843
  },
844
  "sst5": {
845
  "attempted": 600,
846
  "valid": 600,
847
  "failed": 0,
848
  "coverage": 1.0,
849
- "accuracy_all": 0.47333333333333333,
850
- "accuracy_valid": 0.47333333333333333,
851
- "ece_top_label": 0.11648579870821457,
852
- "mean_confidence": 0.5717307193165773,
853
- "brier_hard": 0.6735168048802583,
854
  "brier_hard_n": 600,
855
- "nll_hard": 1.2841517570209144,
856
  "nll_hard_n": 600,
857
  "zero_probability_gold": 0.0,
858
  "zero_probability_gold_n": 600,
859
- "score_mae": 0.6724125988264632,
860
  "score_mae_n": 600,
861
- "within_one_level": 0.7883333333333333,
862
  "within_one_level_n": 600,
863
- "macro_f1": 0.46002232794452175
864
  },
865
  "prompt-injections": {
866
  "attempted": 116,
867
  "valid": 116,
868
  "failed": 0,
869
  "coverage": 1.0,
870
- "accuracy_all": 0.7241379310344828,
871
- "accuracy_valid": 0.7241379310344828,
872
- "ece_top_label": 0.22484482758620694,
873
- "mean_confidence": 0.9346275862068966,
874
- "brier_hard": 0.44889866172413795,
875
  "brier_hard_n": 116,
876
- "nll_hard": 1.682957724359292,
877
  "nll_hard_n": 116,
878
- "zero_probability_gold": 0.02586206896551724,
879
  "zero_probability_gold_n": 116,
880
- "macro_f1": 0.7118012422360249
881
  },
882
  "massive-intent.en": {
883
  "attempted": 300,
884
  "valid": 300,
885
  "failed": 0,
886
  "coverage": 1.0,
887
- "accuracy_all": 0.7733333333333333,
888
- "accuracy_valid": 0.7733333333333333,
889
- "ece_top_label": 0.1762748719939652,
890
- "mean_confidence": 0.9448706790799233,
891
- "brier_hard": 0.3914895786578462,
892
  "brier_hard_n": 300,
893
- "nll_hard": 2.594218156353015,
894
  "nll_hard_n": 300,
895
- "zero_probability_gold": 0.06333333333333334,
896
  "zero_probability_gold_n": 300,
897
- "macro_f1": 0.4611733462473961
898
  },
899
  "xnli.en": {
900
  "attempted": 300,
901
  "valid": 300,
902
  "failed": 0,
903
  "coverage": 1.0,
904
- "accuracy_all": 0.8766666666666667,
905
- "accuracy_valid": 0.8766666666666667,
906
- "ece_top_label": 0.059115758529190925,
907
- "mean_confidence": 0.9151327612640476,
908
- "brier_hard": 0.1964252077913328,
909
  "brier_hard_n": 300,
910
- "nll_hard": 0.36246740512671244,
911
  "nll_hard_n": 300,
912
  "zero_probability_gold": 0.0,
913
  "zero_probability_gold_n": 300,
914
- "macro_f1": 0.8760253913421515
915
  }
916
  }
917
  }
@@ -919,54 +919,54 @@
919
  "summary": {
920
  "typed-decisions": {
921
  "n": 5,
922
- "mean_percent": 70.23,
923
- "sample_variance_pp_squared": 44.56824999999997,
924
- "sample_std_pp": 6.675945625901995
925
  },
926
  "ag-news": {
927
  "n": 5,
928
- "mean_percent": 54.56666666666666,
929
- "sample_variance_pp_squared": 1255.536111111111,
930
- "sample_std_pp": 35.433544997799906
931
  },
932
  "boolq": {
933
  "n": 5,
934
- "mean_percent": 63.266666666666666,
935
- "sample_variance_pp_squared": 256.79999999999995,
936
- "sample_std_pp": 16.0249804992081
937
  },
938
  "emotion": {
939
  "n": 5,
940
- "mean_percent": 30.333333333333332,
941
- "sample_variance_pp_squared": 616.1666666666666,
942
- "sample_std_pp": 24.82270466058577
943
  },
944
  "prompt-injections": {
945
  "n": 5,
946
- "mean_percent": 57.758620689655174,
947
- "sample_variance_pp_squared": 241.5279429250891,
948
- "sample_std_pp": 15.541169290793055
949
  },
950
  "sst5": {
951
  "n": 5,
952
- "mean_percent": 30.53333333333333,
953
- "sample_variance_pp_squared": 185.14444444444447,
954
- "sample_std_pp": 13.606779356057938
955
  },
956
  "massive-intent.en": {
957
  "n": 5,
958
- "mean_percent": 32.53333333333333,
959
- "sample_variance_pp_squared": 1566.1999999999998,
960
- "sample_std_pp": 39.57524478761944
961
  },
962
  "xnli.en": {
963
  "n": 5,
964
- "mean_percent": 55.53333333333334,
965
- "sample_variance_pp_squared": 860.5333333333334,
966
- "sample_std_pp": 29.33484844571953
967
  }
968
  },
969
- "candidate_seed": 1,
970
  "variance_definition": "sample variance, n-1, percentage points squared",
971
  "benchmark_manifest": {
972
  "format_version": 1,
 
12
  "runs": {
13
  "0": {
14
  "seed": 0,
15
+ "selected_epoch": 11,
16
+ "elapsed_s": 612.5042744020466,
17
  "initial_validation": {
18
  "soft_cross_entropy": 1.544869564374288,
19
  "accuracy": 0.36833333333333335,
 
21
  "decisions": 600
22
  },
23
  "selected_validation": {
24
+ "soft_cross_entropy": 0.8771124110619227,
25
+ "accuracy": 0.7116666666666667,
26
+ "brier_soft": 0.0828241604194045,
27
  "decisions": 600
28
  },
29
  "stopping": {
30
  "reason": "early_stopping",
31
+ "epochs_completed": 14,
32
  "patience": 3,
33
  "min_epochs": 10,
34
  "epochs_without_improvement": 3,
 
42
  "valid": 2000,
43
  "failed": 0,
44
  "coverage": 1.0,
45
+ "accuracy_all": 0.7425,
46
+ "accuracy_valid": 0.7425,
47
+ "ece_top_label": 0.19222331993705236,
48
+ "mean_confidence": 0.5502766800629476,
49
+ "brier_hard": 0.41714854154295294,
50
  "brier_hard_n": 2000,
51
+ "nll_hard": 0.7346455747590726,
52
  "nll_hard_n": 2000,
53
  "zero_probability_gold": 0.0,
54
  "zero_probability_gold_n": 2000,
55
+ "soft_accuracy": 0.4642362342348646,
56
  "soft_accuracy_n": 2000,
57
+ "brier_soft": 0.07282817778125876,
58
  "brier_soft_n": 2000,
59
+ "kl_gold_to_prediction": 0.13635658195747785,
60
  "kl_gold_to_prediction_n": 2000,
61
+ "total_variation": 0.18626503447270873,
62
  "total_variation_n": 2000,
63
+ "score_mae": 0.2670681801272688,
64
  "score_mae_n": 800,
65
+ "within_one_level": 0.98125,
66
  "within_one_level_n": 800,
67
+ "macro_f1": 0.6211545025911077
68
  },
69
  "ag-news": {
70
  "attempted": 600,
71
  "valid": 600,
72
  "failed": 0,
73
  "coverage": 1.0,
74
+ "accuracy_all": 0.9416666666666667,
75
+ "accuracy_valid": 0.9416666666666667,
76
+ "ece_top_label": 0.032780915974200804,
77
+ "mean_confidence": 0.9120336837609119,
78
+ "brier_hard": 0.0911117485859852,
79
  "brier_hard_n": 600,
80
+ "nll_hard": 0.18384446182116462,
81
  "nll_hard_n": 600,
82
  "zero_probability_gold": 0.0,
83
  "zero_probability_gold_n": 600,
84
+ "macro_f1": 0.9391643390047506
85
  },
86
  "emotion": {
87
  "attempted": 600,
88
  "valid": 600,
89
  "failed": 0,
90
  "coverage": 1.0,
91
+ "accuracy_all": 0.585,
92
+ "accuracy_valid": 0.585,
93
+ "ece_top_label": 0.3290210797041948,
94
+ "mean_confidence": 0.9044507912979094,
95
+ "brier_hard": 0.7352614358764159,
96
  "brier_hard_n": 600,
97
+ "nll_hard": 2.325969372091789,
98
  "nll_hard_n": 600,
99
+ "zero_probability_gold": 0.01,
100
  "zero_probability_gold_n": 600,
101
+ "macro_f1": 0.4846329953255675
102
  },
103
  "boolq": {
104
  "attempted": 600,
105
  "valid": 600,
106
  "failed": 0,
107
  "coverage": 1.0,
108
+ "accuracy_all": 0.8083333333333333,
109
+ "accuracy_valid": 0.8083333333333333,
110
+ "ece_top_label": 0.09874549999999996,
111
+ "mean_confidence": 0.9070788333333333,
112
+ "brier_hard": 0.2864656637,
113
  "brier_hard_n": 600,
114
+ "nll_hard": 0.45967158478397707,
115
  "nll_hard_n": 600,
116
  "zero_probability_gold": 0.0,
117
  "zero_probability_gold_n": 600,
118
+ "macro_f1": 0.7946275764565816
119
  },
120
  "sst5": {
121
  "attempted": 600,
122
  "valid": 600,
123
  "failed": 0,
124
  "coverage": 1.0,
125
+ "accuracy_all": 0.38166666666666665,
126
+ "accuracy_valid": 0.38166666666666665,
127
+ "ece_top_label": 0.23314279850468647,
128
+ "mean_confidence": 0.614809465171353,
129
+ "brier_hard": 0.7997842513426261,
130
  "brier_hard_n": 600,
131
+ "nll_hard": 1.7175170777887119,
132
  "nll_hard_n": 600,
133
  "zero_probability_gold": 0.0,
134
  "zero_probability_gold_n": 600,
135
+ "score_mae": 0.9144952235712186,
136
  "score_mae_n": 600,
137
+ "within_one_level": 0.6466666666666666,
138
  "within_one_level_n": 600,
139
+ "macro_f1": 0.3319436236284081
140
  },
141
  "prompt-injections": {
142
  "attempted": 116,
143
  "valid": 116,
144
  "failed": 0,
145
  "coverage": 1.0,
146
+ "accuracy_all": 0.6982758620689655,
147
+ "accuracy_valid": 0.6982758620689655,
148
+ "ece_top_label": 0.25387413793103447,
149
+ "mean_confidence": 0.95215,
150
+ "brier_hard": 0.5115509806896552,
151
  "brier_hard_n": 116,
152
+ "nll_hard": 2.710472238032788,
153
  "nll_hard_n": 116,
154
+ "zero_probability_gold": 0.06896551724137931,
155
  "zero_probability_gold_n": 116,
156
+ "macro_f1": 0.6809931641392315
157
  },
158
  "massive-intent.en": {
159
  "attempted": 300,
160
  "valid": 300,
161
  "failed": 0,
162
  "coverage": 1.0,
163
+ "accuracy_all": 0.79,
164
+ "accuracy_valid": 0.79,
165
+ "ece_top_label": 0.16259349215790048,
166
+ "mean_confidence": 0.9477817275562629,
167
+ "brier_hard": 0.3655905743529746,
168
  "brier_hard_n": 300,
169
+ "nll_hard": 2.4460407722040176,
170
  "nll_hard_n": 300,
171
+ "zero_probability_gold": 0.06,
172
  "zero_probability_gold_n": 300,
173
+ "macro_f1": 0.5695724937908466
174
  },
175
  "xnli.en": {
176
  "attempted": 300,
177
  "valid": 300,
178
  "failed": 0,
179
  "coverage": 1.0,
180
+ "accuracy_all": 0.89,
181
+ "accuracy_valid": 0.89,
182
+ "ece_top_label": 0.05910627327283609,
183
+ "mean_confidence": 0.9123050519493905,
184
+ "brier_hard": 0.18251872229879984,
185
  "brier_hard_n": 300,
186
+ "nll_hard": 0.33404125735595414,
187
  "nll_hard_n": 300,
188
  "zero_probability_gold": 0.0,
189
  "zero_probability_gold_n": 300,
190
+ "macro_f1": 0.8905240397272149
191
  }
192
  }
193
  },
194
  "1": {
195
  "seed": 1,
196
+ "selected_epoch": 9,
197
+ "elapsed_s": 541.4585086829611,
198
  "initial_validation": {
199
  "soft_cross_entropy": 1.544869564374288,
200
  "accuracy": 0.36833333333333335,
 
202
  "decisions": 600
203
  },
204
  "selected_validation": {
205
+ "soft_cross_entropy": 0.8583925066391627,
206
+ "accuracy": 0.7466666666666667,
207
+ "brier_soft": 0.07437195796364297,
208
  "decisions": 600
209
  },
210
  "stopping": {
211
  "reason": "early_stopping",
212
+ "epochs_completed": 12,
213
  "patience": 3,
214
  "min_epochs": 10,
215
  "epochs_without_improvement": 3,
 
223
  "valid": 2000,
224
  "failed": 0,
225
  "coverage": 1.0,
226
+ "accuracy_all": 0.7565,
227
+ "accuracy_valid": 0.7565,
228
+ "ece_top_label": 0.19864268503011012,
229
+ "mean_confidence": 0.5578573149698899,
230
+ "brier_hard": 0.40482086772099385,
231
  "brier_hard_n": 2000,
232
+ "nll_hard": 0.7162594531200565,
233
  "nll_hard_n": 2000,
234
  "zero_probability_gold": 0.0,
235
  "zero_probability_gold_n": 2000,
236
+ "soft_accuracy": 0.469737413364262,
237
  "soft_accuracy_n": 2000,
238
+ "brier_soft": 0.0659469654317038,
239
  "brier_soft_n": 2000,
240
+ "kl_gold_to_prediction": 0.12376910578634269,
241
  "kl_gold_to_prediction_n": 2000,
242
+ "total_variation": 0.17750424596946363,
243
  "total_variation_n": 2000,
244
+ "score_mae": 0.24687533223920302,
245
  "score_mae_n": 800,
246
+ "within_one_level": 0.98625,
247
  "within_one_level_n": 800,
248
+ "macro_f1": 0.6300073343403317
249
  },
250
  "ag-news": {
251
  "attempted": 600,
252
  "valid": 600,
253
  "failed": 0,
254
  "coverage": 1.0,
255
+ "accuracy_all": 0.9483333333333334,
256
+ "accuracy_valid": 0.9483333333333334,
257
+ "ece_top_label": 0.03462398162033311,
258
+ "mean_confidence": 0.9188607247165506,
259
+ "brier_hard": 0.08433672632610516,
260
  "brier_hard_n": 600,
261
+ "nll_hard": 0.17417200848974151,
262
  "nll_hard_n": 600,
263
  "zero_probability_gold": 0.0,
264
  "zero_probability_gold_n": 600,
265
+ "macro_f1": 0.9456145576151846
266
  },
267
  "emotion": {
268
  "attempted": 600,
269
  "valid": 600,
270
  "failed": 0,
271
  "coverage": 1.0,
272
+ "accuracy_all": 0.5716666666666667,
273
+ "accuracy_valid": 0.5716666666666667,
274
+ "ece_top_label": 0.33449032496578573,
275
+ "mean_confidence": 0.90098409027152,
276
+ "brier_hard": 0.7328256742449804,
277
  "brier_hard_n": 600,
278
+ "nll_hard": 2.388523752936007,
279
  "nll_hard_n": 600,
280
+ "zero_probability_gold": 0.013333333333333334,
281
  "zero_probability_gold_n": 600,
282
+ "macro_f1": 0.4746076611364094
283
  },
284
  "boolq": {
285
  "attempted": 600,
286
  "valid": 600,
287
  "failed": 0,
288
  "coverage": 1.0,
289
+ "accuracy_all": 0.8266666666666667,
290
+ "accuracy_valid": 0.8266666666666667,
291
+ "ece_top_label": 0.08269850000000001,
292
+ "mean_confidence": 0.9093651666666667,
293
+ "brier_hard": 0.26852242536666665,
294
  "brier_hard_n": 600,
295
+ "nll_hard": 0.44571515698724307,
296
  "nll_hard_n": 600,
297
  "zero_probability_gold": 0.0,
298
  "zero_probability_gold_n": 600,
299
+ "macro_f1": 0.8140998140998141
300
  },
301
  "sst5": {
302
  "attempted": 600,
303
  "valid": 600,
304
  "failed": 0,
305
  "coverage": 1.0,
306
+ "accuracy_all": 0.4483333333333333,
307
+ "accuracy_valid": 0.4483333333333333,
308
+ "ece_top_label": 0.1552169576555402,
309
+ "mean_confidence": 0.5963372951766569,
310
+ "brier_hard": 0.715830700927659,
311
  "brier_hard_n": 600,
312
+ "nll_hard": 1.4312799777899081,
313
  "nll_hard_n": 600,
314
  "zero_probability_gold": 0.0,
315
  "zero_probability_gold_n": 600,
316
+ "score_mae": 0.763809024216731,
317
  "score_mae_n": 600,
318
+ "within_one_level": 0.7216666666666667,
319
  "within_one_level_n": 600,
320
+ "macro_f1": 0.4241661418376961
321
  },
322
  "prompt-injections": {
323
  "attempted": 116,
324
  "valid": 116,
325
  "failed": 0,
326
  "coverage": 1.0,
327
+ "accuracy_all": 0.7844827586206896,
328
+ "accuracy_valid": 0.7844827586206896,
329
+ "ece_top_label": 0.17151379310344822,
330
+ "mean_confidence": 0.9531103448275862,
331
+ "brier_hard": 0.36615182827586207,
332
  "brier_hard_n": 116,
333
+ "nll_hard": 1.9254142815070245,
334
  "nll_hard_n": 116,
335
+ "zero_probability_gold": 0.05172413793103448,
336
  "zero_probability_gold_n": 116,
337
+ "macro_f1": 0.7785414280259642
338
  },
339
  "massive-intent.en": {
340
  "attempted": 300,
341
  "valid": 300,
342
  "failed": 0,
343
  "coverage": 1.0,
344
+ "accuracy_all": 0.7566666666666667,
345
+ "accuracy_valid": 0.7566666666666667,
346
+ "ece_top_label": 0.19351909840865464,
347
+ "mean_confidence": 0.9501857650753213,
348
+ "brier_hard": 0.4290677996735415,
349
  "brier_hard_n": 300,
350
+ "nll_hard": 3.2153508464750185,
351
  "nll_hard_n": 300,
352
+ "zero_probability_gold": 0.08666666666666667,
353
  "zero_probability_gold_n": 300,
354
+ "macro_f1": 0.4943833232448793
355
  },
356
  "xnli.en": {
357
  "attempted": 300,
358
  "valid": 300,
359
  "failed": 0,
360
  "coverage": 1.0,
361
+ "accuracy_all": 0.87,
362
+ "accuracy_valid": 0.87,
363
+ "ece_top_label": 0.07153730747712309,
364
+ "mean_confidence": 0.9147377091848204,
365
+ "brier_hard": 0.22017864577608134,
366
  "brier_hard_n": 300,
367
+ "nll_hard": 0.4059236764304257,
368
  "nll_hard_n": 300,
369
  "zero_probability_gold": 0.0,
370
  "zero_probability_gold_n": 300,
371
+ "macro_f1": 0.8710695955086495
372
  }
373
  }
374
  },
375
  "2": {
376
  "seed": 2,
377
+ "selected_epoch": 13,
378
+ "elapsed_s": 699.7859417049913,
379
  "initial_validation": {
380
  "soft_cross_entropy": 1.544869564374288,
381
  "accuracy": 0.36833333333333335,
 
383
  "decisions": 600
384
  },
385
  "selected_validation": {
386
+ "soft_cross_entropy": 0.8543428432941437,
387
+ "accuracy": 0.7583333333333333,
388
+ "brier_soft": 0.07187043125430743,
389
  "decisions": 600
390
  },
391
  "stopping": {
392
  "reason": "early_stopping",
393
+ "epochs_completed": 16,
394
  "patience": 3,
395
  "min_epochs": 10,
396
  "epochs_without_improvement": 3,
 
404
  "valid": 2000,
405
  "failed": 0,
406
  "coverage": 1.0,
407
+ "accuracy_all": 0.745,
408
+ "accuracy_valid": 0.745,
409
+ "ece_top_label": 0.18594273050266366,
410
+ "mean_confidence": 0.5591527012331047,
411
+ "brier_hard": 0.40870440390133495,
412
  "brier_hard_n": 2000,
413
+ "nll_hard": 0.7205417712302097,
414
  "nll_hard_n": 2000,
415
  "zero_probability_gold": 0.0,
416
  "zero_probability_gold_n": 2000,
417
+ "soft_accuracy": 0.46875735649306255,
418
  "soft_accuracy_n": 2000,
419
+ "brier_soft": 0.06782387317668824,
420
  "brier_soft_n": 2000,
421
+ "kl_gold_to_prediction": 0.12687181845569698,
422
  "kl_gold_to_prediction_n": 2000,
423
+ "total_variation": 0.17935579359201664,
424
  "total_variation_n": 2000,
425
+ "score_mae": 0.25587944352517583,
426
  "score_mae_n": 800,
427
+ "within_one_level": 0.99,
428
  "within_one_level_n": 800,
429
+ "macro_f1": 0.608344198168488
430
  },
431
  "ag-news": {
432
  "attempted": 600,
433
  "valid": 600,
434
  "failed": 0,
435
  "coverage": 1.0,
436
+ "accuracy_all": 0.945,
437
+ "accuracy_valid": 0.945,
438
+ "ece_top_label": 0.03201179295141958,
439
+ "mean_confidence": 0.9180146184255912,
440
+ "brier_hard": 0.08699114920264774,
441
  "brier_hard_n": 600,
442
+ "nll_hard": 0.17458965637122142,
443
  "nll_hard_n": 600,
444
  "zero_probability_gold": 0.0,
445
  "zero_probability_gold_n": 600,
446
+ "macro_f1": 0.9421288699146882
447
  },
448
  "emotion": {
449
  "attempted": 600,
450
  "valid": 600,
451
  "failed": 0,
452
  "coverage": 1.0,
453
+ "accuracy_all": 0.56,
454
+ "accuracy_valid": 0.56,
455
+ "ece_top_label": 0.3501647005433899,
456
+ "mean_confidence": 0.9077857005433899,
457
+ "brier_hard": 0.7581475668473461,
458
  "brier_hard_n": 600,
459
+ "nll_hard": 2.427279492608694,
460
  "nll_hard_n": 600,
461
+ "zero_probability_gold": 0.013333333333333334,
462
  "zero_probability_gold_n": 600,
463
+ "macro_f1": 0.4667304743466239
464
  },
465
  "boolq": {
466
  "attempted": 600,
467
  "valid": 600,
468
  "failed": 0,
469
  "coverage": 1.0,
470
+ "accuracy_all": 0.8183333333333334,
471
+ "accuracy_valid": 0.8183333333333334,
472
+ "ece_top_label": 0.08817149999999999,
473
+ "mean_confidence": 0.9065048333333333,
474
+ "brier_hard": 0.2762812369666667,
475
  "brier_hard_n": 600,
476
+ "nll_hard": 0.453869636876017,
477
  "nll_hard_n": 600,
478
  "zero_probability_gold": 0.0,
479
  "zero_probability_gold_n": 600,
480
+ "macro_f1": 0.8046122269724754
481
  },
482
  "sst5": {
483
  "attempted": 600,
484
  "valid": 600,
485
  "failed": 0,
486
  "coverage": 1.0,
487
+ "accuracy_all": 0.3983333333333333,
488
+ "accuracy_valid": 0.3983333333333333,
489
+ "ece_top_label": 0.21926419101950614,
490
+ "mean_confidence": 0.6138282869290971,
491
+ "brier_hard": 0.7663539453668811,
492
  "brier_hard_n": 600,
493
+ "nll_hard": 1.5341982352412973,
494
  "nll_hard_n": 600,
495
  "zero_probability_gold": 0.0,
496
  "zero_probability_gold_n": 600,
497
+ "score_mae": 0.7954480085445932,
498
  "score_mae_n": 600,
499
+ "within_one_level": 0.695,
500
  "within_one_level_n": 600,
501
+ "macro_f1": 0.3618869575527743
502
  },
503
  "prompt-injections": {
504
  "attempted": 116,
505
  "valid": 116,
506
  "failed": 0,
507
  "coverage": 1.0,
508
+ "accuracy_all": 0.7155172413793104,
509
+ "accuracy_valid": 0.7155172413793104,
510
+ "ece_top_label": 0.2621448275862069,
511
+ "mean_confidence": 0.9488327586206896,
512
+ "brier_hard": 0.4988274203448276,
513
  "brier_hard_n": 116,
514
+ "nll_hard": 2.61023697496577,
515
  "nll_hard_n": 116,
516
+ "zero_probability_gold": 0.06896551724137931,
517
  "zero_probability_gold_n": 116,
518
+ "macro_f1": 0.6992221261884184
519
  },
520
  "massive-intent.en": {
521
  "attempted": 300,
522
  "valid": 300,
523
  "failed": 0,
524
  "coverage": 1.0,
525
+ "accuracy_all": 0.75,
526
+ "accuracy_valid": 0.75,
527
+ "ece_top_label": 0.20336797564979028,
528
+ "mean_confidence": 0.9533679756497903,
529
+ "brier_hard": 0.4358421604675745,
530
  "brier_hard_n": 300,
531
+ "nll_hard": 3.1322756175288426,
532
  "nll_hard_n": 300,
533
+ "zero_probability_gold": 0.08,
534
  "zero_probability_gold_n": 300,
535
+ "macro_f1": 0.5034639202079357
536
  },
537
  "xnli.en": {
538
  "attempted": 300,
539
  "valid": 300,
540
  "failed": 0,
541
  "coverage": 1.0,
542
+ "accuracy_all": 0.8733333333333333,
543
+ "accuracy_valid": 0.8733333333333333,
544
+ "ece_top_label": 0.05108680216338136,
545
+ "mean_confidence": 0.914837261771426,
546
+ "brier_hard": 0.20561652153986584,
547
  "brier_hard_n": 300,
548
+ "nll_hard": 0.3742216981370551,
549
  "nll_hard_n": 300,
550
  "zero_probability_gold": 0.0,
551
  "zero_probability_gold_n": 300,
552
+ "macro_f1": 0.8742255731346232
553
  }
554
  }
555
  },
556
  "3": {
557
  "seed": 3,
558
+ "selected_epoch": 8,
559
+ "elapsed_s": 495.011001034989,
560
  "initial_validation": {
561
  "soft_cross_entropy": 1.544869564374288,
562
  "accuracy": 0.36833333333333335,
 
564
  "decisions": 600
565
  },
566
  "selected_validation": {
567
+ "soft_cross_entropy": 0.8545633804798126,
568
+ "accuracy": 0.77,
569
+ "brier_soft": 0.06962224062221746,
570
  "decisions": 600
571
  },
572
  "stopping": {
573
  "reason": "early_stopping",
574
+ "epochs_completed": 11,
575
  "patience": 3,
576
  "min_epochs": 10,
577
  "epochs_without_improvement": 3,
 
585
  "valid": 2000,
586
  "failed": 0,
587
  "coverage": 1.0,
588
+ "accuracy_all": 0.7655,
589
+ "accuracy_valid": 0.7655,
590
+ "ece_top_label": 0.2096452515650509,
591
+ "mean_confidence": 0.5558547484349491,
592
+ "brier_hard": 0.4011842616732685,
593
  "brier_hard_n": 2000,
594
+ "nll_hard": 0.7113399953872087,
595
  "nll_hard_n": 2000,
596
  "zero_probability_gold": 0.0,
597
  "zero_probability_gold_n": 2000,
598
+ "soft_accuracy": 0.4683026193430771,
599
  "soft_accuracy_n": 2000,
600
+ "brier_soft": 0.06585425455571205,
601
  "brier_soft_n": 2000,
602
+ "kl_gold_to_prediction": 0.12507067261764912,
603
  "kl_gold_to_prediction_n": 2000,
604
+ "total_variation": 0.17848467876854548,
605
  "total_variation_n": 2000,
606
+ "score_mae": 0.24395106365305969,
607
  "score_mae_n": 800,
608
+ "within_one_level": 0.98375,
609
  "within_one_level_n": 800,
610
+ "macro_f1": 0.6399493527359326
611
  },
612
  "ag-news": {
613
  "attempted": 600,
614
  "valid": 600,
615
  "failed": 0,
616
  "coverage": 1.0,
617
+ "accuracy_all": 0.9366666666666666,
618
+ "accuracy_valid": 0.9366666666666666,
619
+ "ece_top_label": 0.0312034834128748,
620
+ "mean_confidence": 0.914535932360556,
621
+ "brier_hard": 0.09381807669545053,
622
  "brier_hard_n": 600,
623
+ "nll_hard": 0.1866731647377733,
624
  "nll_hard_n": 600,
625
  "zero_probability_gold": 0.0,
626
  "zero_probability_gold_n": 600,
627
+ "macro_f1": 0.9336084571946038
628
  },
629
  "emotion": {
630
  "attempted": 600,
631
  "valid": 600,
632
  "failed": 0,
633
  "coverage": 1.0,
634
+ "accuracy_all": 0.565,
635
+ "accuracy_valid": 0.565,
636
+ "ece_top_label": 0.3370553794856075,
637
+ "mean_confidence": 0.8987647229086076,
638
+ "brier_hard": 0.7489678137408704,
639
  "brier_hard_n": 600,
640
+ "nll_hard": 2.3318777367931305,
641
  "nll_hard_n": 600,
642
+ "zero_probability_gold": 0.011666666666666667,
643
  "zero_probability_gold_n": 600,
644
+ "macro_f1": 0.4760859469034663
645
  },
646
  "boolq": {
647
  "attempted": 600,
648
  "valid": 600,
649
  "failed": 0,
650
  "coverage": 1.0,
651
+ "accuracy_all": 0.82,
652
+ "accuracy_valid": 0.82,
653
+ "ece_top_label": 0.09564866666666666,
654
+ "mean_confidence": 0.9069513333333333,
655
+ "brier_hard": 0.2824305872,
656
  "brier_hard_n": 600,
657
+ "nll_hard": 0.4619444556248629,
658
  "nll_hard_n": 600,
659
  "zero_probability_gold": 0.0,
660
  "zero_probability_gold_n": 600,
661
+ "macro_f1": 0.8069498069498069
662
  },
663
  "sst5": {
664
  "attempted": 600,
665
  "valid": 600,
666
  "failed": 0,
667
  "coverage": 1.0,
668
+ "accuracy_all": 0.4116666666666667,
669
+ "accuracy_valid": 0.4116666666666667,
670
+ "ece_top_label": 0.19022023187751014,
671
+ "mean_confidence": 0.6018868985441768,
672
+ "brier_hard": 0.7571439809814903,
673
  "brier_hard_n": 600,
674
+ "nll_hard": 1.5683213482322946,
675
  "nll_hard_n": 600,
676
  "zero_probability_gold": 0.0,
677
  "zero_probability_gold_n": 600,
678
+ "score_mae": 0.8321864596740475,
679
  "score_mae_n": 600,
680
+ "within_one_level": 0.6783333333333333,
681
  "within_one_level_n": 600,
682
+ "macro_f1": 0.3816445513227996
683
  },
684
  "prompt-injections": {
685
  "attempted": 116,
686
  "valid": 116,
687
  "failed": 0,
688
  "coverage": 1.0,
689
+ "accuracy_all": 0.7844827586206896,
690
+ "accuracy_valid": 0.7844827586206896,
691
+ "ece_top_label": 0.17475344827586206,
692
+ "mean_confidence": 0.9523568965517242,
693
+ "brier_hard": 0.3675763103448276,
694
  "brier_hard_n": 116,
695
+ "nll_hard": 1.8418113090371786,
696
  "nll_hard_n": 116,
697
+ "zero_probability_gold": 0.04310344827586207,
698
  "zero_probability_gold_n": 116,
699
+ "macro_f1": 0.7797524113313588
700
  },
701
  "massive-intent.en": {
702
  "attempted": 300,
703
  "valid": 300,
704
  "failed": 0,
705
  "coverage": 1.0,
706
+ "accuracy_all": 0.7733333333333333,
707
+ "accuracy_valid": 0.7733333333333333,
708
+ "ece_top_label": 0.17472328670404882,
709
+ "mean_confidence": 0.9480566200373822,
710
+ "brier_hard": 0.38383190196466177,
711
  "brier_hard_n": 300,
712
+ "nll_hard": 2.4090375945333213,
713
  "nll_hard_n": 300,
714
+ "zero_probability_gold": 0.05,
715
  "zero_probability_gold_n": 300,
716
+ "macro_f1": 0.502438924453748
717
  },
718
  "xnli.en": {
719
  "attempted": 300,
720
  "valid": 300,
721
  "failed": 0,
722
  "coverage": 1.0,
723
+ "accuracy_all": 0.8666666666666667,
724
+ "accuracy_valid": 0.8666666666666667,
725
+ "ece_top_label": 0.054425126999604626,
726
+ "mean_confidence": 0.9150511806609052,
727
+ "brier_hard": 0.2024701804692983,
728
  "brier_hard_n": 300,
729
+ "nll_hard": 0.3723722500840691,
730
  "nll_hard_n": 300,
731
  "zero_probability_gold": 0.0,
732
  "zero_probability_gold_n": 300,
733
+ "macro_f1": 0.8674417053037536
734
  }
735
  }
736
  },
737
  "4": {
738
  "seed": 4,
739
+ "selected_epoch": 8,
740
+ "elapsed_s": 496.1691781419795,
741
  "initial_validation": {
742
  "soft_cross_entropy": 1.544869564374288,
743
  "accuracy": 0.36833333333333335,
 
745
  "decisions": 600
746
  },
747
  "selected_validation": {
748
+ "soft_cross_entropy": 0.8540202794472377,
749
+ "accuracy": 0.76,
750
+ "brier_soft": 0.07019127607345581,
751
  "decisions": 600
752
  },
753
  "stopping": {
754
  "reason": "early_stopping",
755
+ "epochs_completed": 11,
756
  "patience": 3,
757
  "min_epochs": 10,
758
  "epochs_without_improvement": 3,
 
766
  "valid": 2000,
767
  "failed": 0,
768
  "coverage": 1.0,
769
+ "accuracy_all": 0.7545,
770
+ "accuracy_valid": 0.7545,
771
+ "ece_top_label": 0.2042592508573846,
772
+ "mean_confidence": 0.5504920240151281,
773
+ "brier_hard": 0.4113014025244486,
774
  "brier_hard_n": 2000,
775
+ "nll_hard": 0.7247566141198507,
776
  "nll_hard_n": 2000,
777
  "zero_probability_gold": 0.0,
778
  "zero_probability_gold_n": 2000,
779
+ "soft_accuracy": 0.46590212671846004,
780
  "soft_accuracy_n": 2000,
781
+ "brier_soft": 0.06812301484891332,
782
  "brier_soft_n": 2000,
783
+ "kl_gold_to_prediction": 0.12790407031921358,
784
  "kl_gold_to_prediction_n": 2000,
785
+ "total_variation": 0.18054231568172208,
786
  "total_variation_n": 2000,
787
+ "score_mae": 0.25377940453843617,
788
  "score_mae_n": 800,
789
+ "within_one_level": 0.98375,
790
  "within_one_level_n": 800,
791
+ "macro_f1": 0.6337422972118567
792
  },
793
  "ag-news": {
794
  "attempted": 600,
795
  "valid": 600,
796
  "failed": 0,
797
  "coverage": 1.0,
798
+ "accuracy_all": 0.9516666666666667,
799
+ "accuracy_valid": 0.9516666666666667,
800
+ "ece_top_label": 0.03233758800344756,
801
+ "mean_confidence": 0.9199325508493339,
802
+ "brier_hard": 0.08178925980745573,
803
  "brier_hard_n": 600,
804
+ "nll_hard": 0.16974235809511878,
805
  "nll_hard_n": 600,
806
  "zero_probability_gold": 0.0,
807
  "zero_probability_gold_n": 600,
808
+ "macro_f1": 0.9489828512364638
809
  },
810
  "emotion": {
811
  "attempted": 600,
812
  "valid": 600,
813
  "failed": 0,
814
  "coverage": 1.0,
815
+ "accuracy_all": 0.5783333333333334,
816
+ "accuracy_valid": 0.5783333333333334,
817
+ "ece_top_label": 0.3295382068772924,
818
+ "mean_confidence": 0.9035819218093762,
819
+ "brier_hard": 0.7425345488860924,
820
  "brier_hard_n": 600,
821
+ "nll_hard": 2.3910245621280453,
822
  "nll_hard_n": 600,
823
+ "zero_probability_gold": 0.013333333333333334,
824
  "zero_probability_gold_n": 600,
825
+ "macro_f1": 0.48527808236443093
826
  },
827
  "boolq": {
828
  "attempted": 600,
829
  "valid": 600,
830
  "failed": 0,
831
  "coverage": 1.0,
832
+ "accuracy_all": 0.8316666666666667,
833
+ "accuracy_valid": 0.8316666666666667,
834
+ "ece_top_label": 0.09521133333333336,
835
+ "mean_confidence": 0.9097233333333333,
836
+ "brier_hard": 0.27490432886666666,
837
  "brier_hard_n": 600,
838
+ "nll_hard": 0.49707214589381826,
839
  "nll_hard_n": 600,
840
+ "zero_probability_gold": 0.0016666666666666668,
841
  "zero_probability_gold_n": 600,
842
+ "macro_f1": 0.8196294367140413
843
  },
844
  "sst5": {
845
  "attempted": 600,
846
  "valid": 600,
847
  "failed": 0,
848
  "coverage": 1.0,
849
+ "accuracy_all": 0.41,
850
+ "accuracy_valid": 0.41,
851
+ "ece_top_label": 0.20145470242332547,
852
+ "mean_confidence": 0.6098530357566588,
853
+ "brier_hard": 0.7512849781669847,
854
  "brier_hard_n": 600,
855
+ "nll_hard": 1.5195838995022448,
856
  "nll_hard_n": 600,
857
  "zero_probability_gold": 0.0,
858
  "zero_probability_gold_n": 600,
859
+ "score_mae": 0.8012926973404637,
860
  "score_mae_n": 600,
861
+ "within_one_level": 0.7016666666666667,
862
  "within_one_level_n": 600,
863
+ "macro_f1": 0.37608916605464326
864
  },
865
  "prompt-injections": {
866
  "attempted": 116,
867
  "valid": 116,
868
  "failed": 0,
869
  "coverage": 1.0,
870
+ "accuracy_all": 0.7413793103448276,
871
+ "accuracy_valid": 0.7413793103448276,
872
+ "ece_top_label": 0.21426120689655168,
873
+ "mean_confidence": 0.9556405172413793,
874
+ "brier_hard": 0.43549763362068966,
875
  "brier_hard_n": 116,
876
+ "nll_hard": 2.2934410633534186,
877
  "nll_hard_n": 116,
878
+ "zero_probability_gold": 0.0603448275862069,
879
  "zero_probability_gold_n": 116,
880
+ "macro_f1": 0.7276995305164319
881
  },
882
  "massive-intent.en": {
883
  "attempted": 300,
884
  "valid": 300,
885
  "failed": 0,
886
  "coverage": 1.0,
887
+ "accuracy_all": 0.77,
888
+ "accuracy_valid": 0.77,
889
+ "ece_top_label": 0.17971198952007902,
890
+ "mean_confidence": 0.9447345124535086,
891
+ "brier_hard": 0.4058348152163361,
892
  "brier_hard_n": 300,
893
+ "nll_hard": 2.9106532352305194,
894
  "nll_hard_n": 300,
895
+ "zero_probability_gold": 0.07333333333333333,
896
  "zero_probability_gold_n": 300,
897
+ "macro_f1": 0.5030980723143226
898
  },
899
  "xnli.en": {
900
  "attempted": 300,
901
  "valid": 300,
902
  "failed": 0,
903
  "coverage": 1.0,
904
+ "accuracy_all": 0.8633333333333333,
905
+ "accuracy_valid": 0.8633333333333333,
906
+ "ece_top_label": 0.0751956003124493,
907
+ "mean_confidence": 0.9162548569668819,
908
+ "brier_hard": 0.22384958324541668,
909
  "brier_hard_n": 300,
910
+ "nll_hard": 0.4127381776210696,
911
  "nll_hard_n": 300,
912
  "zero_probability_gold": 0.0,
913
  "zero_probability_gold_n": 300,
914
+ "macro_f1": 0.864378848443158
915
  }
916
  }
917
  }
 
919
  "summary": {
920
  "typed-decisions": {
921
  "n": 5,
922
+ "mean_percent": 75.28,
923
+ "sample_variance_pp_squared": 0.8619999999999957,
924
+ "sample_std_pp": 0.9284395510748105
925
  },
926
  "ag-news": {
927
  "n": 5,
928
+ "mean_percent": 94.46666666666667,
929
+ "sample_variance_pp_squared": 0.3388888888888897,
930
+ "sample_std_pp": 0.5821416398857667
931
  },
932
  "boolq": {
933
  "n": 5,
934
+ "mean_percent": 82.10000000000001,
935
+ "sample_variance_pp_squared": 0.7861111111111168,
936
+ "sample_std_pp": 0.8866290718846956
937
  },
938
  "emotion": {
939
  "n": 5,
940
+ "mean_percent": 57.2,
941
+ "sample_variance_pp_squared": 1.0055555555555546,
942
+ "sample_std_pp": 1.0027739304327543
943
  },
944
  "prompt-injections": {
945
  "n": 5,
946
+ "mean_percent": 74.48275862068965,
947
+ "sample_variance_pp_squared": 15.457788347205712,
948
+ "sample_std_pp": 3.93163939689358
949
  },
950
  "sst5": {
951
  "n": 5,
952
+ "mean_percent": 41.0,
953
+ "sample_variance_pp_squared": 6.027777777777775,
954
+ "sample_std_pp": 2.4551533104427055
955
  },
956
  "massive-intent.en": {
957
  "n": 5,
958
+ "mean_percent": 76.8,
959
+ "sample_variance_pp_squared": 2.422222222222218,
960
+ "sample_std_pp": 1.556349003990499
961
  },
962
  "xnli.en": {
963
  "n": 5,
964
+ "mean_percent": 87.26666666666667,
965
+ "sample_variance_pp_squared": 1.0777777777777784,
966
+ "sample_std_pp": 1.038160766826496
967
  }
968
  },
969
+ "candidate_seed": 4,
970
  "variance_definition": "sample variance, n-1, percentage points squared",
971
  "benchmark_manifest": {
972
  "format_version": 1,
release_manifest.json ADDED
@@ -0,0 +1,54 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "published": false,
3
+ "review_status": "export verified; package prepared for user review and upload",
4
+ "files_sha256": {
5
+ "LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
6
+ "README.md": "db6038098fae3817f20fdabc66afb13d7e7cd1306d0b611eea7f70ef84ad08bc",
7
+ "THIRD_PARTY.md": "3dd2683344d45ce64c15f45c723f4b0db720deee84446365fdf4084a73ebeceb",
8
+ "USAGE.md": "038e1d21dca7d2e2da5b086a34e5b446edb8901406609c8c8b0e3407f31ab7d9",
9
+ "adapter.safetensors": "f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081",
10
+ "adapter_config.json": "b26f0cb6ea6c6f1e66f254b832d5000b245d309fc05880f234d0f7489cadb54a",
11
+ "ariadne_bench/__init__.py": "a73c9fdf21413382cb93b523bfb3925aa3fd354e18faeead4650e86a30645411",
12
+ "ariadne_bench/__main__.py": "935a1c1166b0c1ea35a82256345000bf2c73ded718d77773bc27a71ecce28f7d",
13
+ "ariadne_bench/adapters.py": "3bd4820165be077b3b98fe83ac23e649e1ad69186ad5f2a4f5ee7addf1427cf5",
14
+ "ariadne_bench/cli.py": "de267c0502489eb4bb37e4b8f6faa8da8be37a56ced270f754424a40404d1af6",
15
+ "ariadne_bench/datasets.py": "31c722c64895519a4635b501da88862b446384682e3e8ee884495941d6990912",
16
+ "ariadne_bench/experiments/__init__.py": "95ad0a7f30c4e0a194d0a6e2ae8905abae5d622edbb3d4e5d791b4bc460aee5d",
17
+ "ariadne_bench/experiments/align.py": "31e2733dcc223a1e1a6250d41c4158f84091f1cb4067d61388e06b5b01ae34c4",
18
+ "ariadne_bench/experiments/prepare_alignment.py": "c4f39c10ce46ee13c0d89128d18a85ec4a86729e81e070f5834ed01fc1c405a9",
19
+ "ariadne_bench/experiments/spectrum.py": "22d4801b585c539d0d579a9d2a00076e1002fd337bfe31d1906d8ae3fb8d4320",
20
+ "ariadne_bench/frozen_input_interface.py": "253f0186dac6e4a7f4b0fcf3c33449c3b3882cf8be538dca6eeca8d9e4ce9b83",
21
+ "ariadne_bench/full_finetune.py": "6067be8cf837d254b919c841c9a9851e1546a413dbf6324ff242e5d2f56c9734",
22
+ "ariadne_bench/full_input_interface.py": "a8f440a2185085c64896300a243ac52a9619c4d05a1b4e43c50c2c32aa1da996",
23
+ "ariadne_bench/hybrid.py": "90fbbdab9b8157a32854d27c2d93089be6dc88dfe9320426f102eae4e03d9ffd",
24
+ "ariadne_bench/hybrid_control.py": "1b53678bea73fe7608570b78d7215d1dff2e38d6ad3627615f26c06ae74d8b4c",
25
+ "ariadne_bench/interfaces.py": "19700fb8170b2460401900cb96ebdff22bd7bf778ca5132e98d56ba68f344187",
26
+ "ariadne_bench/metrics.py": "fc6fde20c95d5052de52d05d6fb9b6eb69b780c46966f5bf1e517738b276a4b4",
27
+ "ariadne_bench/portable.py": "259f633a8bb6663d0594d9639ea8d80ef493f5a63d4b11444a4003f2f9a318e8",
28
+ "ariadne_bench/reproducibility.py": "5b0c161c596f25277b3326c645df53d5d1f2dc0272da9e77bc4a4c0d56188b6e",
29
+ "ariadne_bench/runner.py": "b20e2af4ba254e2d48a07563356f95728c4807018c3c989b429e9922674c64cd",
30
+ "ariadne_bench/schema.py": "67794f1172c1ef9d9625b84c8096fd48b8fa7cd10e3f88fc13ba3a98ef7c7a76",
31
+ "benchmarks/sources.lock.json": "6e6eaab2778092c8127257fce202c9fa8d4625dd0d029490dd9f137da91d134c",
32
+ "environment.json": "c5b9c796d97f3dec6d242caf79aa5d2ac69c91717f2be6ed5a51627bf6d58c17",
33
+ "evidence/adapter-equivalence.json": "7ffdc6216702ab5de02a1745217b369de9f2f8934e16e5c80dd090a3343e9e84",
34
+ "evidence/audit.json": "7a2954a470b3f94626ca678aef7f9bf1f13551be20e642c1a16a0a90b751bd4d",
35
+ "evidence/candidate-uncertainty.json": "9e97410249531de151864285412ab3bd337b117b3573f6439cf74ae17926cae3",
36
+ "evidence/identity-check.json": "0cc8904dc58e1e0ed49496c7ae01980c52e3a1395d61ef91286cec2c9dcdfa25",
37
+ "evidence/lr-comparison.json": "724a08155ff3f39a335132cffb0248fdb08f06874e84ce6a582c8af9199ac69d",
38
+ "evidence/replay-check.json": "d362a58603077feddc4ad2513c19f8caea25da1a9ab03eec25c5be5c14c3eec0",
39
+ "evidence/seed-0-history.json": "1090a62e8f1256e9510f8c2c913de14954a5e54e14396a962578ebef83b4a6f3",
40
+ "evidence/seed-1-history.json": "6aee05fe087ec5d970329743a06c1cbff3e2a0e8209cdcb2db41649b66b75432",
41
+ "evidence/seed-2-history.json": "71afe93da2e39ef81417074bc1a4559b9d6f070f677e76f74818129633cf3c33",
42
+ "evidence/seed-3-history.json": "4bf723c1659ffc8adeb57d06d51132f4f5af74c811c98feef71de4f6fa3bd66f",
43
+ "evidence/seed-4-history.json": "2609c807813f55ffe5aef7e6c96aafefe1647de1ac589568d7fdc56e46f5c9fe",
44
+ "export-verification.json": "65e46623219d80979a7108e69c0ed89d0b1d5c23a82c4255fee64647694aac3f",
45
+ "licenses/jev-benchmarks-LICENSE": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
46
+ "licenses/laya-LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
47
+ "load_adapter.py": "2f305ef95fef2b2482e9214836fc0071a59cfd6c6d3aaf42cebdbfb131272081",
48
+ "make_training_config.py": "23e9d4acb90141a3d690bc607c861df92137b9457304ae19b8a925398b55536d",
49
+ "metrics.json": "e3cd35993afc309f7ccd99b1a3d96316944e1ab42fd489c7f15465f43ec4a78b",
50
+ "requirements.txt": "285a8ae6529abeefd058f2e9679e0e314b19fda38169e685ea16aec59b053bd9",
51
+ "training_protocol.json": "00b2db7abbd07ba3220a130fe36ed64958f1dfac15773d183f1ae8dc9d8334b3",
52
+ "upload-manifest.json": "613b92d535a5c8fca8dfe1d6535600f6da68ba0bf110b53829d13638614830b6"
53
+ }
54
+ }
training_protocol.json CHANGED
@@ -1,5 +1,5 @@
1
  {
2
- "experiment": "E0i",
3
  "seeds": [
4
  0,
5
  1,
@@ -7,7 +7,7 @@
7
  3,
8
  4
9
  ],
10
- "created_at": "2026-09-27T13:49:50.006241+00:00",
11
  "data": {
12
  "train": {
13
  "format_version": 1,
@@ -67,7 +67,7 @@
67
  "microbatch": 8,
68
  "accumulation": 4,
69
  "effective_batch": 32,
70
- "interface_lr": 0.0003,
71
  "weight_decay": 0.01,
72
  "gradient_clip": 1,
73
  "training_scope": "bridge",
@@ -77,8 +77,23 @@
77
  "device": "cuda:1",
78
  "deterministic": true,
79
  "native_sdk_calibration": true,
80
- "publish_policy": "draft locally; upload only after user and assistant agree",
81
  "baseline_predictions_sha256": "37dff666a00a340cd6e37f775dc60df33a3c6eb78f76bfee7f094443c225b6e8",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
82
  "base_model": "convaiinnovations/laya",
83
  "base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
84
  "training_sources_sha256": {
 
1
  {
2
+ "experiment": "E0j five-seed confirmation",
3
  "seeds": [
4
  0,
5
  1,
 
7
  3,
8
  4
9
  ],
10
+ "created_at": "2026-09-27T15:29:06.749690+00:00",
11
  "data": {
12
  "train": {
13
  "format_version": 1,
 
67
  "microbatch": 8,
68
  "accumulation": 4,
69
  "effective_batch": 32,
70
+ "interface_lr": 0.0001,
71
  "weight_decay": 0.01,
72
  "gradient_clip": 1,
73
  "training_scope": "bridge",
 
77
  "device": "cuda:1",
78
  "deterministic": true,
79
  "native_sdk_calibration": true,
80
+ "publish_policy": "provide user-requested local bundle; preserve existing card; no assistant upload",
81
  "baseline_predictions_sha256": "37dff666a00a340cd6e37f775dc60df33a3c6eb78f76bfee7f094443c225b6e8",
82
+ "run_order": [
83
+ 0,
84
+ 2,
85
+ 3,
86
+ 1,
87
+ 4
88
+ ],
89
+ "seed_0_reference": "completed controlled pilot, unchanged",
90
+ "comparison": "LR is the only training change from E0i; min10/patience3 and all other settings fixed",
91
+ "scope": "adaptive experimental follow-up selected after the seed-0 pilot; not a fresh blind test",
92
+ "replay_evidence": "Two fresh seed-0 LR 1e-4 processes matched losses, metrics and tensors for two epochs; reported seed-0 metric prefix matched. Verified after the sweep.",
93
+ "driver_sha256": {
94
+ "/mnt/storage/ariadne/scripts/run_frozen_linear_lr_confirmation.py": "f7c3841277682cbae43f4eb51fa3cb5fc87fd53068607206240ce137b7e94a9b"
95
+ },
96
+ "release_replay_sha256": "dd3479114e8af11d06708febc1e86b887cc6e3cc5ce1523a249e54b9432412ec",
97
  "base_model": "convaiinnovations/laya",
98
  "base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
99
  "training_sources_sha256": {
upload-manifest.json CHANGED
@@ -1,17 +1,17 @@
1
  {
2
  "published_by_assistant": false,
3
- "candidate_seed": 1,
4
- "selected_epoch": 14,
5
  "format": "compact affine adapter requiring pinned base Laya",
6
- "existing_model_card_preserved": true,
7
- "inference_code_and_weights_match_verified_export": true,
8
- "removed_only_other_seed_loading_entries": true,
9
  "files_sha256": {
10
  "LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
 
11
  "THIRD_PARTY.md": "3dd2683344d45ce64c15f45c723f4b0db720deee84446365fdf4084a73ebeceb",
12
- "USAGE.md": "e6d76a3a450fd589e52cab5190a065cbd783082a98d3524071bcf60a7015ebef",
13
- "adapter.safetensors": "e5ec5e1d494509f63960b6949af15be8390442271805ba45d796a1911b5c239a",
14
- "adapter_config.json": "8d7450b812f696e1dbeb3d7abfa13fbee0cde95aae0349a94e5f22c9b458ed1f",
15
  "ariadne_bench/__init__.py": "a73c9fdf21413382cb93b523bfb3925aa3fd354e18faeead4650e86a30645411",
16
  "ariadne_bench/__main__.py": "935a1c1166b0c1ea35a82256345000bf2c73ded718d77773bc27a71ecce28f7d",
17
  "ariadne_bench/adapters.py": "3bd4820165be077b3b98fe83ac23e649e1ad69186ad5f2a4f5ee7addf1427cf5",
@@ -34,22 +34,24 @@
34
  "ariadne_bench/schema.py": "67794f1172c1ef9d9625b84c8096fd48b8fa7cd10e3f88fc13ba3a98ef7c7a76",
35
  "benchmarks/sources.lock.json": "6e6eaab2778092c8127257fce202c9fa8d4625dd0d029490dd9f137da91d134c",
36
  "environment.json": "c5b9c796d97f3dec6d242caf79aa5d2ac69c91717f2be6ed5a51627bf6d58c17",
37
- "evidence/adapter-equivalence.json": "c77faa691bd273d363dbb196155545b4f72697b9d0f4c80451ffb06b0090d200",
38
- "evidence/audit.json": "1fb474a47923e22374a47e06f545e747420c0a7bd269bbbeaabb5869e89a8d70",
39
- "evidence/candidate-uncertainty.json": "39eba4a205eaea5f13384da24ca8563de364c9edb6858c1db38eb4111960ce40",
40
  "evidence/identity-check.json": "0cc8904dc58e1e0ed49496c7ae01980c52e3a1395d61ef91286cec2c9dcdfa25",
41
- "evidence/replay-check.json": "4f246eb3d6afc0c0b1d306a20b82060ae564d476ec22451322b065edb946d15e",
42
- "evidence/seed-0-history.json": "bc3de468970bf350426d9c4281456e258984838feb4d8dbbfe54c8752ebb24b7",
43
- "evidence/seed-1-history.json": "8ee78b5685138712512b39035ef48a491ed52813c2131a8e9b80715871ab9405",
44
- "evidence/seed-2-history.json": "d9bf7ab3991b85b6e3877b9ebd4ca2fc113cba1716f56da5aa16f4018ac4b197",
45
- "evidence/seed-3-history.json": "7c0f19638e31f28e438af879ceaa68bbe06480e6f7fb72e8e749b4aabb304341",
46
- "evidence/seed-4-history.json": "601c5bd91dcbb004ff8a07621f3f125ba578cc4107a9884e8fe69ac17751a741",
47
- "export-verification.json": "b9f8667626e2facd03f310c0c67c100ddde9481218c51cbead7972fd048d4afb",
 
48
  "licenses/jev-benchmarks-LICENSE": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
49
  "licenses/laya-LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
50
  "load_adapter.py": "2f305ef95fef2b2482e9214836fc0071a59cfd6c6d3aaf42cebdbfb131272081",
51
- "metrics.json": "dd45292f04dc5b81cc958a39f618aa8748b5198e39b5effee6c95520c92d94b6",
 
52
  "requirements.txt": "285a8ae6529abeefd058f2e9679e0e314b19fda38169e685ea16aec59b053bd9",
53
- "training_protocol.json": "dccf81e3d8a70c756a95a88280cbec139eea43bc391d67fa7b40e384a8824f57"
54
  }
55
  }
 
1
  {
2
  "published_by_assistant": false,
3
+ "candidate_seed": 4,
4
+ "selected_epoch": 8,
5
  "format": "compact affine adapter requiring pinned base Laya",
6
+ "includes_model_card": true,
7
+ "all_5116_export_answers_exact": true,
 
8
  "files_sha256": {
9
  "LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
10
+ "README.md": "db6038098fae3817f20fdabc66afb13d7e7cd1306d0b611eea7f70ef84ad08bc",
11
  "THIRD_PARTY.md": "3dd2683344d45ce64c15f45c723f4b0db720deee84446365fdf4084a73ebeceb",
12
+ "USAGE.md": "038e1d21dca7d2e2da5b086a34e5b446edb8901406609c8c8b0e3407f31ab7d9",
13
+ "adapter.safetensors": "f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081",
14
+ "adapter_config.json": "b26f0cb6ea6c6f1e66f254b832d5000b245d309fc05880f234d0f7489cadb54a",
15
  "ariadne_bench/__init__.py": "a73c9fdf21413382cb93b523bfb3925aa3fd354e18faeead4650e86a30645411",
16
  "ariadne_bench/__main__.py": "935a1c1166b0c1ea35a82256345000bf2c73ded718d77773bc27a71ecce28f7d",
17
  "ariadne_bench/adapters.py": "3bd4820165be077b3b98fe83ac23e649e1ad69186ad5f2a4f5ee7addf1427cf5",
 
34
  "ariadne_bench/schema.py": "67794f1172c1ef9d9625b84c8096fd48b8fa7cd10e3f88fc13ba3a98ef7c7a76",
35
  "benchmarks/sources.lock.json": "6e6eaab2778092c8127257fce202c9fa8d4625dd0d029490dd9f137da91d134c",
36
  "environment.json": "c5b9c796d97f3dec6d242caf79aa5d2ac69c91717f2be6ed5a51627bf6d58c17",
37
+ "evidence/adapter-equivalence.json": "7ffdc6216702ab5de02a1745217b369de9f2f8934e16e5c80dd090a3343e9e84",
38
+ "evidence/audit.json": "7a2954a470b3f94626ca678aef7f9bf1f13551be20e642c1a16a0a90b751bd4d",
39
+ "evidence/candidate-uncertainty.json": "9e97410249531de151864285412ab3bd337b117b3573f6439cf74ae17926cae3",
40
  "evidence/identity-check.json": "0cc8904dc58e1e0ed49496c7ae01980c52e3a1395d61ef91286cec2c9dcdfa25",
41
+ "evidence/lr-comparison.json": "724a08155ff3f39a335132cffb0248fdb08f06874e84ce6a582c8af9199ac69d",
42
+ "evidence/replay-check.json": "d362a58603077feddc4ad2513c19f8caea25da1a9ab03eec25c5be5c14c3eec0",
43
+ "evidence/seed-0-history.json": "1090a62e8f1256e9510f8c2c913de14954a5e54e14396a962578ebef83b4a6f3",
44
+ "evidence/seed-1-history.json": "6aee05fe087ec5d970329743a06c1cbff3e2a0e8209cdcb2db41649b66b75432",
45
+ "evidence/seed-2-history.json": "71afe93da2e39ef81417074bc1a4559b9d6f070f677e76f74818129633cf3c33",
46
+ "evidence/seed-3-history.json": "4bf723c1659ffc8adeb57d06d51132f4f5af74c811c98feef71de4f6fa3bd66f",
47
+ "evidence/seed-4-history.json": "2609c807813f55ffe5aef7e6c96aafefe1647de1ac589568d7fdc56e46f5c9fe",
48
+ "export-verification.json": "65e46623219d80979a7108e69c0ed89d0b1d5c23a82c4255fee64647694aac3f",
49
  "licenses/jev-benchmarks-LICENSE": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
50
  "licenses/laya-LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
51
  "load_adapter.py": "2f305ef95fef2b2482e9214836fc0071a59cfd6c6d3aaf42cebdbfb131272081",
52
+ "make_training_config.py": "23e9d4acb90141a3d690bc607c861df92137b9457304ae19b8a925398b55536d",
53
+ "metrics.json": "e3cd35993afc309f7ccd99b1a3d96316944e1ab42fd489c7f15465f43ec4a78b",
54
  "requirements.txt": "285a8ae6529abeefd058f2e9679e0e314b19fda38169e685ea16aec59b053bd9",
55
+ "training_protocol.json": "00b2db7abbd07ba3220a130fe36ed64958f1dfac15773d183f1ae8dc9d8334b3"
56
  }
57
  }