Upload 49 files
Browse files- README.md +121 -261
- USAGE.md +2 -47
- adapter.safetensors +1 -1
- adapter_config.json +4 -4
- evidence/adapter-equivalence.json +6 -1
- evidence/audit.json +10 -10
- evidence/candidate-uncertainty.json +25 -25
- evidence/lr-comparison.json +315 -0
- evidence/replay-check.json +6 -3
- evidence/seed-0-history.json +79 -232
- evidence/seed-1-history.json +67 -152
- evidence/seed-2-history.json +93 -127
- evidence/seed-3-history.json +61 -333
- evidence/seed-4-history.json +75 -58
- export-verification.json +5 -5
- make_training_config.py +37 -0
- metrics.json +391 -391
- release_manifest.json +54 -0
- training_protocol.json +19 -4
- upload-manifest.json +22 -20
README.md
CHANGED
|
@@ -2,327 +2,187 @@
|
|
| 2 |
license: apache-2.0
|
| 3 |
language:
|
| 4 |
- en
|
| 5 |
-
base_model:
|
| 6 |
-
- convaiinnovations/laya
|
| 7 |
-
base_model_relation: adapter
|
| 8 |
pipeline_tag: text-classification
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
tags:
|
| 10 |
- laya
|
| 11 |
-
-
|
| 12 |
-
- system-one
|
| 13 |
-
- classification
|
| 14 |
- adapter
|
| 15 |
-
-
|
| 16 |
-
-
|
| 17 |
-
datasets:
|
| 18 |
-
- LocalLLaMA/typed-decisions
|
| 19 |
-
metrics:
|
| 20 |
-
- accuracy
|
| 21 |
-
- brier_score
|
| 22 |
-
library_name: transformers
|
| 23 |
---
|
| 24 |
-
# Ariadne-Laya-TD
|
| 25 |
|
| 26 |
-
|
| 27 |
|
| 28 |
-
|
| 29 |
|
| 30 |
-
|
| 31 |
-
h' = Wh + b,\qquad W_0 = I,\qquad b_0 = 0
|
| 32 |
-
$$
|
| 33 |
|
| 34 |
-
The
|
| 35 |
|
| 36 |
-
The
|
| 37 |
|
| 38 |
-
|
| 39 |
|
| 40 |
-
|
| 41 |
-
Ariadne learns a small transformation that allows the existing model to use that structure without modifying the encoder itself.
|
| 42 |
-
|
| 43 |
-
This makes specialisms extremely lightweight: a deployment can keep one frozen Laya base resident and switch between task-specific Ariadne interfaces containing only 1.05M parameters—roughly 2 MB in BF16—rather than loading a separate ~421M-parameter model for each specialism.
|
| 44 |
|
| 45 |
## Architecture
|
| 46 |
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
```text
|
| 50 |
-
input
|
| 51 |
-
↓
|
| 52 |
-
Laya embeddings
|
| 53 |
-
↓
|
| 54 |
-
Ariadne linear interface
|
| 55 |
-
↓
|
| 56 |
-
frozen Laya encoder
|
| 57 |
-
↓
|
| 58 |
-
existing decision pathway
|
| 59 |
-
```
|
| 60 |
-
|
| 61 |
-
The interface is deliberately minimal:
|
| 62 |
-
|
| 63 |
-
- one 1024 × 1024 linear transformation
|
| 64 |
-
- one 1024-dimensional bias
|
| 65 |
-
- exact identity initialization
|
| 66 |
-
- **1,049,600 trainable parameters**
|
| 67 |
-
- no LoRA matrices
|
| 68 |
-
- no additional transformer blocks
|
| 69 |
-
- no modification of the Laya encoder during interface training
|
| 70 |
-
|
| 71 |
-
At initialization, Ariadne is exactly equivalent to the unmodified base model.
|
| 72 |
-
|
| 73 |
-
Training therefore learns only how to transform the representation presented to the existing frozen network.
|
| 74 |
-
|
| 75 |
-
## Base model
|
| 76 |
-
|
| 77 |
-
Ariadne-Laya-TD is an adapter for:
|
| 78 |
-
|
| 79 |
-
`convaiinnovations/laya`
|
| 80 |
-
|
| 81 |
-
The underlying Laya base model remains frozen during Ariadne training.
|
| 82 |
-
|
| 83 |
-
The principal comparison model is the official Laya typed-decisions specialist:
|
| 84 |
|
| 85 |
-
|
| 86 |
|
| 87 |
-
|
| 88 |
|
| 89 |
-
##
|
| 90 |
-
|
| 91 |
-
The primary experiment trains Ariadne on the `LocalLLaMA/typed-decisions` training data while keeping the entire Laya base model frozen.
|
| 92 |
-
|
| 93 |
-
Only the 1.05M-parameter interface is updated.
|
| 94 |
|
| 95 |
-
|
| 96 |
|
| 97 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 98 |
|
| 99 |
-
|
| 100 |
|
| 101 |
-
|
| 102 |
|
| 103 |
-
|
|
| 104 |
|---|---:|---:|
|
| 105 |
-
|
|
| 106 |
-
|
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
The goal is not simply to reduce the number of trainable parameters.
|
| 115 |
-
|
| 116 |
-
The frozen-base experiment tests whether task-specific discrimination requires rewriting the internal representation of the model at all.
|
| 117 |
-
|
| 118 |
-
The working hypothesis is:
|
| 119 |
|
| 120 |
-
|
| 121 |
|
| 122 |
-
|
| 123 |
|
| 124 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 125 |
|
| 126 |
-
|
| 127 |
|
| 128 |
-
|
| 129 |
|
| 130 |
-
|
| 131 |
-
|
| 132 |
-
In this experiment:
|
| 133 |
-
|
| 134 |
-
- the specialist itself is frozen
|
| 135 |
-
- only the Ariadne interface is trained
|
| 136 |
-
- five seeds are evaluated under the same benchmark bundle
|
| 137 |
-
|
| 138 |
-
Results:
|
| 139 |
-
|
| 140 |
-
| Method | Trainable parameters | Typed-decisions accuracy |
|
| 141 |
|---|---:|---:|
|
| 142 |
-
|
|
| 143 |
-
|
|
| 144 |
-
|
|
| 145 |
-
|
|
| 146 |
-
|
|
| 147 |
-
|
| 148 |
-
|
|
|
|
| 149 |
|
| 150 |
-
|
| 151 |
|
| 152 |
-
|
| 153 |
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
## Cross-task retention after specialist alignment
|
| 157 |
-
|
| 158 |
-
The post-specialization Ariadne interface was also evaluated on seven additional English benchmarks.
|
| 159 |
-
|
| 160 |
-
| Benchmark | Untouched specialist | Specialist + Ariadne |
|
| 161 |
-
|---|---:|---:|
|
| 162 |
-
| Typed decisions | 76.70% | **78.10 ± 0.31%** |
|
| 163 |
-
| AG News | 94.67% | 94.17 ± 0.37% |
|
| 164 |
-
| BoolQ | 82.00% | 81.10 ± 0.90% |
|
| 165 |
-
| DAIR Emotion | 58.17% | 57.40 ± 0.65% |
|
| 166 |
-
| Prompt injections | 67.24% | 67.41 ± 1.87% |
|
| 167 |
-
| SST-5 | 48.00% | 49.33 ± 1.67% |
|
| 168 |
-
| MASSIVE intent EN | 81.00% | 79.40 ± 1.23% |
|
| 169 |
-
| XNLI EN | 86.33% | 88.33 ± 0.53% |
|
| 170 |
|
| 171 |
-
|
| 172 |
-
|
| 173 |
-
|
| 174 |
-
|
| 175 |
-
|
|
|
|
| 176 |
|
| 177 |
## Training
|
| 178 |
|
| 179 |
-
|
| 180 |
-
|
| 181 |
-
The released Ariadne checkpoint is trained from the official Laya base model.
|
| 182 |
-
|
| 183 |
-
During training:
|
| 184 |
-
|
| 185 |
-
- all Laya parameters remain frozen
|
| 186 |
-
- the frozen encoder is run in evaluation mode
|
| 187 |
-
- only the Ariadne interface is trainable
|
| 188 |
-
- initialization is exact identity
|
| 189 |
-
- training uses soft typed-decision targets
|
| 190 |
-
- checkpoint selection uses validation soft cross-entropy
|
| 191 |
-
|
| 192 |
-
Current training experiments use deterministic execution settings to better separate seed effects from GPU numerical nondeterminism.
|
| 193 |
-
|
| 194 |
-
The final training configuration and five-seed statistics will be added once the current deterministic sweep is complete.
|
| 195 |
-
|
| 196 |
-
### Post-specialist Ariadne experiment
|
| 197 |
-
|
| 198 |
-
The completed post-specialist sweep used:
|
| 199 |
|
| 200 |
-
-
|
| 201 |
-
-
|
| 202 |
-
-
|
| 203 |
-
-
|
| 204 |
-
-
|
| 205 |
-
-
|
| 206 |
-
- objective: soft-target decision cross-entropy
|
| 207 |
-
- seeds: 0, 1, 2, 3, 4
|
| 208 |
-
- checkpoint selection: lowest validation soft cross-entropy
|
| 209 |
|
| 210 |
-
##
|
| 211 |
-
|
| 212 |
-
Ariadne is trained on:
|
| 213 |
-
|
| 214 |
-
`LocalLLaMA/typed-decisions`
|
| 215 |
-
|
| 216 |
-
The primary benchmark contains:
|
| 217 |
|
| 218 |
-
|
| 219 |
-
- **2,000 individual decisions**
|
| 220 |
|
| 221 |
-
|
| 222 |
|
| 223 |
-
|
| 224 |
|
| 225 |
-
|
| 226 |
|
| 227 |
-
|
| 228 |
|
| 229 |
-
|
| 230 |
|
| 231 |
-
|
| 232 |
|
| 233 |
-
|
| 234 |
|
| 235 |
-
-
|
| 236 |
-
- soft accuracy
|
| 237 |
-
- soft cross-entropy
|
| 238 |
-
- hard Brier score
|
| 239 |
-
- soft Brier score
|
| 240 |
-
- expected calibration error (ECE)
|
| 241 |
|
| 242 |
-
|
| 243 |
|
| 244 |
-
|
|
|
|
| 245 |
|
| 246 |
-
|
| 247 |
-
|
| 248 |
-
|
| 249 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 250 |
```
|
| 251 |
|
| 252 |
-
|
| 253 |
-
|
| 254 |
-
Hard decision accuracy and probability calibration should be treated separately.
|
| 255 |
-
|
| 256 |
-
The upstream Laya checkpoint may emit warnings for stored temperature values outside the runtime-supported calibration range. Laya clamps affected values and warns that confidence for those entries should be treated as uncalibrated.
|
| 257 |
-
|
| 258 |
-
Ariadne's improvements in hard accuracy should therefore **not** be interpreted as evidence of improved calibration.
|
| 259 |
|
| 260 |
-
|
| 261 |
|
| 262 |
-
|
| 263 |
|
| 264 |
-
|
| 265 |
-
|
| 266 |
-
|
| 267 |
-
|
| 268 |
-
|
| 269 |
-
|
| 270 |
-
> **Conventional fine-tuning moves the representation toward the task. Ariadne attempts to move the task toward the representation.**
|
| 271 |
-
|
| 272 |
-
## Intended use
|
| 273 |
-
|
| 274 |
-
Ariadne is intended for research and experimentation involving:
|
| 275 |
-
|
| 276 |
-
- representation alignment
|
| 277 |
-
- frozen encoders
|
| 278 |
-
- parameter-efficient adaptation
|
| 279 |
-
- System-1 and selector models
|
| 280 |
-
- typed decision systems
|
| 281 |
-
- reuse of pretrained representations
|
| 282 |
-
- comparison with conventional full fine-tuning
|
| 283 |
-
- comparison with PEFT methods such as LoRA and adapters
|
| 284 |
-
|
| 285 |
-
The approach may be particularly useful when many specialist tasks need to share a large common frozen encoder.
|
| 286 |
-
|
| 287 |
-
## Limitations
|
| 288 |
-
|
| 289 |
-
- The current release focuses on English typed decisions.
|
| 290 |
-
- The primary benchmark consists of synthetic decision workflows.
|
| 291 |
-
- Frozen-base training currently shows significant convergence sensitivity.
|
| 292 |
-
- Individual high-performing runs should not be confused with multi-seed average performance.
|
| 293 |
-
- The deterministic frozen-base five-seed sweep is still being completed.
|
| 294 |
-
- The typed-decisions test set was inspected during development.
|
| 295 |
-
- Cross-task results should not be assumed to generalize to arbitrary domains.
|
| 296 |
-
- Hard accuracy does not imply good probability calibration.
|
| 297 |
-
- Ariadne depends on the upstream Laya model and is not a standalone foundation model.
|
| 298 |
-
- The model should not be used for safety-critical decisions without independent task-specific validation.
|
| 299 |
-
|
| 300 |
-
## Reproducibility
|
| 301 |
-
|
| 302 |
-
The experimental artifacts include:
|
| 303 |
-
|
| 304 |
-
- deterministic and non-deterministic training histories
|
| 305 |
-
- per-seed validation metrics
|
| 306 |
-
- per-task predictions
|
| 307 |
-
- checkpoint hashes
|
| 308 |
-
- frozen-parameter verification
|
| 309 |
-
- machine-readable benchmark reports
|
| 310 |
-
- per-seed/per-task CSV results
|
| 311 |
-
|
| 312 |
-
The test suite verifies that interface-only experiments do not alter the frozen upstream model parameters.
|
| 313 |
-
|
| 314 |
-
The current audited experimental suite includes **43 passing tests**.
|
| 315 |
-
|
| 316 |
-
## Relationship to Laya
|
| 317 |
|
| 318 |
-
|
| 319 |
|
| 320 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 321 |
|
| 322 |
-
|
|
|
|
|
|
|
|
|
|
| 323 |
|
| 324 |
-
##
|
| 325 |
|
| 326 |
-
|
| 327 |
|
| 328 |
-
|
|
|
|
| 2 |
license: apache-2.0
|
| 3 |
language:
|
| 4 |
- en
|
|
|
|
|
|
|
|
|
|
| 5 |
pipeline_tag: text-classification
|
| 6 |
+
base_model: convaiinnovations/laya
|
| 7 |
+
base_model_relation: adapter
|
| 8 |
+
datasets:
|
| 9 |
+
- LocalLLaMA/typed-decisions
|
| 10 |
tags:
|
| 11 |
- laya
|
| 12 |
+
- modernbert
|
|
|
|
|
|
|
| 13 |
- adapter
|
| 14 |
+
- typed-decisions
|
| 15 |
+
- custom-code
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
---
|
|
|
|
| 17 |
|
| 18 |
+
# Ariadne-Laya-TD — stable training with a small input adapter
|
| 19 |
|
| 20 |
+
**A small typed-decision adapter for frozen Laya, with a more stable training recipe.**
|
| 21 |
|
| 22 |
+
A 1,049,600-parameter affine input adapter trained on typed decisions while every parameter of base Laya stays frozen. It requires the pinned base model; the small adapter file is not a standalone model.
|
|
|
|
|
|
|
| 23 |
|
| 24 |
+
The included checkpoint scores **75.45%** on the local Typed Decisions test, compared with **36.25%** for the matched base. It is seed 4, epoch 8, selected by validation loss. This release uses LR 0.0001.
|
| 25 |
|
| 26 |
+
Across five deterministic runs, typed-decision accuracy was **75.28% ± 0.93 percentage points** (sample SD). The matched native base scored **36.25%**. The table below reports all eight tasks and every seed.
|
| 27 |
|
| 28 |
+
**Lower learning rate improved training stability.** Changing LR from 3e-4 to 1e-4 raised mean typed-decision accuracy from 70.23% to 75.28% and reduced the across-seed standard deviation from 6.68 to 0.93 points. Sample variance fell by 98.1% in this five-seed comparison. The complete learning-rate comparison is below.
|
| 29 |
|
| 30 |
+
**Consistency across seeds.** Typed-decision accuracy ranged from 74.25% to 76.55%. AG News stayed between 93.67% and 95.17%, and MASSIVE between 75.00% and 79.00%. These are measurements of this recipe on five seeds and the fixed benchmark bundle.
|
|
|
|
|
|
|
|
|
|
| 31 |
|
| 32 |
## Architecture
|
| 33 |
|
| 34 |
+
The map is `x′ = Wx + b`, initialized with `W = I`, `b = 0`, immediately after Laya's original token embedding and normalization. Original embeddings, normalization, all 28 ModernBERT blocks, and decision heads remain frozen and in evaluation mode. Gradients pass through the frozen stack to update only the map. There is no added LLM and no extra BERT forward pass.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
|
| 36 |
+
The combined model has 422,343,427 parameters; 1,049,600 were trainable (about 0.25%). Each FP32 adapter checkpoint is approximately 4.2 MB. The original model and tokenizer are downloaded separately.
|
| 37 |
|
| 38 |
+
Base: [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya/tree/55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851), revision `55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851`. Source weights were hashed before training and at checkpoint saves; all pretrained tensors remained unchanged.
|
| 39 |
|
| 40 |
+
## Evaluation
|
|
|
|
|
|
|
|
|
|
|
|
|
| 41 |
|
| 42 |
+
All scores below are local measurements on identical frozen cases. Accuracy is the highest-probability label accuracy, including ordinal questions; failed decisions count as incorrect. Values are percentages; variance is sample variance in percentage-points squared. Five seeds are used for the mean and SD. These are not the upstream model card's historical benchmark numbers.
|
| 43 |
|
| 44 |
+
| Benchmark | Decisions | Native base | Adapter mean ± SD | Variance | Selected checkpoint |
|
| 45 |
+
|---|---:|---:|---:|---:|---:|
|
| 46 |
+
| Typed decisions | 2000 | 36.25 | 75.28 ± 0.93 | 0.8620 | 75.45 |
|
| 47 |
+
| AG News | 600 | 94.83 | 94.47 ± 0.58 | 0.3389 | 95.17 |
|
| 48 |
+
| BoolQ | 600 | 83.00 | 82.10 ± 0.89 | 0.7861 | 83.17 |
|
| 49 |
+
| DAIR Emotion | 600 | 57.33 | 57.20 ± 1.00 | 1.0056 | 57.83 |
|
| 50 |
+
| Prompt injections | 116 | 68.97 | 74.48 ± 3.93 | 15.4578 | 74.14 |
|
| 51 |
+
| SST-5 | 600 | 37.00 | 41.00 ± 2.46 | 6.0278 | 41.00 |
|
| 52 |
+
| MASSIVE intent EN | 300 | 78.67 | 76.80 ± 1.56 | 2.4222 | 77.00 |
|
| 53 |
+
| XNLI EN | 300 | 86.00 | 87.27 ± 1.04 | 1.0778 | 86.33 |
|
| 54 |
|
| 55 |
+
### What changed from the first release
|
| 56 |
|
| 57 |
+
The earlier adapter used LR 3e-4. This version uses 1e-4 with the same seeds, data, initialization, batch, stopping rule and deterministic runtime. The table compares all five runs of each recipe. Different stopping epochs are a consequence of the shared validation rule.
|
| 58 |
|
| 59 |
+
| Benchmark | Earlier LR 3e-4 mean ± SD | This LR 1e-4 mean ± SD |
|
| 60 |
|---|---:|---:|
|
| 61 |
+
| Typed decisions | 70.23 ± 6.68 | 75.28 ± 0.93 |
|
| 62 |
+
| AG News | 54.57 ± 35.43 | 94.47 ± 0.58 |
|
| 63 |
+
| BoolQ | 63.27 ± 16.02 | 82.10 ± 0.89 |
|
| 64 |
+
| DAIR Emotion | 30.33 ± 24.82 | 57.20 ± 1.00 |
|
| 65 |
+
| Prompt injections | 57.76 ± 15.54 | 74.48 ± 3.93 |
|
| 66 |
+
| SST-5 | 30.53 ± 13.61 | 41.00 ± 2.46 |
|
| 67 |
+
| MASSIVE intent EN | 32.53 ± 39.58 | 76.80 ± 1.56 |
|
| 68 |
+
| XNLI EN | 55.53 ± 29.33 | 87.27 ± 1.04 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 69 |
|
| 70 |
+
The earlier published checkpoint was seed 1, epoch 14. The new validation-selected checkpoint improves AG News from 93.17% to 95.17%, BoolQ from 79.83% to 83.17%, and MASSIVE from 74.33% to 77.00%. Its typed-decision score is 75.45%, compared with 76.95% for the earlier checkpoint. The five-seed means above measure the improvement in the training recipe.
|
| 71 |
|
| 72 |
+
### Individual seeds
|
| 73 |
|
| 74 |
+
| Seed | Typed | AG News | BoolQ | Emotion | Injection | SST-5 | MASSIVE | XNLI | Best epoch | Stop epoch |
|
| 75 |
+
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 76 |
+
| 0 | 74.25 | 94.17 | 80.83 | 58.50 | 69.83 | 38.17 | 79.00 | 89.00 | 11 | 14 |
|
| 77 |
+
| 1 | 75.65 | 94.83 | 82.67 | 57.17 | 78.45 | 44.83 | 75.67 | 87.00 | 9 | 12 |
|
| 78 |
+
| 2 | 74.50 | 94.50 | 81.83 | 56.00 | 71.55 | 39.83 | 75.00 | 87.33 | 13 | 16 |
|
| 79 |
+
| 3 | 76.55 | 93.67 | 82.00 | 56.50 | 78.45 | 41.17 | 77.33 | 86.67 | 8 | 11 |
|
| 80 |
+
| 4 | 75.45 | 95.17 | 83.17 | 57.83 | 74.14 | 41.00 | 77.00 | 86.33 | 8 | 11 |
|
| 81 |
|
| 82 |
+
### Selected checkpoint versus base: paired uncertainty
|
| 83 |
|
| 84 |
+
These descriptive 95% intervals resample cases 10,000 times, keeping each typed case's five decisions together. They describe benchmark-case uncertainty for the selected model, separately from the across-seed SD above. They have no multiple-comparison correction and were not used to select the checkpoint.
|
| 85 |
|
| 86 |
+
| Benchmark | Change from base (pp) | Paired 95% interval (pp) |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 87 |
|---|---:|---:|
|
| 88 |
+
| Typed decisions | +39.20 | [+35.90, +42.50] |
|
| 89 |
+
| AG News | +0.33 | [-0.50, +1.33] |
|
| 90 |
+
| BoolQ | +0.17 | [-1.50, +1.83] |
|
| 91 |
+
| DAIR Emotion | +0.50 | [-1.50, +2.67] |
|
| 92 |
+
| Prompt injections | +5.17 | [+1.72, +9.48] |
|
| 93 |
+
| SST-5 | +4.00 | [+1.00, +6.83] |
|
| 94 |
+
| MASSIVE intent EN | -1.67 | [-4.33, +1.00] |
|
| 95 |
+
| XNLI EN | +0.33 | [-1.67, +2.33] |
|
| 96 |
|
| 97 |
+
### Checkpoint selection
|
| 98 |
|
| 99 |
+
The default package contains **seed 4, epoch 8**, selected by the lowest validation soft cross-entropy across the five seeds (0.854020). This selection rule was recorded before the sweep. Test performance was not used to choose the checkpoint. Only the selected adapter weights are included; all five seeds’ results are reported.
|
| 100 |
|
| 101 |
+
### Benchmark scope
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 102 |
|
| 103 |
+
- Typed Decisions: 400 test cases / 2,000 decisions from four synthetic workflows. Training and validation case IDs are disjoint from test IDs. The test task itself is the adaptation task.
|
| 104 |
+
- AG News, BoolQ, Emotion and SST-5: fixed prefixes of 600 examples each. Prompt injections: all 116 test examples. MASSIVE and XNLI: first 300 English test examples each.
|
| 105 |
+
- MASSIVE uses the gold intent plus 19 seeded distractors (20 options), not all intents. XNLI uses three labels.
|
| 106 |
+
- These seven retention tasks were excluded from adapter training; some were present in the base model's training. These are held-out examples for this adaptation, not proof of wholly unseen-task generalization.
|
| 107 |
+
- The test suite was inspected during earlier exploratory experiments. This recipe was selected adaptively after earlier results; seed 0 began as the lower-LR pilot and was retained for the five-seed confirmation. Checkpoint selection used validation only, but the research process is not a blind test. Earlier runs with changing stopping rules are excluded from the main table.
|
| 108 |
+
- No downstream deployment, multilingual retention, or broad safety claim is established by these small English benchmark subsets.
|
| 109 |
|
| 110 |
## Training
|
| 111 |
|
| 112 |
+
Training data: [LocalLLaMA/typed-decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions), revision `f7a2487edd7a043a5441a5e9ccc7fe5ddbd9ebe8`. A case-level split with seed 42 yields 1,080 training cases (5,400 decisions) and 120 validation cases (600 decisions). Questions from one case never cross splits.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 113 |
|
| 114 |
+
- Seeds 0, 1, 2, 3, 4; RTX PRO 6000 Blackwell only, sequential runs.
|
| 115 |
+
- Every map starts at the same identity transformation and native dropout stays off. Seeds control training-example shuffling; this comparison measures sensitivity to data order with deterministic computation.
|
| 116 |
+
- AdamW, LR 0.0001, weight decay 0.01, gradient clipping at norm 1; effective batch 32 (microbatch 8 × accumulation 4). No learning-rate warmup or scheduler.
|
| 117 |
+
- Soft-target cross-entropy on raw logits. The pretrained stack stays in eval mode; its dropout is off. The affine map is computed in FP32; the remaining CUDA forward uses BF16 autocast.
|
| 118 |
+
- Minimum 10 completed epochs; stop after three consecutive epochs without a strictly lower validation loss. Patience counts from epoch 1, with stopping gated by the minimum. No fixed maximum. Keep the best trained checkpoint; epoch zero is a separate baseline.
|
| 119 |
+
- Input/head token budgets: 1024/256, shared by native baseline and adapter. These differ from the original base config's defaults, so this card reports matched-budget local baselines.
|
|
|
|
|
|
|
|
|
|
| 120 |
|
| 121 |
+
## Reproducibility
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 122 |
|
| 123 |
+
Strict deterministic algorithms were enabled with errors on unsupported operations, cuDNN benchmarking and cuDNN SDPA were disabled, and float32 matmul precision was set to highest. Python, NumPy and PyTorch seeds were set. Processes used `PYTHONHASHSEED=0` and `CUBLAS_WORKSPACE_CONFIG=:4096:8`.
|
|
|
|
| 124 |
|
| 125 |
+
Two independent two-epoch seed-0 verification runs (338 optimizer updates each) matched losses, validation metrics, and checkpoint tensors bit for bit. The reported seed-0 run matched that two-epoch metric prefix. This verifies the checked environment and prefix; it does not promise bitwise agreement across GPUs, libraries or operating systems. See `environment.json` and `evidence/replay-check.json`.
|
| 126 |
|
| 127 |
+
Before training, the identity adapter matched every native-base answer object on all 5,116 benchmark decisions. All five final runs were independently rescored and checked for unchanged pretrained weights. Export verification is recorded separately in `export-verification.json` after testing the portable package.
|
| 128 |
|
| 129 |
+
## Confidence and temperature warning
|
| 130 |
|
| 131 |
+
The pinned base checkpoint ships `choice:11+` temperature 0.1005828. Laya 0.3.20 clamps it to 0.5 and emits a warning on load. Both baseline and adapters use that same behavior. Only the 300 MASSIVE examples use this bucket in this suite. These temperatures are not used in the training loss.
|
| 132 |
|
| 133 |
+
Temperature scaling preserves class order mathematically but changes probability sharpness and calibration metrics. Calibration was not refitted after adaptation. Confidence calibration should be evaluated for the intended application; the base model's calibration claims have not been established for this adapter.
|
| 134 |
|
| 135 |
+
## Load locally
|
| 136 |
|
| 137 |
+
This is a custom adapter. Use the bundled loader: `laya.load(repo_id)` and `AutoModel.from_pretrained(repo_id)` do not install this input transformation.
|
| 138 |
|
| 139 |
+
Download the repository files with the Hugging Face UI or `hf download GoatHerder/Ariadne-Laya-TD --local-dir ariadne-laya-td`, then change into that directory. The learning rate and selected seed are recorded in `training_protocol.json` and `adapter_config.json`.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 140 |
|
| 141 |
+
Install a PyTorch build for your GPU and the versions in `requirements.txt`. The measured environment used PyTorch 2.13.0+cu130; package versions and CUDA/cuDNN details are in `environment.json`. Run from this downloaded repository directory:
|
| 142 |
|
| 143 |
+
```python
|
| 144 |
+
from load_adapter import load
|
| 145 |
|
| 146 |
+
model = load(device="cuda:0") # defaults to the validation-selected seed
|
| 147 |
+
answer = model.predict("The customer says the delivery arrived damaged.", {
|
| 148 |
+
"route": {"type": "choice", "instructions": "Choose the support queue.",
|
| 149 |
+
"criteria": {"delivery": "Delivery and damaged items",
|
| 150 |
+
"billing": "Payments and invoices"}}
|
| 151 |
+
})
|
| 152 |
+
print(answer)
|
| 153 |
+
model.close()
|
| 154 |
```
|
| 155 |
|
| 156 |
+
Only the validation-selected seed is included in this compact package. Use `device='cpu'` for CPU inference. The loader downloads the pinned base, verifies base configuration/tokenizer hashes and model tensors, checks adapter weights, then freezes all parameters for inference. `local_files_only=True` uses cached weights; `base_path` can point to an existing copy of the pinned base.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 157 |
|
| 158 |
+
## Reproduce preparation and one training run
|
| 159 |
|
| 160 |
+
The bundled `ariadne_bench` source contains the actual tokenizer/prompt preparation, training loop, adapters and scorer. Dataset files are downloaded from their pinned sources; evaluation data is not redistributed.
|
| 161 |
|
| 162 |
+
```bash
|
| 163 |
+
python -m ariadne_bench.experiments.prepare_alignment --out data/alignment
|
| 164 |
+
python -m ariadne_bench prepare --profile laya-core \
|
| 165 |
+
--suites typed-decisions,ag-news,emotion,boolq,sst5,prompt-injections,massive-intent.en,xnli.en \
|
| 166 |
+
--out data/heldout-eight
|
| 167 |
+
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 168 |
|
| 169 |
+
The expected case-file hashes and immutable dataset revisions are in `training_protocol.json` and `benchmarks/sources.lock.json`. `make_training_config.py` resolves the pinned base into a local training config:
|
| 170 |
|
| 171 |
+
```bash
|
| 172 |
+
python make_training_config.py --seed 0 --device cuda:0 --out train-config.json
|
| 173 |
+
PYTHONHASHSEED=0 CUBLAS_WORKSPACE_CONFIG=:4096:8 USE_TF=0 OMP_NUM_THREADS=4 \
|
| 174 |
+
python -m ariadne_bench.experiments.align --config train-config.json \
|
| 175 |
+
--train data/alignment/train --validation data/alignment/validation \
|
| 176 |
+
--out training-seed-0 --epochs 0 --min-epochs 10 --early-stopping-patience 3 \
|
| 177 |
+
--batch-size 8 --accumulation 4 --lr 0.0001 --keep-best-only
|
| 178 |
|
| 179 |
+
PYTHONHASHSEED=0 CUBLAS_WORKSPACE_CONFIG=:4096:8 USE_TF=0 OMP_NUM_THREADS=4 \
|
| 180 |
+
python -m ariadne_bench run --data data/heldout-eight \
|
| 181 |
+
--config training-seed-0/benchmark-config.json --out evaluation-seed-0 --warmup 5
|
| 182 |
+
```
|
| 183 |
|
| 184 |
+
## License and attribution
|
| 185 |
|
| 186 |
+
Apache-2.0. The [base-model card](https://huggingface.co/convaiinnovations/laya/blob/55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851/README.md) and [Typed Decisions dataset card](https://huggingface.co/datasets/LocalLLaMA/typed-decisions/blob/f7a2487edd7a043a5441a5e9ccc7fe5ddbd9ebe8/README.md) list Apache-2.0. Base weights are downloaded from their original repository. Evaluation datasets retain their own licenses. Prompt preparation credits and upstream notices are in `THIRD_PARTY.md` and `licenses/`.
|
| 187 |
|
| 188 |
+
This is an independent adaptation of Laya. The upstream authors did not produce these adapter weights or experimental results.
|
USAGE.md
CHANGED
|
@@ -1,48 +1,3 @@
|
|
| 1 |
-
#
|
| 2 |
|
| 3 |
-
This package
|
| 4 |
-
|
| 5 |
-
The ongoing lower-learning-rate experiment is separate. These are the already verified weights, not a claim that the training stability investigation is complete.
|
| 6 |
-
|
| 7 |
-
## Upload
|
| 8 |
-
|
| 9 |
-
Extract the ZIP and upload the files and folders **at the root of your existing Hugging Face model repository**. There is deliberately no README.md in this bundle, so it can coexist with your existing model card. The ZIP is a transport archive; upload its extracted contents for normal use.
|
| 10 |
-
|
| 11 |
-
This is a custom input adapter. Loading requires the bundled loader; bare `laya.load(repo_id)` and `AutoModel.from_pretrained(repo_id)` do not install the extra affine layer.
|
| 12 |
-
|
| 13 |
-
## Load a downloaded copy
|
| 14 |
-
|
| 15 |
-
Install an appropriate PyTorch build and the versions in `requirements.txt`. The measured GPU environment is recorded in `environment.json`. From the downloaded repository directory:
|
| 16 |
-
|
| 17 |
-
```python
|
| 18 |
-
from load_adapter import load
|
| 19 |
-
|
| 20 |
-
model = load(device="cuda:0")
|
| 21 |
-
answer = model.predict("The delivery arrived damaged.", {
|
| 22 |
-
"route": {
|
| 23 |
-
"type": "choice",
|
| 24 |
-
"instructions": "Choose the customer-support queue.",
|
| 25 |
-
"criteria": {"delivery": "Delivery and damaged items", "billing": "Payments and invoices"},
|
| 26 |
-
}
|
| 27 |
-
})
|
| 28 |
-
print(answer)
|
| 29 |
-
model.close()
|
| 30 |
-
```
|
| 31 |
-
|
| 32 |
-
After you upload, first download your repository with `huggingface_hub.snapshot_download("YOUR_ACCOUNT/YOUR_REPOSITORY")`, then run the example from that downloaded directory. Use `device="cpu"` for CPU inference. `base_path` can point to the pinned base snapshot already on disk; `local_files_only=True` uses cached base files.
|
| 33 |
-
|
| 34 |
-
The loader verifies the adapter checksum, base configuration/tokenizer hashes and pretrained tensor hash. Strict determinism is enabled by default and requires the CUBLAS workspace environment to be set before CUDA is initialized. A fresh process handles this automatically. Numerical agreement across other hardware or software is not guaranteed.
|
| 35 |
-
|
| 36 |
-
## Measured performance and limits
|
| 37 |
-
|
| 38 |
-
Selected checkpoint accuracy: typed decisions **76.95%**, AG News **93.17%**, BoolQ **79.83%**, Emotion **57.33%**, prompt injections **71.55%**, SST-5 **42.00%**, MASSIVE EN **74.33%**, XNLI EN **87.67%**. Native base typed accuracy was 36.25% under the matched protocol.
|
| 39 |
-
|
| 40 |
-
Across all five training seeds, typed accuracy was **70.23% ± 6.68 pp** (sample SD). Three seeds had severe retention losses. The selected model also lost 3.17 pp on BoolQ and 4.33 pp on MASSIVE versus base. Full results and paired uncertainty are in `metrics.json` and `evidence/candidate-uncertainty.json`. Other seeds' scores are included for transparency; this bundle contains only seed 1's weights.
|
| 41 |
-
|
| 42 |
-
The benchmark uses 2,000 typed decisions plus fixed subsets of seven other tasks, totaling 5,116 decisions. MASSIVE uses 20 candidate intents. The same test subsets were inspected in earlier experiments. These results describe an experimental task adapter, not established broad task improvement.
|
| 43 |
-
|
| 44 |
-
The base checkpoint's choice:11+ temperature is clamped from 0.1005828 to 0.5 by Laya 0.3.20. This expected load warning concerns confidence calibration; raw-logit training and class ordering are unaffected. Calibration was not refitted after adaptation.
|
| 45 |
-
|
| 46 |
-
All 5,116 answer objects matched the original selected checkpoint when the portable export was tested from outside the workspace. `export-verification.json` records that check. `upload-manifest.json` verifies that this bundle keeps the same weights and inference code, with the other seeds removed from its loading configuration.
|
| 47 |
-
|
| 48 |
-
License: Apache-2.0. Base model: `convaiinnovations/laya`, pinned to revision `55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851`. Attribution is in `THIRD_PARTY.md` and `licenses/`.
|
|
|
|
| 1 |
+
# Using this adapter
|
| 2 |
|
| 3 |
+
This package uses LR 0.0001, seed 4, epoch 8. See [the model card](README.md#load-locally) for the loading example, all five seeds' results, requirements and limitations. Load with the bundled `load_adapter.py`; the adapter weights require the pinned base Laya.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
adapter.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 4198616
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081
|
| 3 |
size 4198616
|
adapter_config.json
CHANGED
|
@@ -4,7 +4,7 @@
|
|
| 4 |
"position": "after native embeddings, before ModernBERT block 0",
|
| 5 |
"base_model": "convaiinnovations/laya",
|
| 6 |
"base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
|
| 7 |
-
"candidate_seed":
|
| 8 |
"max_len": 1024,
|
| 9 |
"head_max_len": 256,
|
| 10 |
"trainable_parameters_during_training": 1049600,
|
|
@@ -16,10 +16,10 @@
|
|
| 16 |
"tokenizer/tokenizer_config.json": "50044de60daaa73df97d262e15a40d4faf0160e7d742df64b377877a1320dd12"
|
| 17 |
},
|
| 18 |
"seeds": {
|
| 19 |
-
"
|
| 20 |
"weights_file": "adapter.safetensors",
|
| 21 |
-
"weights_sha256": "
|
| 22 |
-
"selected_epoch":
|
| 23 |
}
|
| 24 |
}
|
| 25 |
}
|
|
|
|
| 4 |
"position": "after native embeddings, before ModernBERT block 0",
|
| 5 |
"base_model": "convaiinnovations/laya",
|
| 6 |
"base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
|
| 7 |
+
"candidate_seed": 4,
|
| 8 |
"max_len": 1024,
|
| 9 |
"head_max_len": 256,
|
| 10 |
"trainable_parameters_during_training": 1049600,
|
|
|
|
| 16 |
"tokenizer/tokenizer_config.json": "50044de60daaa73df97d262e15a40d4faf0160e7d742df64b377877a1320dd12"
|
| 17 |
},
|
| 18 |
"seeds": {
|
| 19 |
+
"4": {
|
| 20 |
"weights_file": "adapter.safetensors",
|
| 21 |
+
"weights_sha256": "f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081",
|
| 22 |
+
"selected_epoch": 8
|
| 23 |
}
|
| 24 |
}
|
| 25 |
}
|
evidence/adapter-equivalence.json
CHANGED
|
@@ -35,5 +35,10 @@
|
|
| 35 |
"forward_exact": true,
|
| 36 |
"exact_match": true
|
| 37 |
}
|
| 38 |
-
]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
}
|
|
|
|
| 35 |
"forward_exact": true,
|
| 36 |
"exact_match": true
|
| 37 |
}
|
| 38 |
+
],
|
| 39 |
+
"reuse_provenance": {
|
| 40 |
+
"experiment": "E0i",
|
| 41 |
+
"source_sha256": "c77faa691bd273d363dbb196155545b4f72697b9d0f4c80451ffb06b0090d200",
|
| 42 |
+
"reason": "unchanged adapter computation and strict runtime; this forward/backward diagnostic has no optimizer step or learning-rate dependence"
|
| 43 |
+
}
|
| 44 |
}
|
evidence/audit.json
CHANGED
|
@@ -5,52 +5,52 @@
|
|
| 5 |
"runs": [
|
| 6 |
{
|
| 7 |
"seed": 0,
|
| 8 |
-
"checkpoint_sha256": "
|
| 9 |
"checkpoint_bytes": 4198616,
|
| 10 |
"all_pretrained_tensors_unchanged": true,
|
| 11 |
"independent_scoring_matches": true,
|
| 12 |
"min_epochs": 10,
|
| 13 |
-
"stop_epoch":
|
| 14 |
"minimum_epoch_rule_verified": true
|
| 15 |
},
|
| 16 |
{
|
| 17 |
"seed": 1,
|
| 18 |
-
"checkpoint_sha256": "
|
| 19 |
"checkpoint_bytes": 4198616,
|
| 20 |
"all_pretrained_tensors_unchanged": true,
|
| 21 |
"independent_scoring_matches": true,
|
| 22 |
"min_epochs": 10,
|
| 23 |
-
"stop_epoch":
|
| 24 |
"minimum_epoch_rule_verified": true
|
| 25 |
},
|
| 26 |
{
|
| 27 |
"seed": 2,
|
| 28 |
-
"checkpoint_sha256": "
|
| 29 |
"checkpoint_bytes": 4198616,
|
| 30 |
"all_pretrained_tensors_unchanged": true,
|
| 31 |
"independent_scoring_matches": true,
|
| 32 |
"min_epochs": 10,
|
| 33 |
-
"stop_epoch":
|
| 34 |
"minimum_epoch_rule_verified": true
|
| 35 |
},
|
| 36 |
{
|
| 37 |
"seed": 3,
|
| 38 |
-
"checkpoint_sha256": "
|
| 39 |
"checkpoint_bytes": 4198616,
|
| 40 |
"all_pretrained_tensors_unchanged": true,
|
| 41 |
"independent_scoring_matches": true,
|
| 42 |
"min_epochs": 10,
|
| 43 |
-
"stop_epoch":
|
| 44 |
"minimum_epoch_rule_verified": true
|
| 45 |
},
|
| 46 |
{
|
| 47 |
"seed": 4,
|
| 48 |
-
"checkpoint_sha256": "
|
| 49 |
"checkpoint_bytes": 4198616,
|
| 50 |
"all_pretrained_tensors_unchanged": true,
|
| 51 |
"independent_scoring_matches": true,
|
| 52 |
"min_epochs": 10,
|
| 53 |
-
"stop_epoch":
|
| 54 |
"minimum_epoch_rule_verified": true
|
| 55 |
}
|
| 56 |
],
|
|
|
|
| 5 |
"runs": [
|
| 6 |
{
|
| 7 |
"seed": 0,
|
| 8 |
+
"checkpoint_sha256": "4794ce39ed5fe621ef8deb7e40900590951f7382e9c7e15543f8d0ef4b41df6b",
|
| 9 |
"checkpoint_bytes": 4198616,
|
| 10 |
"all_pretrained_tensors_unchanged": true,
|
| 11 |
"independent_scoring_matches": true,
|
| 12 |
"min_epochs": 10,
|
| 13 |
+
"stop_epoch": 14,
|
| 14 |
"minimum_epoch_rule_verified": true
|
| 15 |
},
|
| 16 |
{
|
| 17 |
"seed": 1,
|
| 18 |
+
"checkpoint_sha256": "0441f00f59c24e9b6efe2e01e71fefc31ca5e4b0da70ed2a5a2051a385782ddc",
|
| 19 |
"checkpoint_bytes": 4198616,
|
| 20 |
"all_pretrained_tensors_unchanged": true,
|
| 21 |
"independent_scoring_matches": true,
|
| 22 |
"min_epochs": 10,
|
| 23 |
+
"stop_epoch": 12,
|
| 24 |
"minimum_epoch_rule_verified": true
|
| 25 |
},
|
| 26 |
{
|
| 27 |
"seed": 2,
|
| 28 |
+
"checkpoint_sha256": "1a4cebb05017866daa8f339671599ac21ff3125fcbfac8605845813ef540c007",
|
| 29 |
"checkpoint_bytes": 4198616,
|
| 30 |
"all_pretrained_tensors_unchanged": true,
|
| 31 |
"independent_scoring_matches": true,
|
| 32 |
"min_epochs": 10,
|
| 33 |
+
"stop_epoch": 16,
|
| 34 |
"minimum_epoch_rule_verified": true
|
| 35 |
},
|
| 36 |
{
|
| 37 |
"seed": 3,
|
| 38 |
+
"checkpoint_sha256": "1ce28c37f42b247b9bdb51d126bc189161f33d8ff28990b922a3041e5a3ff911",
|
| 39 |
"checkpoint_bytes": 4198616,
|
| 40 |
"all_pretrained_tensors_unchanged": true,
|
| 41 |
"independent_scoring_matches": true,
|
| 42 |
"min_epochs": 10,
|
| 43 |
+
"stop_epoch": 11,
|
| 44 |
"minimum_epoch_rule_verified": true
|
| 45 |
},
|
| 46 |
{
|
| 47 |
"seed": 4,
|
| 48 |
+
"checkpoint_sha256": "f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081",
|
| 49 |
"checkpoint_bytes": 4198616,
|
| 50 |
"all_pretrained_tensors_unchanged": true,
|
| 51 |
"independent_scoring_matches": true,
|
| 52 |
"min_epochs": 10,
|
| 53 |
+
"stop_epoch": 11,
|
| 54 |
"minimum_epoch_rule_verified": true
|
| 55 |
}
|
| 56 |
],
|
evidence/candidate-uncertainty.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
{
|
| 2 |
-
"candidate_seed":
|
| 3 |
-
"selected_epoch":
|
| 4 |
"benchmark_cases_sha256": "10fb671245cdac1ee09e9c4d2b9b09829fbf1abd06e0954fddee0b9ec8519854",
|
| 5 |
"resamples": 10000,
|
| 6 |
"bootstrap_seed_per_suite": 42,
|
|
@@ -10,66 +10,66 @@
|
|
| 10 |
"suites": {
|
| 11 |
"typed-decisions": {
|
| 12 |
"cases": 400,
|
| 13 |
-
"accuracy_difference_pp":
|
| 14 |
"case_bootstrap_95_interval_pp": [
|
| 15 |
-
|
| 16 |
-
|
| 17 |
]
|
| 18 |
},
|
| 19 |
"ag-news": {
|
| 20 |
"cases": 600,
|
| 21 |
-
"accuracy_difference_pp":
|
| 22 |
"case_bootstrap_95_interval_pp": [
|
| 23 |
-
-
|
| 24 |
-
|
| 25 |
]
|
| 26 |
},
|
| 27 |
"boolq": {
|
| 28 |
"cases": 600,
|
| 29 |
-
"accuracy_difference_pp":
|
| 30 |
"case_bootstrap_95_interval_pp": [
|
| 31 |
-
-
|
| 32 |
-
|
| 33 |
]
|
| 34 |
},
|
| 35 |
"emotion": {
|
| 36 |
"cases": 600,
|
| 37 |
-
"accuracy_difference_pp": 0.
|
| 38 |
"case_bootstrap_95_interval_pp": [
|
| 39 |
-
-
|
| 40 |
2.6666666666666665
|
| 41 |
]
|
| 42 |
},
|
| 43 |
"prompt-injections": {
|
| 44 |
"cases": 116,
|
| 45 |
-
"accuracy_difference_pp":
|
| 46 |
"case_bootstrap_95_interval_pp": [
|
| 47 |
-
|
| 48 |
-
|
| 49 |
]
|
| 50 |
},
|
| 51 |
"sst5": {
|
| 52 |
"cases": 600,
|
| 53 |
-
"accuracy_difference_pp":
|
| 54 |
"case_bootstrap_95_interval_pp": [
|
| 55 |
-
1.
|
| 56 |
-
|
| 57 |
]
|
| 58 |
},
|
| 59 |
"massive-intent.en": {
|
| 60 |
"cases": 300,
|
| 61 |
-
"accuracy_difference_pp": -
|
| 62 |
"case_bootstrap_95_interval_pp": [
|
| 63 |
-
-
|
| 64 |
-
|
| 65 |
]
|
| 66 |
},
|
| 67 |
"xnli.en": {
|
| 68 |
"cases": 300,
|
| 69 |
-
"accuracy_difference_pp":
|
| 70 |
"case_bootstrap_95_interval_pp": [
|
| 71 |
-
-1.
|
| 72 |
-
|
| 73 |
]
|
| 74 |
}
|
| 75 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"candidate_seed": 4,
|
| 3 |
+
"selected_epoch": 8,
|
| 4 |
"benchmark_cases_sha256": "10fb671245cdac1ee09e9c4d2b9b09829fbf1abd06e0954fddee0b9ec8519854",
|
| 5 |
"resamples": 10000,
|
| 6 |
"bootstrap_seed_per_suite": 42,
|
|
|
|
| 10 |
"suites": {
|
| 11 |
"typed-decisions": {
|
| 12 |
"cases": 400,
|
| 13 |
+
"accuracy_difference_pp": 39.2,
|
| 14 |
"case_bootstrap_95_interval_pp": [
|
| 15 |
+
35.9,
|
| 16 |
+
42.5
|
| 17 |
]
|
| 18 |
},
|
| 19 |
"ag-news": {
|
| 20 |
"cases": 600,
|
| 21 |
+
"accuracy_difference_pp": 0.3333333333333333,
|
| 22 |
"case_bootstrap_95_interval_pp": [
|
| 23 |
+
-0.5,
|
| 24 |
+
1.3333333333333333
|
| 25 |
]
|
| 26 |
},
|
| 27 |
"boolq": {
|
| 28 |
"cases": 600,
|
| 29 |
+
"accuracy_difference_pp": 0.16666666666666666,
|
| 30 |
"case_bootstrap_95_interval_pp": [
|
| 31 |
+
-1.5,
|
| 32 |
+
1.8333333333333333
|
| 33 |
]
|
| 34 |
},
|
| 35 |
"emotion": {
|
| 36 |
"cases": 600,
|
| 37 |
+
"accuracy_difference_pp": 0.5,
|
| 38 |
"case_bootstrap_95_interval_pp": [
|
| 39 |
+
-1.5,
|
| 40 |
2.6666666666666665
|
| 41 |
]
|
| 42 |
},
|
| 43 |
"prompt-injections": {
|
| 44 |
"cases": 116,
|
| 45 |
+
"accuracy_difference_pp": 5.172413793103448,
|
| 46 |
"case_bootstrap_95_interval_pp": [
|
| 47 |
+
1.7241379310344827,
|
| 48 |
+
9.482758620689655
|
| 49 |
]
|
| 50 |
},
|
| 51 |
"sst5": {
|
| 52 |
"cases": 600,
|
| 53 |
+
"accuracy_difference_pp": 4.0,
|
| 54 |
"case_bootstrap_95_interval_pp": [
|
| 55 |
+
1.0,
|
| 56 |
+
6.833333333333333
|
| 57 |
]
|
| 58 |
},
|
| 59 |
"massive-intent.en": {
|
| 60 |
"cases": 300,
|
| 61 |
+
"accuracy_difference_pp": -1.6666666666666667,
|
| 62 |
"case_bootstrap_95_interval_pp": [
|
| 63 |
+
-4.333333333333333,
|
| 64 |
+
1.0
|
| 65 |
]
|
| 66 |
},
|
| 67 |
"xnli.en": {
|
| 68 |
"cases": 300,
|
| 69 |
+
"accuracy_difference_pp": 0.3333333333333333,
|
| 70 |
"case_bootstrap_95_interval_pp": [
|
| 71 |
+
-1.6666666666666667,
|
| 72 |
+
2.3333333333333335
|
| 73 |
]
|
| 74 |
}
|
| 75 |
}
|
evidence/lr-comparison.json
ADDED
|
@@ -0,0 +1,315 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"lr_3e4": {
|
| 3 |
+
"typed-decisions": {
|
| 4 |
+
"n": 5,
|
| 5 |
+
"mean_percent": 70.23,
|
| 6 |
+
"sample_variance_pp_squared": 44.56824999999997,
|
| 7 |
+
"sample_std_pp": 6.675945625901995
|
| 8 |
+
},
|
| 9 |
+
"ag-news": {
|
| 10 |
+
"n": 5,
|
| 11 |
+
"mean_percent": 54.56666666666666,
|
| 12 |
+
"sample_variance_pp_squared": 1255.536111111111,
|
| 13 |
+
"sample_std_pp": 35.433544997799906
|
| 14 |
+
},
|
| 15 |
+
"boolq": {
|
| 16 |
+
"n": 5,
|
| 17 |
+
"mean_percent": 63.266666666666666,
|
| 18 |
+
"sample_variance_pp_squared": 256.79999999999995,
|
| 19 |
+
"sample_std_pp": 16.0249804992081
|
| 20 |
+
},
|
| 21 |
+
"emotion": {
|
| 22 |
+
"n": 5,
|
| 23 |
+
"mean_percent": 30.333333333333332,
|
| 24 |
+
"sample_variance_pp_squared": 616.1666666666666,
|
| 25 |
+
"sample_std_pp": 24.82270466058577
|
| 26 |
+
},
|
| 27 |
+
"prompt-injections": {
|
| 28 |
+
"n": 5,
|
| 29 |
+
"mean_percent": 57.758620689655174,
|
| 30 |
+
"sample_variance_pp_squared": 241.5279429250891,
|
| 31 |
+
"sample_std_pp": 15.541169290793055
|
| 32 |
+
},
|
| 33 |
+
"sst5": {
|
| 34 |
+
"n": 5,
|
| 35 |
+
"mean_percent": 30.53333333333333,
|
| 36 |
+
"sample_variance_pp_squared": 185.14444444444447,
|
| 37 |
+
"sample_std_pp": 13.606779356057938
|
| 38 |
+
},
|
| 39 |
+
"massive-intent.en": {
|
| 40 |
+
"n": 5,
|
| 41 |
+
"mean_percent": 32.53333333333333,
|
| 42 |
+
"sample_variance_pp_squared": 1566.1999999999998,
|
| 43 |
+
"sample_std_pp": 39.57524478761944
|
| 44 |
+
},
|
| 45 |
+
"xnli.en": {
|
| 46 |
+
"n": 5,
|
| 47 |
+
"mean_percent": 55.53333333333334,
|
| 48 |
+
"sample_variance_pp_squared": 860.5333333333334,
|
| 49 |
+
"sample_std_pp": 29.33484844571953
|
| 50 |
+
}
|
| 51 |
+
},
|
| 52 |
+
"lr_1e4": {
|
| 53 |
+
"typed-decisions": {
|
| 54 |
+
"n": 5,
|
| 55 |
+
"mean_percent": 75.28,
|
| 56 |
+
"sample_variance_pp_squared": 0.8619999999999957,
|
| 57 |
+
"sample_std_pp": 0.9284395510748105
|
| 58 |
+
},
|
| 59 |
+
"ag-news": {
|
| 60 |
+
"n": 5,
|
| 61 |
+
"mean_percent": 94.46666666666667,
|
| 62 |
+
"sample_variance_pp_squared": 0.3388888888888897,
|
| 63 |
+
"sample_std_pp": 0.5821416398857667
|
| 64 |
+
},
|
| 65 |
+
"boolq": {
|
| 66 |
+
"n": 5,
|
| 67 |
+
"mean_percent": 82.10000000000001,
|
| 68 |
+
"sample_variance_pp_squared": 0.7861111111111168,
|
| 69 |
+
"sample_std_pp": 0.8866290718846956
|
| 70 |
+
},
|
| 71 |
+
"emotion": {
|
| 72 |
+
"n": 5,
|
| 73 |
+
"mean_percent": 57.2,
|
| 74 |
+
"sample_variance_pp_squared": 1.0055555555555546,
|
| 75 |
+
"sample_std_pp": 1.0027739304327543
|
| 76 |
+
},
|
| 77 |
+
"prompt-injections": {
|
| 78 |
+
"n": 5,
|
| 79 |
+
"mean_percent": 74.48275862068965,
|
| 80 |
+
"sample_variance_pp_squared": 15.457788347205712,
|
| 81 |
+
"sample_std_pp": 3.93163939689358
|
| 82 |
+
},
|
| 83 |
+
"sst5": {
|
| 84 |
+
"n": 5,
|
| 85 |
+
"mean_percent": 41.0,
|
| 86 |
+
"sample_variance_pp_squared": 6.027777777777775,
|
| 87 |
+
"sample_std_pp": 2.4551533104427055
|
| 88 |
+
},
|
| 89 |
+
"massive-intent.en": {
|
| 90 |
+
"n": 5,
|
| 91 |
+
"mean_percent": 76.8,
|
| 92 |
+
"sample_variance_pp_squared": 2.422222222222218,
|
| 93 |
+
"sample_std_pp": 1.556349003990499
|
| 94 |
+
},
|
| 95 |
+
"xnli.en": {
|
| 96 |
+
"n": 5,
|
| 97 |
+
"mean_percent": 87.26666666666667,
|
| 98 |
+
"sample_variance_pp_squared": 1.0777777777777784,
|
| 99 |
+
"sample_std_pp": 1.038160766826496
|
| 100 |
+
}
|
| 101 |
+
},
|
| 102 |
+
"paired_seed_accuracy_differences_pp": {
|
| 103 |
+
"0": {
|
| 104 |
+
"typed-decisions": 13.250000000000007,
|
| 105 |
+
"ag-news": 62.83333333333333,
|
| 106 |
+
"emotion": 38.99999999999999,
|
| 107 |
+
"boolq": 24.16666666666667,
|
| 108 |
+
"sst5": 23.166666666666664,
|
| 109 |
+
"prompt-injections": 16.37931034482759,
|
| 110 |
+
"massive-intent.en": 76.66666666666667,
|
| 111 |
+
"xnli.en": 54.666666666666664
|
| 112 |
+
},
|
| 113 |
+
"1": {
|
| 114 |
+
"typed-decisions": -1.3000000000000012,
|
| 115 |
+
"ag-news": 1.6666666666666718,
|
| 116 |
+
"emotion": -0.16666666666667052,
|
| 117 |
+
"boolq": 2.833333333333332,
|
| 118 |
+
"sst5": 2.833333333333332,
|
| 119 |
+
"prompt-injections": 6.896551724137923,
|
| 120 |
+
"massive-intent.en": 1.333333333333342,
|
| 121 |
+
"xnli.en": -0.666666666666671
|
| 122 |
+
},
|
| 123 |
+
"2": {
|
| 124 |
+
"typed-decisions": 8.399999999999997,
|
| 125 |
+
"ag-news": 68.16666666666666,
|
| 126 |
+
"emotion": 51.16666666666667,
|
| 127 |
+
"boolq": 32.0,
|
| 128 |
+
"sst5": 17.166666666666668,
|
| 129 |
+
"prompt-injections": 37.06896551724138,
|
| 130 |
+
"massive-intent.en": 68.66666666666667,
|
| 131 |
+
"xnli.en": 53.0
|
| 132 |
+
},
|
| 133 |
+
"3": {
|
| 134 |
+
"typed-decisions": 5.099999999999993,
|
| 135 |
+
"ag-news": 65.16666666666666,
|
| 136 |
+
"emotion": 42.99999999999999,
|
| 137 |
+
"boolq": 33.166666666666664,
|
| 138 |
+
"sst5": 15.500000000000004,
|
| 139 |
+
"prompt-injections": 21.55172413793103,
|
| 140 |
+
"massive-intent.en": 75.0,
|
| 141 |
+
"xnli.en": 53.0
|
| 142 |
+
},
|
| 143 |
+
"4": {
|
| 144 |
+
"typed-decisions": -0.20000000000000018,
|
| 145 |
+
"ag-news": 1.6666666666666607,
|
| 146 |
+
"emotion": 1.333333333333342,
|
| 147 |
+
"boolq": 2.0000000000000018,
|
| 148 |
+
"sst5": -6.333333333333336,
|
| 149 |
+
"prompt-injections": 1.7241379310344862,
|
| 150 |
+
"massive-intent.en": -0.33333333333332993,
|
| 151 |
+
"xnli.en": -1.333333333333342
|
| 152 |
+
}
|
| 153 |
+
},
|
| 154 |
+
"previous_release": {
|
| 155 |
+
"seed": 1,
|
| 156 |
+
"epoch": 14,
|
| 157 |
+
"interface_lr": 0.0003,
|
| 158 |
+
"suites": {
|
| 159 |
+
"typed-decisions": {
|
| 160 |
+
"attempted": 2000,
|
| 161 |
+
"valid": 2000,
|
| 162 |
+
"failed": 0,
|
| 163 |
+
"coverage": 1.0,
|
| 164 |
+
"accuracy_all": 0.7695,
|
| 165 |
+
"accuracy_valid": 0.7695,
|
| 166 |
+
"ece_top_label": 0.2142237525419539,
|
| 167 |
+
"mean_confidence": 0.555276247458046,
|
| 168 |
+
"brier_hard": 0.3977842163576973,
|
| 169 |
+
"brier_hard_n": 2000,
|
| 170 |
+
"nll_hard": 0.7027635604099817,
|
| 171 |
+
"nll_hard_n": 2000,
|
| 172 |
+
"zero_probability_gold": 0.0,
|
| 173 |
+
"zero_probability_gold_n": 2000,
|
| 174 |
+
"soft_accuracy": 0.47103605907759827,
|
| 175 |
+
"soft_accuracy_n": 2000,
|
| 176 |
+
"brier_soft": 0.06255403710877039,
|
| 177 |
+
"brier_soft_n": 2000,
|
| 178 |
+
"kl_gold_to_prediction": 0.11643450461088396,
|
| 179 |
+
"kl_gold_to_prediction_n": 2000,
|
| 180 |
+
"total_variation": 0.17146599324503523,
|
| 181 |
+
"total_variation_n": 2000,
|
| 182 |
+
"score_mae": 0.2299944493828981,
|
| 183 |
+
"score_mae_n": 800,
|
| 184 |
+
"within_one_level": 0.98875,
|
| 185 |
+
"within_one_level_n": 800,
|
| 186 |
+
"macro_f1": 0.6487623885802034
|
| 187 |
+
},
|
| 188 |
+
"ag-news": {
|
| 189 |
+
"attempted": 600,
|
| 190 |
+
"valid": 600,
|
| 191 |
+
"failed": 0,
|
| 192 |
+
"coverage": 1.0,
|
| 193 |
+
"accuracy_all": 0.9316666666666666,
|
| 194 |
+
"accuracy_valid": 0.9316666666666666,
|
| 195 |
+
"ece_top_label": 0.03387674147756575,
|
| 196 |
+
"mean_confidence": 0.904708562141314,
|
| 197 |
+
"brier_hard": 0.10911433540423408,
|
| 198 |
+
"brier_hard_n": 600,
|
| 199 |
+
"nll_hard": 0.21361121678114303,
|
| 200 |
+
"nll_hard_n": 600,
|
| 201 |
+
"zero_probability_gold": 0.0,
|
| 202 |
+
"zero_probability_gold_n": 600,
|
| 203 |
+
"macro_f1": 0.9283987145646934
|
| 204 |
+
},
|
| 205 |
+
"emotion": {
|
| 206 |
+
"attempted": 600,
|
| 207 |
+
"valid": 600,
|
| 208 |
+
"failed": 0,
|
| 209 |
+
"coverage": 1.0,
|
| 210 |
+
"accuracy_all": 0.5733333333333334,
|
| 211 |
+
"accuracy_valid": 0.5733333333333334,
|
| 212 |
+
"ece_top_label": 0.3169210441851814,
|
| 213 |
+
"mean_confidence": 0.8880587177380197,
|
| 214 |
+
"brier_hard": 0.7264584486728032,
|
| 215 |
+
"brier_hard_n": 600,
|
| 216 |
+
"nll_hard": 2.146154603203373,
|
| 217 |
+
"nll_hard_n": 600,
|
| 218 |
+
"zero_probability_gold": 0.008333333333333333,
|
| 219 |
+
"zero_probability_gold_n": 600,
|
| 220 |
+
"macro_f1": 0.4862286665693279
|
| 221 |
+
},
|
| 222 |
+
"boolq": {
|
| 223 |
+
"attempted": 600,
|
| 224 |
+
"valid": 600,
|
| 225 |
+
"failed": 0,
|
| 226 |
+
"coverage": 1.0,
|
| 227 |
+
"accuracy_all": 0.7983333333333333,
|
| 228 |
+
"accuracy_valid": 0.7983333333333333,
|
| 229 |
+
"ece_top_label": 0.09755433333333334,
|
| 230 |
+
"mean_confidence": 0.8900093333333333,
|
| 231 |
+
"brier_hard": 0.29822958146666667,
|
| 232 |
+
"brier_hard_n": 600,
|
| 233 |
+
"nll_hard": 0.4711253960780479,
|
| 234 |
+
"nll_hard_n": 600,
|
| 235 |
+
"zero_probability_gold": 0.0,
|
| 236 |
+
"zero_probability_gold_n": 600,
|
| 237 |
+
"macro_f1": 0.7813983878883868
|
| 238 |
+
},
|
| 239 |
+
"sst5": {
|
| 240 |
+
"attempted": 600,
|
| 241 |
+
"valid": 600,
|
| 242 |
+
"failed": 0,
|
| 243 |
+
"coverage": 1.0,
|
| 244 |
+
"accuracy_all": 0.42,
|
| 245 |
+
"accuracy_valid": 0.42,
|
| 246 |
+
"ece_top_label": 0.1643168667044709,
|
| 247 |
+
"mean_confidence": 0.5761336745094923,
|
| 248 |
+
"brier_hard": 0.722245716690302,
|
| 249 |
+
"brier_hard_n": 600,
|
| 250 |
+
"nll_hard": 1.412054753482227,
|
| 251 |
+
"nll_hard_n": 600,
|
| 252 |
+
"zero_probability_gold": 0.0,
|
| 253 |
+
"zero_probability_gold_n": 600,
|
| 254 |
+
"score_mae": 0.7360310433394476,
|
| 255 |
+
"score_mae_n": 600,
|
| 256 |
+
"within_one_level": 0.7466666666666667,
|
| 257 |
+
"within_one_level_n": 600,
|
| 258 |
+
"macro_f1": 0.41387920942262574
|
| 259 |
+
},
|
| 260 |
+
"prompt-injections": {
|
| 261 |
+
"attempted": 116,
|
| 262 |
+
"valid": 116,
|
| 263 |
+
"failed": 0,
|
| 264 |
+
"coverage": 1.0,
|
| 265 |
+
"accuracy_all": 0.7155172413793104,
|
| 266 |
+
"accuracy_valid": 0.7155172413793104,
|
| 267 |
+
"ece_top_label": 0.20010517241379305,
|
| 268 |
+
"mean_confidence": 0.9105,
|
| 269 |
+
"brier_hard": 0.43610344172413795,
|
| 270 |
+
"brier_hard_n": 116,
|
| 271 |
+
"nll_hard": 1.2355666268201304,
|
| 272 |
+
"nll_hard_n": 116,
|
| 273 |
+
"zero_probability_gold": 0.008620689655172414,
|
| 274 |
+
"zero_probability_gold_n": 116,
|
| 275 |
+
"macro_f1": 0.7038756091900673
|
| 276 |
+
},
|
| 277 |
+
"massive-intent.en": {
|
| 278 |
+
"attempted": 300,
|
| 279 |
+
"valid": 300,
|
| 280 |
+
"failed": 0,
|
| 281 |
+
"coverage": 1.0,
|
| 282 |
+
"accuracy_all": 0.7433333333333333,
|
| 283 |
+
"accuracy_valid": 0.7433333333333333,
|
| 284 |
+
"ece_top_label": 0.1906540683935104,
|
| 285 |
+
"mean_confidence": 0.9339874017268438,
|
| 286 |
+
"brier_hard": 0.43706654758878133,
|
| 287 |
+
"brier_hard_n": 300,
|
| 288 |
+
"nll_hard": 2.7801978001907903,
|
| 289 |
+
"nll_hard_n": 300,
|
| 290 |
+
"zero_probability_gold": 0.07,
|
| 291 |
+
"zero_probability_gold_n": 300,
|
| 292 |
+
"macro_f1": 0.44346732036749054
|
| 293 |
+
},
|
| 294 |
+
"xnli.en": {
|
| 295 |
+
"attempted": 300,
|
| 296 |
+
"valid": 300,
|
| 297 |
+
"failed": 0,
|
| 298 |
+
"coverage": 1.0,
|
| 299 |
+
"accuracy_all": 0.8766666666666667,
|
| 300 |
+
"accuracy_valid": 0.8766666666666667,
|
| 301 |
+
"ece_top_label": 0.055158283453726184,
|
| 302 |
+
"mean_confidence": 0.913468962679153,
|
| 303 |
+
"brier_hard": 0.19180470102420566,
|
| 304 |
+
"brier_hard_n": 300,
|
| 305 |
+
"nll_hard": 0.3708980570966955,
|
| 306 |
+
"nll_hard_n": 300,
|
| 307 |
+
"zero_probability_gold": 0.0,
|
| 308 |
+
"zero_probability_gold_n": 300,
|
| 309 |
+
"macro_f1": 0.8772028178860477
|
| 310 |
+
}
|
| 311 |
+
},
|
| 312 |
+
"repository_revision": "eaea15162eac93240e5d70f263920eea3093e6c2"
|
| 313 |
+
},
|
| 314 |
+
"scope": "adaptive follow-up on a repeatedly inspected benchmark bundle"
|
| 315 |
+
}
|
evidence/replay-check.json
CHANGED
|
@@ -1,15 +1,17 @@
|
|
| 1 |
{
|
| 2 |
"passed": true,
|
| 3 |
"seed": 0,
|
|
|
|
| 4 |
"separate_processes": true,
|
| 5 |
"epochs_per_process": 2,
|
| 6 |
"optimizer_updates_per_process": 338,
|
| 7 |
"initial_validation_exact": true,
|
| 8 |
"epoch_losses_and_metrics_exact": true,
|
| 9 |
"selected_checkpoint_tensors_bitwise_equal": true,
|
|
|
|
| 10 |
"checkpoint_sha256": [
|
| 11 |
-
"
|
| 12 |
-
"
|
| 13 |
],
|
| 14 |
"reproducibility": {
|
| 15 |
"deterministic_algorithms": true,
|
|
@@ -28,5 +30,6 @@
|
|
| 28 |
"cudnn_version": 92000,
|
| 29 |
"scope": "fixed hardware/runtime; cross-platform bitwise agreement is not promised"
|
| 30 |
},
|
| 31 |
-
"
|
|
|
|
| 32 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"passed": true,
|
| 3 |
"seed": 0,
|
| 4 |
+
"interface_lr": 0.0001,
|
| 5 |
"separate_processes": true,
|
| 6 |
"epochs_per_process": 2,
|
| 7 |
"optimizer_updates_per_process": 338,
|
| 8 |
"initial_validation_exact": true,
|
| 9 |
"epoch_losses_and_metrics_exact": true,
|
| 10 |
"selected_checkpoint_tensors_bitwise_equal": true,
|
| 11 |
+
"reported_seed_0_two_epoch_metrics_exact": true,
|
| 12 |
"checkpoint_sha256": [
|
| 13 |
+
"3c52574d070140a55de8929922302c193d4493f76f35a502edf419ee5005881e",
|
| 14 |
+
"3c52574d070140a55de8929922302c193d4493f76f35a502edf419ee5005881e"
|
| 15 |
],
|
| 16 |
"reproducibility": {
|
| 17 |
"deterministic_algorithms": true,
|
|
|
|
| 30 |
"cudnn_version": 92000,
|
| 31 |
"scope": "fixed hardware/runtime; cross-platform bitwise agreement is not promised"
|
| 32 |
},
|
| 33 |
+
"timing": "verification after the completed five-seed sweep",
|
| 34 |
+
"limit": "same runtime, hardware, data and two-epoch prefix; not a full-trajectory or cross-platform replay"
|
| 35 |
}
|
evidence/seed-0-history.json
CHANGED
|
@@ -9,399 +9,246 @@
|
|
| 9 |
"epochs": [
|
| 10 |
{
|
| 11 |
"epoch": 1,
|
| 12 |
-
"train_soft_cross_entropy": 1.
|
| 13 |
"validation": {
|
| 14 |
-
"soft_cross_entropy": 1.
|
| 15 |
-
"accuracy": 0.
|
| 16 |
-
"brier_soft": 0.
|
| 17 |
"decisions": 600
|
| 18 |
},
|
| 19 |
"early_stopping": {
|
| 20 |
"patience": 3,
|
| 21 |
"min_epochs": 10,
|
| 22 |
"epochs_without_improvement": 0,
|
| 23 |
-
"best_validation_loss": 1.
|
| 24 |
"improved": true
|
| 25 |
}
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"epoch": 2,
|
| 29 |
-
"train_soft_cross_entropy":
|
| 30 |
"validation": {
|
| 31 |
-
"soft_cross_entropy": 1.
|
| 32 |
-
"accuracy": 0.
|
| 33 |
-
"brier_soft": 0.
|
| 34 |
"decisions": 600
|
| 35 |
},
|
| 36 |
"early_stopping": {
|
| 37 |
"patience": 3,
|
| 38 |
"min_epochs": 10,
|
| 39 |
"epochs_without_improvement": 0,
|
| 40 |
-
"best_validation_loss": 1.
|
| 41 |
"improved": true
|
| 42 |
}
|
| 43 |
},
|
| 44 |
{
|
| 45 |
"epoch": 3,
|
| 46 |
-
"train_soft_cross_entropy":
|
| 47 |
"validation": {
|
| 48 |
-
"soft_cross_entropy":
|
| 49 |
-
"accuracy": 0.
|
| 50 |
-
"brier_soft": 0.
|
| 51 |
"decisions": 600
|
| 52 |
},
|
| 53 |
"early_stopping": {
|
| 54 |
"patience": 3,
|
| 55 |
"min_epochs": 10,
|
| 56 |
"epochs_without_improvement": 0,
|
| 57 |
-
"best_validation_loss":
|
| 58 |
"improved": true
|
| 59 |
}
|
| 60 |
},
|
| 61 |
{
|
| 62 |
"epoch": 4,
|
| 63 |
-
"train_soft_cross_entropy":
|
| 64 |
"validation": {
|
| 65 |
-
"soft_cross_entropy": 1.
|
| 66 |
-
"accuracy": 0.
|
| 67 |
-
"brier_soft": 0.
|
| 68 |
"decisions": 600
|
| 69 |
},
|
| 70 |
"early_stopping": {
|
| 71 |
"patience": 3,
|
| 72 |
"min_epochs": 10,
|
| 73 |
"epochs_without_improvement": 1,
|
| 74 |
-
"best_validation_loss":
|
| 75 |
"improved": false
|
| 76 |
}
|
| 77 |
},
|
| 78 |
{
|
| 79 |
"epoch": 5,
|
| 80 |
-
"train_soft_cross_entropy":
|
| 81 |
"validation": {
|
| 82 |
-
"soft_cross_entropy":
|
| 83 |
-
"accuracy": 0.
|
| 84 |
-
"brier_soft": 0.
|
| 85 |
"decisions": 600
|
| 86 |
},
|
| 87 |
"early_stopping": {
|
| 88 |
"patience": 3,
|
| 89 |
"min_epochs": 10,
|
| 90 |
-
"epochs_without_improvement":
|
| 91 |
-
"best_validation_loss":
|
| 92 |
-
"improved":
|
| 93 |
}
|
| 94 |
},
|
| 95 |
{
|
| 96 |
"epoch": 6,
|
| 97 |
-
"train_soft_cross_entropy":
|
| 98 |
"validation": {
|
| 99 |
-
"soft_cross_entropy":
|
| 100 |
-
"accuracy": 0.
|
| 101 |
-
"brier_soft": 0.
|
| 102 |
"decisions": 600
|
| 103 |
},
|
| 104 |
"early_stopping": {
|
| 105 |
"patience": 3,
|
| 106 |
"min_epochs": 10,
|
| 107 |
"epochs_without_improvement": 0,
|
| 108 |
-
"best_validation_loss":
|
| 109 |
"improved": true
|
| 110 |
}
|
| 111 |
},
|
| 112 |
{
|
| 113 |
"epoch": 7,
|
| 114 |
-
"train_soft_cross_entropy":
|
| 115 |
"validation": {
|
| 116 |
-
"soft_cross_entropy":
|
| 117 |
-
"accuracy": 0.
|
| 118 |
-
"brier_soft": 0.
|
| 119 |
"decisions": 600
|
| 120 |
},
|
| 121 |
"early_stopping": {
|
| 122 |
"patience": 3,
|
| 123 |
"min_epochs": 10,
|
| 124 |
"epochs_without_improvement": 0,
|
| 125 |
-
"best_validation_loss":
|
| 126 |
"improved": true
|
| 127 |
}
|
| 128 |
},
|
| 129 |
{
|
| 130 |
"epoch": 8,
|
| 131 |
-
"train_soft_cross_entropy":
|
| 132 |
-
"validation": {
|
| 133 |
-
"soft_cross_entropy": 1.0927943936983744,
|
| 134 |
-
"accuracy": 0.4866666666666667,
|
| 135 |
-
"brier_soft": 0.20212342927853266,
|
| 136 |
-
"decisions": 600
|
| 137 |
-
},
|
| 138 |
-
"early_stopping": {
|
| 139 |
-
"patience": 3,
|
| 140 |
-
"min_epochs": 10,
|
| 141 |
-
"epochs_without_improvement": 0,
|
| 142 |
-
"best_validation_loss": 1.0927943936983744,
|
| 143 |
-
"improved": true
|
| 144 |
-
}
|
| 145 |
-
},
|
| 146 |
-
{
|
| 147 |
-
"epoch": 9,
|
| 148 |
-
"train_soft_cross_entropy": 1.0959268622045164,
|
| 149 |
"validation": {
|
| 150 |
-
"soft_cross_entropy":
|
| 151 |
-
"accuracy": 0.
|
| 152 |
-
"brier_soft": 0.
|
| 153 |
-
"decisions": 600
|
| 154 |
-
},
|
| 155 |
-
"early_stopping": {
|
| 156 |
-
"patience": 3,
|
| 157 |
-
"min_epochs": 10,
|
| 158 |
-
"epochs_without_improvement": 0,
|
| 159 |
-
"best_validation_loss": 1.0886747964223227,
|
| 160 |
-
"improved": true
|
| 161 |
-
}
|
| 162 |
-
},
|
| 163 |
-
{
|
| 164 |
-
"epoch": 10,
|
| 165 |
-
"train_soft_cross_entropy": 1.0879168816849036,
|
| 166 |
-
"validation": {
|
| 167 |
-
"soft_cross_entropy": 1.0776302735010783,
|
| 168 |
-
"accuracy": 0.5133333333333333,
|
| 169 |
-
"brier_soft": 0.19176260660092037,
|
| 170 |
-
"decisions": 600
|
| 171 |
-
},
|
| 172 |
-
"early_stopping": {
|
| 173 |
-
"patience": 3,
|
| 174 |
-
"min_epochs": 10,
|
| 175 |
-
"epochs_without_improvement": 0,
|
| 176 |
-
"best_validation_loss": 1.0776302735010783,
|
| 177 |
-
"improved": true
|
| 178 |
-
}
|
| 179 |
-
},
|
| 180 |
-
{
|
| 181 |
-
"epoch": 11,
|
| 182 |
-
"train_soft_cross_entropy": 1.0794475747037817,
|
| 183 |
-
"validation": {
|
| 184 |
-
"soft_cross_entropy": 1.0989113728205362,
|
| 185 |
-
"accuracy": 0.49666666666666665,
|
| 186 |
-
"brier_soft": 0.20668872609734534,
|
| 187 |
"decisions": 600
|
| 188 |
},
|
| 189 |
"early_stopping": {
|
| 190 |
"patience": 3,
|
| 191 |
"min_epochs": 10,
|
| 192 |
"epochs_without_improvement": 1,
|
| 193 |
-
"best_validation_loss":
|
| 194 |
"improved": false
|
| 195 |
}
|
| 196 |
},
|
| 197 |
{
|
| 198 |
-
"epoch":
|
| 199 |
-
"train_soft_cross_entropy":
|
| 200 |
-
"validation": {
|
| 201 |
-
"soft_cross_entropy": 1.0758668931325277,
|
| 202 |
-
"accuracy": 0.5016666666666667,
|
| 203 |
-
"brier_soft": 0.19317058285077413,
|
| 204 |
-
"decisions": 600
|
| 205 |
-
},
|
| 206 |
-
"early_stopping": {
|
| 207 |
-
"patience": 3,
|
| 208 |
-
"min_epochs": 10,
|
| 209 |
-
"epochs_without_improvement": 0,
|
| 210 |
-
"best_validation_loss": 1.0758668931325277,
|
| 211 |
-
"improved": true
|
| 212 |
-
}
|
| 213 |
-
},
|
| 214 |
-
{
|
| 215 |
-
"epoch": 13,
|
| 216 |
-
"train_soft_cross_entropy": 1.052984880871243,
|
| 217 |
-
"validation": {
|
| 218 |
-
"soft_cross_entropy": 1.0600487383206685,
|
| 219 |
-
"accuracy": 0.5316666666666666,
|
| 220 |
-
"brier_soft": 0.18257314254840215,
|
| 221 |
-
"decisions": 600
|
| 222 |
-
},
|
| 223 |
-
"early_stopping": {
|
| 224 |
-
"patience": 3,
|
| 225 |
-
"min_epochs": 10,
|
| 226 |
-
"epochs_without_improvement": 0,
|
| 227 |
-
"best_validation_loss": 1.0600487383206685,
|
| 228 |
-
"improved": true
|
| 229 |
-
}
|
| 230 |
-
},
|
| 231 |
-
{
|
| 232 |
-
"epoch": 14,
|
| 233 |
-
"train_soft_cross_entropy": 1.0401919462062694,
|
| 234 |
-
"validation": {
|
| 235 |
-
"soft_cross_entropy": 1.0525298221906025,
|
| 236 |
-
"accuracy": 0.5366666666666666,
|
| 237 |
-
"brier_soft": 0.1811353324353695,
|
| 238 |
-
"decisions": 600
|
| 239 |
-
},
|
| 240 |
-
"early_stopping": {
|
| 241 |
-
"patience": 3,
|
| 242 |
-
"min_epochs": 10,
|
| 243 |
-
"epochs_without_improvement": 0,
|
| 244 |
-
"best_validation_loss": 1.0525298221906025,
|
| 245 |
-
"improved": true
|
| 246 |
-
}
|
| 247 |
-
},
|
| 248 |
-
{
|
| 249 |
-
"epoch": 15,
|
| 250 |
-
"train_soft_cross_entropy": 1.0262903751267327,
|
| 251 |
-
"validation": {
|
| 252 |
-
"soft_cross_entropy": 1.0581993921597799,
|
| 253 |
-
"accuracy": 0.5233333333333333,
|
| 254 |
-
"brier_soft": 0.18550985043247542,
|
| 255 |
-
"decisions": 600
|
| 256 |
-
},
|
| 257 |
-
"early_stopping": {
|
| 258 |
-
"patience": 3,
|
| 259 |
-
"min_epochs": 10,
|
| 260 |
-
"epochs_without_improvement": 1,
|
| 261 |
-
"best_validation_loss": 1.0525298221906025,
|
| 262 |
-
"improved": false
|
| 263 |
-
}
|
| 264 |
-
},
|
| 265 |
-
{
|
| 266 |
-
"epoch": 16,
|
| 267 |
-
"train_soft_cross_entropy": 1.0127459985238534,
|
| 268 |
-
"validation": {
|
| 269 |
-
"soft_cross_entropy": 1.0545752588907877,
|
| 270 |
-
"accuracy": 0.54,
|
| 271 |
-
"brier_soft": 0.18113683501879374,
|
| 272 |
-
"decisions": 600
|
| 273 |
-
},
|
| 274 |
-
"early_stopping": {
|
| 275 |
-
"patience": 3,
|
| 276 |
-
"min_epochs": 10,
|
| 277 |
-
"epochs_without_improvement": 2,
|
| 278 |
-
"best_validation_loss": 1.0525298221906025,
|
| 279 |
-
"improved": false
|
| 280 |
-
}
|
| 281 |
-
},
|
| 282 |
-
{
|
| 283 |
-
"epoch": 17,
|
| 284 |
-
"train_soft_cross_entropy": 1.0096984642523306,
|
| 285 |
-
"validation": {
|
| 286 |
-
"soft_cross_entropy": 1.0468252456188203,
|
| 287 |
-
"accuracy": 0.5483333333333333,
|
| 288 |
-
"brier_soft": 0.17509076982736588,
|
| 289 |
-
"decisions": 600
|
| 290 |
-
},
|
| 291 |
-
"early_stopping": {
|
| 292 |
-
"patience": 3,
|
| 293 |
-
"min_epochs": 10,
|
| 294 |
-
"epochs_without_improvement": 0,
|
| 295 |
-
"best_validation_loss": 1.0468252456188203,
|
| 296 |
-
"improved": true
|
| 297 |
-
}
|
| 298 |
-
},
|
| 299 |
-
{
|
| 300 |
-
"epoch": 18,
|
| 301 |
-
"train_soft_cross_entropy": 1.0022709012914588,
|
| 302 |
"validation": {
|
| 303 |
-
"soft_cross_entropy":
|
| 304 |
-
"accuracy": 0.
|
| 305 |
-
"brier_soft": 0.
|
| 306 |
"decisions": 600
|
| 307 |
},
|
| 308 |
"early_stopping": {
|
| 309 |
"patience": 3,
|
| 310 |
"min_epochs": 10,
|
| 311 |
"epochs_without_improvement": 0,
|
| 312 |
-
"best_validation_loss":
|
| 313 |
"improved": true
|
| 314 |
}
|
| 315 |
},
|
| 316 |
{
|
| 317 |
-
"epoch":
|
| 318 |
-
"train_soft_cross_entropy": 0.
|
| 319 |
"validation": {
|
| 320 |
-
"soft_cross_entropy":
|
| 321 |
-
"accuracy": 0.
|
| 322 |
-
"brier_soft": 0.
|
| 323 |
"decisions": 600
|
| 324 |
},
|
| 325 |
"early_stopping": {
|
| 326 |
"patience": 3,
|
| 327 |
"min_epochs": 10,
|
| 328 |
"epochs_without_improvement": 1,
|
| 329 |
-
"best_validation_loss":
|
| 330 |
"improved": false
|
| 331 |
}
|
| 332 |
},
|
| 333 |
{
|
| 334 |
-
"epoch":
|
| 335 |
-
"train_soft_cross_entropy": 0.
|
| 336 |
"validation": {
|
| 337 |
-
"soft_cross_entropy":
|
| 338 |
-
"accuracy": 0.
|
| 339 |
-
"brier_soft": 0.
|
| 340 |
"decisions": 600
|
| 341 |
},
|
| 342 |
"early_stopping": {
|
| 343 |
"patience": 3,
|
| 344 |
"min_epochs": 10,
|
| 345 |
"epochs_without_improvement": 0,
|
| 346 |
-
"best_validation_loss":
|
| 347 |
"improved": true
|
| 348 |
}
|
| 349 |
},
|
| 350 |
{
|
| 351 |
-
"epoch":
|
| 352 |
-
"train_soft_cross_entropy": 0.
|
| 353 |
"validation": {
|
| 354 |
-
"soft_cross_entropy":
|
| 355 |
-
"accuracy": 0.
|
| 356 |
-
"brier_soft": 0.
|
| 357 |
"decisions": 600
|
| 358 |
},
|
| 359 |
"early_stopping": {
|
| 360 |
"patience": 3,
|
| 361 |
"min_epochs": 10,
|
| 362 |
"epochs_without_improvement": 1,
|
| 363 |
-
"best_validation_loss":
|
| 364 |
"improved": false
|
| 365 |
}
|
| 366 |
},
|
| 367 |
{
|
| 368 |
-
"epoch":
|
| 369 |
-
"train_soft_cross_entropy": 0.
|
| 370 |
"validation": {
|
| 371 |
-
"soft_cross_entropy":
|
| 372 |
-
"accuracy": 0.
|
| 373 |
-
"brier_soft": 0.
|
| 374 |
"decisions": 600
|
| 375 |
},
|
| 376 |
"early_stopping": {
|
| 377 |
"patience": 3,
|
| 378 |
"min_epochs": 10,
|
| 379 |
"epochs_without_improvement": 2,
|
| 380 |
-
"best_validation_loss":
|
| 381 |
"improved": false
|
| 382 |
}
|
| 383 |
},
|
| 384 |
{
|
| 385 |
-
"epoch":
|
| 386 |
-
"train_soft_cross_entropy": 0.
|
| 387 |
"validation": {
|
| 388 |
-
"soft_cross_entropy":
|
| 389 |
-
"accuracy": 0.
|
| 390 |
-
"brier_soft": 0.
|
| 391 |
"decisions": 600
|
| 392 |
},
|
| 393 |
"early_stopping": {
|
| 394 |
"patience": 3,
|
| 395 |
"min_epochs": 10,
|
| 396 |
"epochs_without_improvement": 3,
|
| 397 |
-
"best_validation_loss":
|
| 398 |
"improved": false
|
| 399 |
}
|
| 400 |
}
|
| 401 |
],
|
| 402 |
"stopping": {
|
| 403 |
"reason": "early_stopping",
|
| 404 |
-
"epochs_completed":
|
| 405 |
"patience": 3,
|
| 406 |
"min_epochs": 10,
|
| 407 |
"epochs_without_improvement": 3,
|
|
|
|
| 9 |
"epochs": [
|
| 10 |
{
|
| 11 |
"epoch": 1,
|
| 12 |
+
"train_soft_cross_entropy": 1.087307652897305,
|
| 13 |
"validation": {
|
| 14 |
+
"soft_cross_entropy": 1.0358176565170287,
|
| 15 |
+
"accuracy": 0.5516666666666666,
|
| 16 |
+
"brier_soft": 0.17500559210777283,
|
| 17 |
"decisions": 600
|
| 18 |
},
|
| 19 |
"early_stopping": {
|
| 20 |
"patience": 3,
|
| 21 |
"min_epochs": 10,
|
| 22 |
"epochs_without_improvement": 0,
|
| 23 |
+
"best_validation_loss": 1.0358176565170287,
|
| 24 |
"improved": true
|
| 25 |
}
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"epoch": 2,
|
| 29 |
+
"train_soft_cross_entropy": 0.9998172007666694,
|
| 30 |
"validation": {
|
| 31 |
+
"soft_cross_entropy": 1.013402551015218,
|
| 32 |
+
"accuracy": 0.6066666666666667,
|
| 33 |
+
"brier_soft": 0.16418362177908422,
|
| 34 |
"decisions": 600
|
| 35 |
},
|
| 36 |
"early_stopping": {
|
| 37 |
"patience": 3,
|
| 38 |
"min_epochs": 10,
|
| 39 |
"epochs_without_improvement": 0,
|
| 40 |
+
"best_validation_loss": 1.013402551015218,
|
| 41 |
"improved": true
|
| 42 |
}
|
| 43 |
},
|
| 44 |
{
|
| 45 |
"epoch": 3,
|
| 46 |
+
"train_soft_cross_entropy": 0.961990624445456,
|
| 47 |
"validation": {
|
| 48 |
+
"soft_cross_entropy": 0.9978651634852092,
|
| 49 |
+
"accuracy": 0.6116666666666667,
|
| 50 |
+
"brier_soft": 0.15542398323615392,
|
| 51 |
"decisions": 600
|
| 52 |
},
|
| 53 |
"early_stopping": {
|
| 54 |
"patience": 3,
|
| 55 |
"min_epochs": 10,
|
| 56 |
"epochs_without_improvement": 0,
|
| 57 |
+
"best_validation_loss": 0.9978651634852092,
|
| 58 |
"improved": true
|
| 59 |
}
|
| 60 |
},
|
| 61 |
{
|
| 62 |
"epoch": 4,
|
| 63 |
+
"train_soft_cross_entropy": 0.9616141191235295,
|
| 64 |
"validation": {
|
| 65 |
+
"soft_cross_entropy": 1.1574800237019858,
|
| 66 |
+
"accuracy": 0.48,
|
| 67 |
+
"brier_soft": 0.23033222297827402,
|
| 68 |
"decisions": 600
|
| 69 |
},
|
| 70 |
"early_stopping": {
|
| 71 |
"patience": 3,
|
| 72 |
"min_epochs": 10,
|
| 73 |
"epochs_without_improvement": 1,
|
| 74 |
+
"best_validation_loss": 0.9978651634852092,
|
| 75 |
"improved": false
|
| 76 |
}
|
| 77 |
},
|
| 78 |
{
|
| 79 |
"epoch": 5,
|
| 80 |
+
"train_soft_cross_entropy": 0.9278180778468097,
|
| 81 |
"validation": {
|
| 82 |
+
"soft_cross_entropy": 0.9557415008544922,
|
| 83 |
+
"accuracy": 0.6833333333333333,
|
| 84 |
+
"brier_soft": 0.12863569288204113,
|
| 85 |
"decisions": 600
|
| 86 |
},
|
| 87 |
"early_stopping": {
|
| 88 |
"patience": 3,
|
| 89 |
"min_epochs": 10,
|
| 90 |
+
"epochs_without_improvement": 0,
|
| 91 |
+
"best_validation_loss": 0.9557415008544922,
|
| 92 |
+
"improved": true
|
| 93 |
}
|
| 94 |
},
|
| 95 |
{
|
| 96 |
"epoch": 6,
|
| 97 |
+
"train_soft_cross_entropy": 0.8944051215383741,
|
| 98 |
"validation": {
|
| 99 |
+
"soft_cross_entropy": 0.9157323372364045,
|
| 100 |
+
"accuracy": 0.6916666666666667,
|
| 101 |
+
"brier_soft": 0.1038101188838482,
|
| 102 |
"decisions": 600
|
| 103 |
},
|
| 104 |
"early_stopping": {
|
| 105 |
"patience": 3,
|
| 106 |
"min_epochs": 10,
|
| 107 |
"epochs_without_improvement": 0,
|
| 108 |
+
"best_validation_loss": 0.9157323372364045,
|
| 109 |
"improved": true
|
| 110 |
}
|
| 111 |
},
|
| 112 |
{
|
| 113 |
"epoch": 7,
|
| 114 |
+
"train_soft_cross_entropy": 0.858607970961818,
|
| 115 |
"validation": {
|
| 116 |
+
"soft_cross_entropy": 0.9070962381362915,
|
| 117 |
+
"accuracy": 0.7016666666666667,
|
| 118 |
+
"brier_soft": 0.09971169379850228,
|
| 119 |
"decisions": 600
|
| 120 |
},
|
| 121 |
"early_stopping": {
|
| 122 |
"patience": 3,
|
| 123 |
"min_epochs": 10,
|
| 124 |
"epochs_without_improvement": 0,
|
| 125 |
+
"best_validation_loss": 0.9070962381362915,
|
| 126 |
"improved": true
|
| 127 |
}
|
| 128 |
},
|
| 129 |
{
|
| 130 |
"epoch": 8,
|
| 131 |
+
"train_soft_cross_entropy": 0.8490360851199539,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 132 |
"validation": {
|
| 133 |
+
"soft_cross_entropy": 0.9083957560857137,
|
| 134 |
+
"accuracy": 0.7116666666666667,
|
| 135 |
+
"brier_soft": 0.10123874122897784,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 136 |
"decisions": 600
|
| 137 |
},
|
| 138 |
"early_stopping": {
|
| 139 |
"patience": 3,
|
| 140 |
"min_epochs": 10,
|
| 141 |
"epochs_without_improvement": 1,
|
| 142 |
+
"best_validation_loss": 0.9070962381362915,
|
| 143 |
"improved": false
|
| 144 |
}
|
| 145 |
},
|
| 146 |
{
|
| 147 |
+
"epoch": 9,
|
| 148 |
+
"train_soft_cross_entropy": 0.8271938294393045,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 149 |
"validation": {
|
| 150 |
+
"soft_cross_entropy": 0.8859062685569128,
|
| 151 |
+
"accuracy": 0.7233333333333334,
|
| 152 |
+
"brier_soft": 0.08866231886049111,
|
| 153 |
"decisions": 600
|
| 154 |
},
|
| 155 |
"early_stopping": {
|
| 156 |
"patience": 3,
|
| 157 |
"min_epochs": 10,
|
| 158 |
"epochs_without_improvement": 0,
|
| 159 |
+
"best_validation_loss": 0.8859062685569128,
|
| 160 |
"improved": true
|
| 161 |
}
|
| 162 |
},
|
| 163 |
{
|
| 164 |
+
"epoch": 10,
|
| 165 |
+
"train_soft_cross_entropy": 0.8117916690861737,
|
| 166 |
"validation": {
|
| 167 |
+
"soft_cross_entropy": 0.8980106647809346,
|
| 168 |
+
"accuracy": 0.72,
|
| 169 |
+
"brier_soft": 0.09496299955993891,
|
| 170 |
"decisions": 600
|
| 171 |
},
|
| 172 |
"early_stopping": {
|
| 173 |
"patience": 3,
|
| 174 |
"min_epochs": 10,
|
| 175 |
"epochs_without_improvement": 1,
|
| 176 |
+
"best_validation_loss": 0.8859062685569128,
|
| 177 |
"improved": false
|
| 178 |
}
|
| 179 |
},
|
| 180 |
{
|
| 181 |
+
"epoch": 11,
|
| 182 |
+
"train_soft_cross_entropy": 0.8015254776124601,
|
| 183 |
"validation": {
|
| 184 |
+
"soft_cross_entropy": 0.8771124110619227,
|
| 185 |
+
"accuracy": 0.7116666666666667,
|
| 186 |
+
"brier_soft": 0.0828241604194045,
|
| 187 |
"decisions": 600
|
| 188 |
},
|
| 189 |
"early_stopping": {
|
| 190 |
"patience": 3,
|
| 191 |
"min_epochs": 10,
|
| 192 |
"epochs_without_improvement": 0,
|
| 193 |
+
"best_validation_loss": 0.8771124110619227,
|
| 194 |
"improved": true
|
| 195 |
}
|
| 196 |
},
|
| 197 |
{
|
| 198 |
+
"epoch": 12,
|
| 199 |
+
"train_soft_cross_entropy": 0.7930714415620874,
|
| 200 |
"validation": {
|
| 201 |
+
"soft_cross_entropy": 0.8903550871213277,
|
| 202 |
+
"accuracy": 0.7233333333333334,
|
| 203 |
+
"brier_soft": 0.09005908486122886,
|
| 204 |
"decisions": 600
|
| 205 |
},
|
| 206 |
"early_stopping": {
|
| 207 |
"patience": 3,
|
| 208 |
"min_epochs": 10,
|
| 209 |
"epochs_without_improvement": 1,
|
| 210 |
+
"best_validation_loss": 0.8771124110619227,
|
| 211 |
"improved": false
|
| 212 |
}
|
| 213 |
},
|
| 214 |
{
|
| 215 |
+
"epoch": 13,
|
| 216 |
+
"train_soft_cross_entropy": 0.7864397717405248,
|
| 217 |
"validation": {
|
| 218 |
+
"soft_cross_entropy": 0.8776129017273585,
|
| 219 |
+
"accuracy": 0.7283333333333334,
|
| 220 |
+
"brier_soft": 0.08296463950847587,
|
| 221 |
"decisions": 600
|
| 222 |
},
|
| 223 |
"early_stopping": {
|
| 224 |
"patience": 3,
|
| 225 |
"min_epochs": 10,
|
| 226 |
"epochs_without_improvement": 2,
|
| 227 |
+
"best_validation_loss": 0.8771124110619227,
|
| 228 |
"improved": false
|
| 229 |
}
|
| 230 |
},
|
| 231 |
{
|
| 232 |
+
"epoch": 14,
|
| 233 |
+
"train_soft_cross_entropy": 0.7832702258339634,
|
| 234 |
"validation": {
|
| 235 |
+
"soft_cross_entropy": 0.8776572781801224,
|
| 236 |
+
"accuracy": 0.7466666666666667,
|
| 237 |
+
"brier_soft": 0.08209236452666421,
|
| 238 |
"decisions": 600
|
| 239 |
},
|
| 240 |
"early_stopping": {
|
| 241 |
"patience": 3,
|
| 242 |
"min_epochs": 10,
|
| 243 |
"epochs_without_improvement": 3,
|
| 244 |
+
"best_validation_loss": 0.8771124110619227,
|
| 245 |
"improved": false
|
| 246 |
}
|
| 247 |
}
|
| 248 |
],
|
| 249 |
"stopping": {
|
| 250 |
"reason": "early_stopping",
|
| 251 |
+
"epochs_completed": 14,
|
| 252 |
"patience": 3,
|
| 253 |
"min_epochs": 10,
|
| 254 |
"epochs_without_improvement": 3,
|
evidence/seed-1-history.json
CHANGED
|
@@ -9,297 +9,212 @@
|
|
| 9 |
"epochs": [
|
| 10 |
{
|
| 11 |
"epoch": 1,
|
| 12 |
-
"train_soft_cross_entropy": 1.
|
| 13 |
"validation": {
|
| 14 |
-
"soft_cross_entropy": 0.
|
| 15 |
-
"accuracy": 0.
|
| 16 |
-
"brier_soft": 0.
|
| 17 |
"decisions": 600
|
| 18 |
},
|
| 19 |
"early_stopping": {
|
| 20 |
"patience": 3,
|
| 21 |
"min_epochs": 10,
|
| 22 |
"epochs_without_improvement": 0,
|
| 23 |
-
"best_validation_loss": 0.
|
| 24 |
"improved": true
|
| 25 |
}
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"epoch": 2,
|
| 29 |
-
"train_soft_cross_entropy": 0.
|
| 30 |
"validation": {
|
| 31 |
-
"soft_cross_entropy": 0.
|
| 32 |
-
"accuracy": 0.
|
| 33 |
-
"brier_soft": 0.
|
| 34 |
"decisions": 600
|
| 35 |
},
|
| 36 |
"early_stopping": {
|
| 37 |
"patience": 3,
|
| 38 |
"min_epochs": 10,
|
| 39 |
"epochs_without_improvement": 0,
|
| 40 |
-
"best_validation_loss": 0.
|
| 41 |
"improved": true
|
| 42 |
}
|
| 43 |
},
|
| 44 |
{
|
| 45 |
"epoch": 3,
|
| 46 |
-
"train_soft_cross_entropy": 0.
|
| 47 |
"validation": {
|
| 48 |
-
"soft_cross_entropy": 0.
|
| 49 |
-
"accuracy": 0.
|
| 50 |
-
"brier_soft": 0.
|
| 51 |
"decisions": 600
|
| 52 |
},
|
| 53 |
"early_stopping": {
|
| 54 |
"patience": 3,
|
| 55 |
"min_epochs": 10,
|
| 56 |
"epochs_without_improvement": 0,
|
| 57 |
-
"best_validation_loss": 0.
|
| 58 |
"improved": true
|
| 59 |
}
|
| 60 |
},
|
| 61 |
{
|
| 62 |
"epoch": 4,
|
| 63 |
-
"train_soft_cross_entropy": 0.
|
| 64 |
"validation": {
|
| 65 |
-
"soft_cross_entropy": 0.
|
| 66 |
-
"accuracy": 0.
|
| 67 |
-
"brier_soft": 0.
|
| 68 |
"decisions": 600
|
| 69 |
},
|
| 70 |
"early_stopping": {
|
| 71 |
"patience": 3,
|
| 72 |
"min_epochs": 10,
|
| 73 |
"epochs_without_improvement": 0,
|
| 74 |
-
"best_validation_loss": 0.
|
| 75 |
"improved": true
|
| 76 |
}
|
| 77 |
},
|
| 78 |
{
|
| 79 |
"epoch": 5,
|
| 80 |
-
"train_soft_cross_entropy": 0.
|
| 81 |
"validation": {
|
| 82 |
-
"soft_cross_entropy": 0.
|
| 83 |
-
"accuracy": 0.
|
| 84 |
-
"brier_soft": 0.
|
| 85 |
"decisions": 600
|
| 86 |
},
|
| 87 |
"early_stopping": {
|
| 88 |
"patience": 3,
|
| 89 |
"min_epochs": 10,
|
| 90 |
-
"epochs_without_improvement":
|
| 91 |
-
"best_validation_loss": 0.
|
| 92 |
-
"improved":
|
| 93 |
}
|
| 94 |
},
|
| 95 |
{
|
| 96 |
"epoch": 6,
|
| 97 |
-
"train_soft_cross_entropy": 0.
|
| 98 |
"validation": {
|
| 99 |
-
"soft_cross_entropy": 0.
|
| 100 |
-
"accuracy": 0.
|
| 101 |
-
"brier_soft": 0.
|
| 102 |
"decisions": 600
|
| 103 |
},
|
| 104 |
"early_stopping": {
|
| 105 |
"patience": 3,
|
| 106 |
"min_epochs": 10,
|
| 107 |
-
"epochs_without_improvement":
|
| 108 |
-
"best_validation_loss": 0.
|
| 109 |
"improved": false
|
| 110 |
}
|
| 111 |
},
|
| 112 |
{
|
| 113 |
"epoch": 7,
|
| 114 |
-
"train_soft_cross_entropy": 0.
|
| 115 |
"validation": {
|
| 116 |
-
"soft_cross_entropy": 0.
|
| 117 |
"accuracy": 0.7466666666666667,
|
| 118 |
-
"brier_soft": 0.
|
| 119 |
"decisions": 600
|
| 120 |
},
|
| 121 |
"early_stopping": {
|
| 122 |
"patience": 3,
|
| 123 |
"min_epochs": 10,
|
| 124 |
-
"epochs_without_improvement":
|
| 125 |
-
"best_validation_loss": 0.
|
| 126 |
-
"improved":
|
| 127 |
}
|
| 128 |
},
|
| 129 |
{
|
| 130 |
"epoch": 8,
|
| 131 |
-
"train_soft_cross_entropy": 0.
|
| 132 |
"validation": {
|
| 133 |
-
"soft_cross_entropy": 0.
|
| 134 |
-
"accuracy": 0.
|
| 135 |
-
"brier_soft": 0.
|
| 136 |
"decisions": 600
|
| 137 |
},
|
| 138 |
"early_stopping": {
|
| 139 |
"patience": 3,
|
| 140 |
"min_epochs": 10,
|
| 141 |
-
"epochs_without_improvement":
|
| 142 |
-
"best_validation_loss": 0.
|
| 143 |
-
"improved":
|
| 144 |
}
|
| 145 |
},
|
| 146 |
{
|
| 147 |
"epoch": 9,
|
| 148 |
-
"train_soft_cross_entropy": 0.
|
| 149 |
"validation": {
|
| 150 |
-
"soft_cross_entropy": 0.
|
| 151 |
-
"accuracy": 0.
|
| 152 |
-
"brier_soft": 0.
|
| 153 |
"decisions": 600
|
| 154 |
},
|
| 155 |
"early_stopping": {
|
| 156 |
"patience": 3,
|
| 157 |
"min_epochs": 10,
|
| 158 |
"epochs_without_improvement": 0,
|
| 159 |
-
"best_validation_loss": 0.
|
| 160 |
"improved": true
|
| 161 |
}
|
| 162 |
},
|
| 163 |
{
|
| 164 |
"epoch": 10,
|
| 165 |
-
"train_soft_cross_entropy": 0.
|
| 166 |
"validation": {
|
| 167 |
-
"soft_cross_entropy": 0.
|
| 168 |
-
"accuracy": 0.
|
| 169 |
-
"brier_soft": 0.
|
| 170 |
"decisions": 600
|
| 171 |
},
|
| 172 |
"early_stopping": {
|
| 173 |
"patience": 3,
|
| 174 |
"min_epochs": 10,
|
| 175 |
"epochs_without_improvement": 1,
|
| 176 |
-
"best_validation_loss": 0.
|
| 177 |
"improved": false
|
| 178 |
}
|
| 179 |
},
|
| 180 |
{
|
| 181 |
"epoch": 11,
|
| 182 |
-
"train_soft_cross_entropy": 0.
|
| 183 |
"validation": {
|
| 184 |
-
"soft_cross_entropy": 0.
|
| 185 |
-
"accuracy": 0.
|
| 186 |
-
"brier_soft": 0.
|
| 187 |
"decisions": 600
|
| 188 |
},
|
| 189 |
"early_stopping": {
|
| 190 |
"patience": 3,
|
| 191 |
"min_epochs": 10,
|
| 192 |
"epochs_without_improvement": 2,
|
| 193 |
-
"best_validation_loss": 0.
|
| 194 |
"improved": false
|
| 195 |
}
|
| 196 |
},
|
| 197 |
{
|
| 198 |
"epoch": 12,
|
| 199 |
-
"train_soft_cross_entropy": 0.
|
| 200 |
-
"validation": {
|
| 201 |
-
"soft_cross_entropy": 0.842797059615453,
|
| 202 |
-
"accuracy": 0.7766666666666666,
|
| 203 |
-
"brier_soft": 0.06609537469533583,
|
| 204 |
-
"decisions": 600
|
| 205 |
-
},
|
| 206 |
-
"early_stopping": {
|
| 207 |
-
"patience": 3,
|
| 208 |
-
"min_epochs": 10,
|
| 209 |
-
"epochs_without_improvement": 0,
|
| 210 |
-
"best_validation_loss": 0.842797059615453,
|
| 211 |
-
"improved": true
|
| 212 |
-
}
|
| 213 |
-
},
|
| 214 |
-
{
|
| 215 |
-
"epoch": 13,
|
| 216 |
-
"train_soft_cross_entropy": 0.7722371352601934,
|
| 217 |
-
"validation": {
|
| 218 |
-
"soft_cross_entropy": 0.848033101161321,
|
| 219 |
-
"accuracy": 0.7633333333333333,
|
| 220 |
-
"brier_soft": 0.06775808438037832,
|
| 221 |
-
"decisions": 600
|
| 222 |
-
},
|
| 223 |
-
"early_stopping": {
|
| 224 |
-
"patience": 3,
|
| 225 |
-
"min_epochs": 10,
|
| 226 |
-
"epochs_without_improvement": 1,
|
| 227 |
-
"best_validation_loss": 0.842797059615453,
|
| 228 |
-
"improved": false
|
| 229 |
-
}
|
| 230 |
-
},
|
| 231 |
-
{
|
| 232 |
-
"epoch": 14,
|
| 233 |
-
"train_soft_cross_entropy": 0.7689926949695305,
|
| 234 |
-
"validation": {
|
| 235 |
-
"soft_cross_entropy": 0.8377540612220764,
|
| 236 |
-
"accuracy": 0.765,
|
| 237 |
-
"brier_soft": 0.06276483290052662,
|
| 238 |
-
"decisions": 600
|
| 239 |
-
},
|
| 240 |
-
"early_stopping": {
|
| 241 |
-
"patience": 3,
|
| 242 |
-
"min_epochs": 10,
|
| 243 |
-
"epochs_without_improvement": 0,
|
| 244 |
-
"best_validation_loss": 0.8377540612220764,
|
| 245 |
-
"improved": true
|
| 246 |
-
}
|
| 247 |
-
},
|
| 248 |
-
{
|
| 249 |
-
"epoch": 15,
|
| 250 |
-
"train_soft_cross_entropy": 0.7658979249000549,
|
| 251 |
-
"validation": {
|
| 252 |
-
"soft_cross_entropy": 0.8519912085930507,
|
| 253 |
-
"accuracy": 0.765,
|
| 254 |
-
"brier_soft": 0.07007166295622785,
|
| 255 |
-
"decisions": 600
|
| 256 |
-
},
|
| 257 |
-
"early_stopping": {
|
| 258 |
-
"patience": 3,
|
| 259 |
-
"min_epochs": 10,
|
| 260 |
-
"epochs_without_improvement": 1,
|
| 261 |
-
"best_validation_loss": 0.8377540612220764,
|
| 262 |
-
"improved": false
|
| 263 |
-
}
|
| 264 |
-
},
|
| 265 |
-
{
|
| 266 |
-
"epoch": 16,
|
| 267 |
-
"train_soft_cross_entropy": 0.7651108237107594,
|
| 268 |
-
"validation": {
|
| 269 |
-
"soft_cross_entropy": 0.8408213845888773,
|
| 270 |
-
"accuracy": 0.7633333333333333,
|
| 271 |
-
"brier_soft": 0.06544813718336324,
|
| 272 |
-
"decisions": 600
|
| 273 |
-
},
|
| 274 |
-
"early_stopping": {
|
| 275 |
-
"patience": 3,
|
| 276 |
-
"min_epochs": 10,
|
| 277 |
-
"epochs_without_improvement": 2,
|
| 278 |
-
"best_validation_loss": 0.8377540612220764,
|
| 279 |
-
"improved": false
|
| 280 |
-
}
|
| 281 |
-
},
|
| 282 |
-
{
|
| 283 |
-
"epoch": 17,
|
| 284 |
-
"train_soft_cross_entropy": 0.7817914198063038,
|
| 285 |
"validation": {
|
| 286 |
-
"soft_cross_entropy": 0.
|
| 287 |
-
"accuracy": 0.
|
| 288 |
-
"brier_soft": 0.
|
| 289 |
"decisions": 600
|
| 290 |
},
|
| 291 |
"early_stopping": {
|
| 292 |
"patience": 3,
|
| 293 |
"min_epochs": 10,
|
| 294 |
"epochs_without_improvement": 3,
|
| 295 |
-
"best_validation_loss": 0.
|
| 296 |
"improved": false
|
| 297 |
}
|
| 298 |
}
|
| 299 |
],
|
| 300 |
"stopping": {
|
| 301 |
"reason": "early_stopping",
|
| 302 |
-
"epochs_completed":
|
| 303 |
"patience": 3,
|
| 304 |
"min_epochs": 10,
|
| 305 |
"epochs_without_improvement": 3,
|
|
|
|
| 9 |
"epochs": [
|
| 10 |
{
|
| 11 |
"epoch": 1,
|
| 12 |
+
"train_soft_cross_entropy": 1.0518495603843971,
|
| 13 |
"validation": {
|
| 14 |
+
"soft_cross_entropy": 0.9761401589711507,
|
| 15 |
+
"accuracy": 0.6666666666666666,
|
| 16 |
+
"brier_soft": 0.14186986642579238,
|
| 17 |
"decisions": 600
|
| 18 |
},
|
| 19 |
"early_stopping": {
|
| 20 |
"patience": 3,
|
| 21 |
"min_epochs": 10,
|
| 22 |
"epochs_without_improvement": 0,
|
| 23 |
+
"best_validation_loss": 0.9761401589711507,
|
| 24 |
"improved": true
|
| 25 |
}
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"epoch": 2,
|
| 29 |
+
"train_soft_cross_entropy": 0.912936895776678,
|
| 30 |
"validation": {
|
| 31 |
+
"soft_cross_entropy": 0.9058315515518188,
|
| 32 |
+
"accuracy": 0.7133333333333334,
|
| 33 |
+
"brier_soft": 0.09862347106138865,
|
| 34 |
"decisions": 600
|
| 35 |
},
|
| 36 |
"early_stopping": {
|
| 37 |
"patience": 3,
|
| 38 |
"min_epochs": 10,
|
| 39 |
"epochs_without_improvement": 0,
|
| 40 |
+
"best_validation_loss": 0.9058315515518188,
|
| 41 |
"improved": true
|
| 42 |
}
|
| 43 |
},
|
| 44 |
{
|
| 45 |
"epoch": 3,
|
| 46 |
+
"train_soft_cross_entropy": 0.8595328189708569,
|
| 47 |
"validation": {
|
| 48 |
+
"soft_cross_entropy": 0.8869859429200491,
|
| 49 |
+
"accuracy": 0.7333333333333333,
|
| 50 |
+
"brier_soft": 0.08666441923628251,
|
| 51 |
"decisions": 600
|
| 52 |
},
|
| 53 |
"early_stopping": {
|
| 54 |
"patience": 3,
|
| 55 |
"min_epochs": 10,
|
| 56 |
"epochs_without_improvement": 0,
|
| 57 |
+
"best_validation_loss": 0.8869859429200491,
|
| 58 |
"improved": true
|
| 59 |
}
|
| 60 |
},
|
| 61 |
{
|
| 62 |
"epoch": 4,
|
| 63 |
+
"train_soft_cross_entropy": 0.8311581687574033,
|
| 64 |
"validation": {
|
| 65 |
+
"soft_cross_entropy": 0.875844070315361,
|
| 66 |
+
"accuracy": 0.7466666666666667,
|
| 67 |
+
"brier_soft": 0.08233745716512203,
|
| 68 |
"decisions": 600
|
| 69 |
},
|
| 70 |
"early_stopping": {
|
| 71 |
"patience": 3,
|
| 72 |
"min_epochs": 10,
|
| 73 |
"epochs_without_improvement": 0,
|
| 74 |
+
"best_validation_loss": 0.875844070315361,
|
| 75 |
"improved": true
|
| 76 |
}
|
| 77 |
},
|
| 78 |
{
|
| 79 |
"epoch": 5,
|
| 80 |
+
"train_soft_cross_entropy": 0.812676587899526,
|
| 81 |
"validation": {
|
| 82 |
+
"soft_cross_entropy": 0.8657626557350159,
|
| 83 |
+
"accuracy": 0.74,
|
| 84 |
+
"brier_soft": 0.07910436561020712,
|
| 85 |
"decisions": 600
|
| 86 |
},
|
| 87 |
"early_stopping": {
|
| 88 |
"patience": 3,
|
| 89 |
"min_epochs": 10,
|
| 90 |
+
"epochs_without_improvement": 0,
|
| 91 |
+
"best_validation_loss": 0.8657626557350159,
|
| 92 |
+
"improved": true
|
| 93 |
}
|
| 94 |
},
|
| 95 |
{
|
| 96 |
"epoch": 6,
|
| 97 |
+
"train_soft_cross_entropy": 0.7997228077605919,
|
| 98 |
"validation": {
|
| 99 |
+
"soft_cross_entropy": 0.8763724052906037,
|
| 100 |
+
"accuracy": 0.7433333333333333,
|
| 101 |
+
"brier_soft": 0.08383749471356472,
|
| 102 |
"decisions": 600
|
| 103 |
},
|
| 104 |
"early_stopping": {
|
| 105 |
"patience": 3,
|
| 106 |
"min_epochs": 10,
|
| 107 |
+
"epochs_without_improvement": 1,
|
| 108 |
+
"best_validation_loss": 0.8657626557350159,
|
| 109 |
"improved": false
|
| 110 |
}
|
| 111 |
},
|
| 112 |
{
|
| 113 |
"epoch": 7,
|
| 114 |
+
"train_soft_cross_entropy": 0.7919920165892,
|
| 115 |
"validation": {
|
| 116 |
+
"soft_cross_entropy": 0.8720610602696737,
|
| 117 |
"accuracy": 0.7466666666666667,
|
| 118 |
+
"brier_soft": 0.08182110945383708,
|
| 119 |
"decisions": 600
|
| 120 |
},
|
| 121 |
"early_stopping": {
|
| 122 |
"patience": 3,
|
| 123 |
"min_epochs": 10,
|
| 124 |
+
"epochs_without_improvement": 2,
|
| 125 |
+
"best_validation_loss": 0.8657626557350159,
|
| 126 |
+
"improved": false
|
| 127 |
}
|
| 128 |
},
|
| 129 |
{
|
| 130 |
"epoch": 8,
|
| 131 |
+
"train_soft_cross_entropy": 0.7865607251061334,
|
| 132 |
"validation": {
|
| 133 |
+
"soft_cross_entropy": 0.8708612742026647,
|
| 134 |
+
"accuracy": 0.735,
|
| 135 |
+
"brier_soft": 0.08140341289962331,
|
| 136 |
"decisions": 600
|
| 137 |
},
|
| 138 |
"early_stopping": {
|
| 139 |
"patience": 3,
|
| 140 |
"min_epochs": 10,
|
| 141 |
+
"epochs_without_improvement": 3,
|
| 142 |
+
"best_validation_loss": 0.8657626557350159,
|
| 143 |
+
"improved": false
|
| 144 |
}
|
| 145 |
},
|
| 146 |
{
|
| 147 |
"epoch": 9,
|
| 148 |
+
"train_soft_cross_entropy": 0.781722169496395,
|
| 149 |
"validation": {
|
| 150 |
+
"soft_cross_entropy": 0.8583925066391627,
|
| 151 |
+
"accuracy": 0.7466666666666667,
|
| 152 |
+
"brier_soft": 0.07437195796364297,
|
| 153 |
"decisions": 600
|
| 154 |
},
|
| 155 |
"early_stopping": {
|
| 156 |
"patience": 3,
|
| 157 |
"min_epochs": 10,
|
| 158 |
"epochs_without_improvement": 0,
|
| 159 |
+
"best_validation_loss": 0.8583925066391627,
|
| 160 |
"improved": true
|
| 161 |
}
|
| 162 |
},
|
| 163 |
{
|
| 164 |
"epoch": 10,
|
| 165 |
+
"train_soft_cross_entropy": 0.7776996397530591,
|
| 166 |
"validation": {
|
| 167 |
+
"soft_cross_entropy": 0.8674025861422221,
|
| 168 |
+
"accuracy": 0.7433333333333333,
|
| 169 |
+
"brier_soft": 0.07951540602991979,
|
| 170 |
"decisions": 600
|
| 171 |
},
|
| 172 |
"early_stopping": {
|
| 173 |
"patience": 3,
|
| 174 |
"min_epochs": 10,
|
| 175 |
"epochs_without_improvement": 1,
|
| 176 |
+
"best_validation_loss": 0.8583925066391627,
|
| 177 |
"improved": false
|
| 178 |
}
|
| 179 |
},
|
| 180 |
{
|
| 181 |
"epoch": 11,
|
| 182 |
+
"train_soft_cross_entropy": 0.7751605097894315,
|
| 183 |
"validation": {
|
| 184 |
+
"soft_cross_entropy": 0.86817025522391,
|
| 185 |
+
"accuracy": 0.7433333333333333,
|
| 186 |
+
"brier_soft": 0.07928484301393231,
|
| 187 |
"decisions": 600
|
| 188 |
},
|
| 189 |
"early_stopping": {
|
| 190 |
"patience": 3,
|
| 191 |
"min_epochs": 10,
|
| 192 |
"epochs_without_improvement": 2,
|
| 193 |
+
"best_validation_loss": 0.8583925066391627,
|
| 194 |
"improved": false
|
| 195 |
}
|
| 196 |
},
|
| 197 |
{
|
| 198 |
"epoch": 12,
|
| 199 |
+
"train_soft_cross_entropy": 0.7721157057638521,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 200 |
"validation": {
|
| 201 |
+
"soft_cross_entropy": 0.8638703493277232,
|
| 202 |
+
"accuracy": 0.7583333333333333,
|
| 203 |
+
"brier_soft": 0.0763223198801279,
|
| 204 |
"decisions": 600
|
| 205 |
},
|
| 206 |
"early_stopping": {
|
| 207 |
"patience": 3,
|
| 208 |
"min_epochs": 10,
|
| 209 |
"epochs_without_improvement": 3,
|
| 210 |
+
"best_validation_loss": 0.8583925066391627,
|
| 211 |
"improved": false
|
| 212 |
}
|
| 213 |
}
|
| 214 |
],
|
| 215 |
"stopping": {
|
| 216 |
"reason": "early_stopping",
|
| 217 |
+
"epochs_completed": 12,
|
| 218 |
"patience": 3,
|
| 219 |
"min_epochs": 10,
|
| 220 |
"epochs_without_improvement": 3,
|
evidence/seed-2-history.json
CHANGED
|
@@ -9,314 +9,280 @@
|
|
| 9 |
"epochs": [
|
| 10 |
{
|
| 11 |
"epoch": 1,
|
| 12 |
-
"train_soft_cross_entropy": 1.
|
| 13 |
"validation": {
|
| 14 |
-
"soft_cross_entropy":
|
| 15 |
-
"accuracy": 0.
|
| 16 |
-
"brier_soft": 0.
|
| 17 |
"decisions": 600
|
| 18 |
},
|
| 19 |
"early_stopping": {
|
| 20 |
"patience": 3,
|
| 21 |
"min_epochs": 10,
|
| 22 |
"epochs_without_improvement": 0,
|
| 23 |
-
"best_validation_loss":
|
| 24 |
"improved": true
|
| 25 |
}
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"epoch": 2,
|
| 29 |
-
"train_soft_cross_entropy":
|
| 30 |
"validation": {
|
| 31 |
-
"soft_cross_entropy":
|
| 32 |
-
"accuracy": 0.
|
| 33 |
-
"brier_soft": 0.
|
| 34 |
"decisions": 600
|
| 35 |
},
|
| 36 |
"early_stopping": {
|
| 37 |
"patience": 3,
|
| 38 |
"min_epochs": 10,
|
| 39 |
"epochs_without_improvement": 0,
|
| 40 |
-
"best_validation_loss":
|
| 41 |
"improved": true
|
| 42 |
}
|
| 43 |
},
|
| 44 |
{
|
| 45 |
"epoch": 3,
|
| 46 |
-
"train_soft_cross_entropy":
|
| 47 |
"validation": {
|
| 48 |
-
"soft_cross_entropy":
|
| 49 |
-
"accuracy": 0.
|
| 50 |
-
"brier_soft": 0.
|
| 51 |
"decisions": 600
|
| 52 |
},
|
| 53 |
"early_stopping": {
|
| 54 |
"patience": 3,
|
| 55 |
"min_epochs": 10,
|
| 56 |
"epochs_without_improvement": 0,
|
| 57 |
-
"best_validation_loss":
|
| 58 |
"improved": true
|
| 59 |
}
|
| 60 |
},
|
| 61 |
{
|
| 62 |
"epoch": 4,
|
| 63 |
-
"train_soft_cross_entropy":
|
| 64 |
"validation": {
|
| 65 |
-
"soft_cross_entropy":
|
| 66 |
-
"accuracy": 0.
|
| 67 |
-
"brier_soft": 0.
|
| 68 |
"decisions": 600
|
| 69 |
},
|
| 70 |
"early_stopping": {
|
| 71 |
"patience": 3,
|
| 72 |
"min_epochs": 10,
|
| 73 |
"epochs_without_improvement": 0,
|
| 74 |
-
"best_validation_loss":
|
| 75 |
"improved": true
|
| 76 |
}
|
| 77 |
},
|
| 78 |
{
|
| 79 |
"epoch": 5,
|
| 80 |
-
"train_soft_cross_entropy":
|
| 81 |
"validation": {
|
| 82 |
-
"soft_cross_entropy":
|
| 83 |
-
"accuracy": 0.
|
| 84 |
-
"brier_soft": 0.
|
| 85 |
"decisions": 600
|
| 86 |
},
|
| 87 |
"early_stopping": {
|
| 88 |
"patience": 3,
|
| 89 |
"min_epochs": 10,
|
| 90 |
-
"epochs_without_improvement":
|
| 91 |
-
"best_validation_loss":
|
| 92 |
-
"improved":
|
| 93 |
}
|
| 94 |
},
|
| 95 |
{
|
| 96 |
"epoch": 6,
|
| 97 |
-
"train_soft_cross_entropy":
|
| 98 |
"validation": {
|
| 99 |
-
"soft_cross_entropy":
|
| 100 |
-
"accuracy": 0.
|
| 101 |
-
"brier_soft": 0.
|
| 102 |
"decisions": 600
|
| 103 |
},
|
| 104 |
"early_stopping": {
|
| 105 |
"patience": 3,
|
| 106 |
"min_epochs": 10,
|
| 107 |
"epochs_without_improvement": 0,
|
| 108 |
-
"best_validation_loss":
|
| 109 |
"improved": true
|
| 110 |
}
|
| 111 |
},
|
| 112 |
{
|
| 113 |
"epoch": 7,
|
| 114 |
-
"train_soft_cross_entropy":
|
| 115 |
"validation": {
|
| 116 |
-
"soft_cross_entropy":
|
| 117 |
-
"accuracy": 0.
|
| 118 |
-
"brier_soft": 0.
|
| 119 |
"decisions": 600
|
| 120 |
},
|
| 121 |
"early_stopping": {
|
| 122 |
"patience": 3,
|
| 123 |
"min_epochs": 10,
|
| 124 |
-
"epochs_without_improvement":
|
| 125 |
-
"best_validation_loss":
|
| 126 |
-
"improved":
|
| 127 |
}
|
| 128 |
},
|
| 129 |
{
|
| 130 |
"epoch": 8,
|
| 131 |
-
"train_soft_cross_entropy": 0.
|
| 132 |
"validation": {
|
| 133 |
-
"soft_cross_entropy":
|
| 134 |
-
"accuracy": 0.
|
| 135 |
-
"brier_soft": 0.
|
| 136 |
"decisions": 600
|
| 137 |
},
|
| 138 |
"early_stopping": {
|
| 139 |
"patience": 3,
|
| 140 |
"min_epochs": 10,
|
| 141 |
-
"epochs_without_improvement":
|
| 142 |
-
"best_validation_loss":
|
| 143 |
-
"improved":
|
| 144 |
}
|
| 145 |
},
|
| 146 |
{
|
| 147 |
"epoch": 9,
|
| 148 |
-
"train_soft_cross_entropy": 0.
|
| 149 |
"validation": {
|
| 150 |
-
"soft_cross_entropy":
|
| 151 |
-
"accuracy": 0.
|
| 152 |
-
"brier_soft": 0.
|
| 153 |
"decisions": 600
|
| 154 |
},
|
| 155 |
"early_stopping": {
|
| 156 |
"patience": 3,
|
| 157 |
"min_epochs": 10,
|
| 158 |
-
"epochs_without_improvement":
|
| 159 |
-
"best_validation_loss":
|
| 160 |
"improved": false
|
| 161 |
}
|
| 162 |
},
|
| 163 |
{
|
| 164 |
"epoch": 10,
|
| 165 |
-
"train_soft_cross_entropy": 0.
|
| 166 |
"validation": {
|
| 167 |
-
"soft_cross_entropy":
|
| 168 |
-
"accuracy": 0.
|
| 169 |
-
"brier_soft": 0.
|
| 170 |
"decisions": 600
|
| 171 |
},
|
| 172 |
"early_stopping": {
|
| 173 |
"patience": 3,
|
| 174 |
"min_epochs": 10,
|
| 175 |
"epochs_without_improvement": 0,
|
| 176 |
-
"best_validation_loss":
|
| 177 |
"improved": true
|
| 178 |
}
|
| 179 |
},
|
| 180 |
{
|
| 181 |
"epoch": 11,
|
| 182 |
-
"train_soft_cross_entropy": 0.
|
| 183 |
-
"validation": {
|
| 184 |
-
"soft_cross_entropy": 1.0138032054901123,
|
| 185 |
-
"accuracy": 0.6166666666666667,
|
| 186 |
-
"brier_soft": 0.165894419302543,
|
| 187 |
-
"decisions": 600
|
| 188 |
-
},
|
| 189 |
-
"early_stopping": {
|
| 190 |
-
"patience": 3,
|
| 191 |
-
"min_epochs": 10,
|
| 192 |
-
"epochs_without_improvement": 0,
|
| 193 |
-
"best_validation_loss": 1.0138032054901123,
|
| 194 |
-
"improved": true
|
| 195 |
-
}
|
| 196 |
-
},
|
| 197 |
-
{
|
| 198 |
-
"epoch": 12,
|
| 199 |
-
"train_soft_cross_entropy": 0.9309376364284091,
|
| 200 |
"validation": {
|
| 201 |
-
"soft_cross_entropy": 0.
|
| 202 |
-
"accuracy": 0.
|
| 203 |
-
"brier_soft": 0.
|
| 204 |
-
"decisions": 600
|
| 205 |
-
},
|
| 206 |
-
"early_stopping": {
|
| 207 |
-
"patience": 3,
|
| 208 |
-
"min_epochs": 10,
|
| 209 |
-
"epochs_without_improvement": 0,
|
| 210 |
-
"best_validation_loss": 0.9929217569033305,
|
| 211 |
-
"improved": true
|
| 212 |
-
}
|
| 213 |
-
},
|
| 214 |
-
{
|
| 215 |
-
"epoch": 13,
|
| 216 |
-
"train_soft_cross_entropy": 0.9144250082969666,
|
| 217 |
-
"validation": {
|
| 218 |
-
"soft_cross_entropy": 0.9964303588867187,
|
| 219 |
-
"accuracy": 0.6266666666666667,
|
| 220 |
-
"brier_soft": 0.1560174826408426,
|
| 221 |
"decisions": 600
|
| 222 |
},
|
| 223 |
"early_stopping": {
|
| 224 |
"patience": 3,
|
| 225 |
"min_epochs": 10,
|
| 226 |
"epochs_without_improvement": 1,
|
| 227 |
-
"best_validation_loss": 0.
|
| 228 |
"improved": false
|
| 229 |
}
|
| 230 |
},
|
| 231 |
{
|
| 232 |
-
"epoch":
|
| 233 |
-
"train_soft_cross_entropy": 0.
|
| 234 |
"validation": {
|
| 235 |
-
"soft_cross_entropy":
|
| 236 |
-
"accuracy": 0.
|
| 237 |
-
"brier_soft": 0.
|
| 238 |
"decisions": 600
|
| 239 |
},
|
| 240 |
"early_stopping": {
|
| 241 |
"patience": 3,
|
| 242 |
"min_epochs": 10,
|
| 243 |
"epochs_without_improvement": 2,
|
| 244 |
-
"best_validation_loss": 0.
|
| 245 |
"improved": false
|
| 246 |
}
|
| 247 |
},
|
| 248 |
{
|
| 249 |
-
"epoch":
|
| 250 |
-
"train_soft_cross_entropy": 0.
|
| 251 |
"validation": {
|
| 252 |
-
"soft_cross_entropy": 0.
|
| 253 |
-
"accuracy": 0.
|
| 254 |
-
"brier_soft": 0.
|
| 255 |
"decisions": 600
|
| 256 |
},
|
| 257 |
"early_stopping": {
|
| 258 |
"patience": 3,
|
| 259 |
"min_epochs": 10,
|
| 260 |
"epochs_without_improvement": 0,
|
| 261 |
-
"best_validation_loss": 0.
|
| 262 |
"improved": true
|
| 263 |
}
|
| 264 |
},
|
| 265 |
{
|
| 266 |
-
"epoch":
|
| 267 |
-
"train_soft_cross_entropy": 0.
|
| 268 |
"validation": {
|
| 269 |
-
"soft_cross_entropy": 0.
|
| 270 |
-
"accuracy": 0.
|
| 271 |
-
"brier_soft": 0.
|
| 272 |
"decisions": 600
|
| 273 |
},
|
| 274 |
"early_stopping": {
|
| 275 |
"patience": 3,
|
| 276 |
"min_epochs": 10,
|
| 277 |
"epochs_without_improvement": 1,
|
| 278 |
-
"best_validation_loss": 0.
|
| 279 |
"improved": false
|
| 280 |
}
|
| 281 |
},
|
| 282 |
{
|
| 283 |
-
"epoch":
|
| 284 |
-
"train_soft_cross_entropy": 0.
|
| 285 |
"validation": {
|
| 286 |
-
"soft_cross_entropy":
|
| 287 |
-
"accuracy": 0.
|
| 288 |
-
"brier_soft": 0.
|
| 289 |
"decisions": 600
|
| 290 |
},
|
| 291 |
"early_stopping": {
|
| 292 |
"patience": 3,
|
| 293 |
"min_epochs": 10,
|
| 294 |
"epochs_without_improvement": 2,
|
| 295 |
-
"best_validation_loss": 0.
|
| 296 |
"improved": false
|
| 297 |
}
|
| 298 |
},
|
| 299 |
{
|
| 300 |
-
"epoch":
|
| 301 |
-
"train_soft_cross_entropy": 0.
|
| 302 |
"validation": {
|
| 303 |
-
"soft_cross_entropy":
|
| 304 |
-
"accuracy": 0.
|
| 305 |
-
"brier_soft": 0.
|
| 306 |
"decisions": 600
|
| 307 |
},
|
| 308 |
"early_stopping": {
|
| 309 |
"patience": 3,
|
| 310 |
"min_epochs": 10,
|
| 311 |
"epochs_without_improvement": 3,
|
| 312 |
-
"best_validation_loss": 0.
|
| 313 |
"improved": false
|
| 314 |
}
|
| 315 |
}
|
| 316 |
],
|
| 317 |
"stopping": {
|
| 318 |
"reason": "early_stopping",
|
| 319 |
-
"epochs_completed":
|
| 320 |
"patience": 3,
|
| 321 |
"min_epochs": 10,
|
| 322 |
"epochs_without_improvement": 3,
|
|
|
|
| 9 |
"epochs": [
|
| 10 |
{
|
| 11 |
"epoch": 1,
|
| 12 |
+
"train_soft_cross_entropy": 1.0231783160456904,
|
| 13 |
"validation": {
|
| 14 |
+
"soft_cross_entropy": 0.9674482258160909,
|
| 15 |
+
"accuracy": 0.6516666666666666,
|
| 16 |
+
"brier_soft": 0.139310811907053,
|
| 17 |
"decisions": 600
|
| 18 |
},
|
| 19 |
"early_stopping": {
|
| 20 |
"patience": 3,
|
| 21 |
"min_epochs": 10,
|
| 22 |
"epochs_without_improvement": 0,
|
| 23 |
+
"best_validation_loss": 0.9674482258160909,
|
| 24 |
"improved": true
|
| 25 |
}
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"epoch": 2,
|
| 29 |
+
"train_soft_cross_entropy": 0.8996694993972778,
|
| 30 |
"validation": {
|
| 31 |
+
"soft_cross_entropy": 0.8974683928489685,
|
| 32 |
+
"accuracy": 0.75,
|
| 33 |
+
"brier_soft": 0.09312087532132864,
|
| 34 |
"decisions": 600
|
| 35 |
},
|
| 36 |
"early_stopping": {
|
| 37 |
"patience": 3,
|
| 38 |
"min_epochs": 10,
|
| 39 |
"epochs_without_improvement": 0,
|
| 40 |
+
"best_validation_loss": 0.8974683928489685,
|
| 41 |
"improved": true
|
| 42 |
}
|
| 43 |
},
|
| 44 |
{
|
| 45 |
"epoch": 3,
|
| 46 |
+
"train_soft_cross_entropy": 0.8483260071277618,
|
| 47 |
"validation": {
|
| 48 |
+
"soft_cross_entropy": 0.8809152638912201,
|
| 49 |
+
"accuracy": 0.7433333333333333,
|
| 50 |
+
"brier_soft": 0.08576706000914176,
|
| 51 |
"decisions": 600
|
| 52 |
},
|
| 53 |
"early_stopping": {
|
| 54 |
"patience": 3,
|
| 55 |
"min_epochs": 10,
|
| 56 |
"epochs_without_improvement": 0,
|
| 57 |
+
"best_validation_loss": 0.8809152638912201,
|
| 58 |
"improved": true
|
| 59 |
}
|
| 60 |
},
|
| 61 |
{
|
| 62 |
"epoch": 4,
|
| 63 |
+
"train_soft_cross_entropy": 0.8239163917523843,
|
| 64 |
"validation": {
|
| 65 |
+
"soft_cross_entropy": 0.8699432893594106,
|
| 66 |
+
"accuracy": 0.75,
|
| 67 |
+
"brier_soft": 0.07744032381723324,
|
| 68 |
"decisions": 600
|
| 69 |
},
|
| 70 |
"early_stopping": {
|
| 71 |
"patience": 3,
|
| 72 |
"min_epochs": 10,
|
| 73 |
"epochs_without_improvement": 0,
|
| 74 |
+
"best_validation_loss": 0.8699432893594106,
|
| 75 |
"improved": true
|
| 76 |
}
|
| 77 |
},
|
| 78 |
{
|
| 79 |
"epoch": 5,
|
| 80 |
+
"train_soft_cross_entropy": 0.8087966615623898,
|
| 81 |
"validation": {
|
| 82 |
+
"soft_cross_entropy": 0.8749503823121388,
|
| 83 |
+
"accuracy": 0.725,
|
| 84 |
+
"brier_soft": 0.08202173060427109,
|
| 85 |
"decisions": 600
|
| 86 |
},
|
| 87 |
"early_stopping": {
|
| 88 |
"patience": 3,
|
| 89 |
"min_epochs": 10,
|
| 90 |
+
"epochs_without_improvement": 1,
|
| 91 |
+
"best_validation_loss": 0.8699432893594106,
|
| 92 |
+
"improved": false
|
| 93 |
}
|
| 94 |
},
|
| 95 |
{
|
| 96 |
"epoch": 6,
|
| 97 |
+
"train_soft_cross_entropy": 0.798634327782525,
|
| 98 |
"validation": {
|
| 99 |
+
"soft_cross_entropy": 0.8646231130758921,
|
| 100 |
+
"accuracy": 0.7633333333333333,
|
| 101 |
+
"brier_soft": 0.07594866358985504,
|
| 102 |
"decisions": 600
|
| 103 |
},
|
| 104 |
"early_stopping": {
|
| 105 |
"patience": 3,
|
| 106 |
"min_epochs": 10,
|
| 107 |
"epochs_without_improvement": 0,
|
| 108 |
+
"best_validation_loss": 0.8646231130758921,
|
| 109 |
"improved": true
|
| 110 |
}
|
| 111 |
},
|
| 112 |
{
|
| 113 |
"epoch": 7,
|
| 114 |
+
"train_soft_cross_entropy": 0.7900045536182545,
|
| 115 |
"validation": {
|
| 116 |
+
"soft_cross_entropy": 0.8598153670628865,
|
| 117 |
+
"accuracy": 0.755,
|
| 118 |
+
"brier_soft": 0.07338624291742842,
|
| 119 |
"decisions": 600
|
| 120 |
},
|
| 121 |
"early_stopping": {
|
| 122 |
"patience": 3,
|
| 123 |
"min_epochs": 10,
|
| 124 |
+
"epochs_without_improvement": 0,
|
| 125 |
+
"best_validation_loss": 0.8598153670628865,
|
| 126 |
+
"improved": true
|
| 127 |
}
|
| 128 |
},
|
| 129 |
{
|
| 130 |
"epoch": 8,
|
| 131 |
+
"train_soft_cross_entropy": 0.7848307268266325,
|
| 132 |
"validation": {
|
| 133 |
+
"soft_cross_entropy": 0.8584532356262207,
|
| 134 |
+
"accuracy": 0.7483333333333333,
|
| 135 |
+
"brier_soft": 0.07417237816999356,
|
| 136 |
"decisions": 600
|
| 137 |
},
|
| 138 |
"early_stopping": {
|
| 139 |
"patience": 3,
|
| 140 |
"min_epochs": 10,
|
| 141 |
+
"epochs_without_improvement": 0,
|
| 142 |
+
"best_validation_loss": 0.8584532356262207,
|
| 143 |
+
"improved": true
|
| 144 |
}
|
| 145 |
},
|
| 146 |
{
|
| 147 |
"epoch": 9,
|
| 148 |
+
"train_soft_cross_entropy": 0.7802449210043306,
|
| 149 |
"validation": {
|
| 150 |
+
"soft_cross_entropy": 0.863755419254303,
|
| 151 |
+
"accuracy": 0.7433333333333333,
|
| 152 |
+
"brier_soft": 0.07594012685120105,
|
| 153 |
"decisions": 600
|
| 154 |
},
|
| 155 |
"early_stopping": {
|
| 156 |
"patience": 3,
|
| 157 |
"min_epochs": 10,
|
| 158 |
+
"epochs_without_improvement": 1,
|
| 159 |
+
"best_validation_loss": 0.8584532356262207,
|
| 160 |
"improved": false
|
| 161 |
}
|
| 162 |
},
|
| 163 |
{
|
| 164 |
"epoch": 10,
|
| 165 |
+
"train_soft_cross_entropy": 0.777841743098365,
|
| 166 |
"validation": {
|
| 167 |
+
"soft_cross_entropy": 0.8546390326817831,
|
| 168 |
+
"accuracy": 0.7566666666666667,
|
| 169 |
+
"brier_soft": 0.07055099235226711,
|
| 170 |
"decisions": 600
|
| 171 |
},
|
| 172 |
"early_stopping": {
|
| 173 |
"patience": 3,
|
| 174 |
"min_epochs": 10,
|
| 175 |
"epochs_without_improvement": 0,
|
| 176 |
+
"best_validation_loss": 0.8546390326817831,
|
| 177 |
"improved": true
|
| 178 |
}
|
| 179 |
},
|
| 180 |
{
|
| 181 |
"epoch": 11,
|
| 182 |
+
"train_soft_cross_entropy": 0.7738726913045954,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 183 |
"validation": {
|
| 184 |
+
"soft_cross_entropy": 0.8676082201798757,
|
| 185 |
+
"accuracy": 0.7466666666666667,
|
| 186 |
+
"brier_soft": 0.07854201994836331,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 187 |
"decisions": 600
|
| 188 |
},
|
| 189 |
"early_stopping": {
|
| 190 |
"patience": 3,
|
| 191 |
"min_epochs": 10,
|
| 192 |
"epochs_without_improvement": 1,
|
| 193 |
+
"best_validation_loss": 0.8546390326817831,
|
| 194 |
"improved": false
|
| 195 |
}
|
| 196 |
},
|
| 197 |
{
|
| 198 |
+
"epoch": 12,
|
| 199 |
+
"train_soft_cross_entropy": 0.7718432058228387,
|
| 200 |
"validation": {
|
| 201 |
+
"soft_cross_entropy": 0.8577940479914348,
|
| 202 |
+
"accuracy": 0.7366666666666667,
|
| 203 |
+
"brier_soft": 0.07410156331335505,
|
| 204 |
"decisions": 600
|
| 205 |
},
|
| 206 |
"early_stopping": {
|
| 207 |
"patience": 3,
|
| 208 |
"min_epochs": 10,
|
| 209 |
"epochs_without_improvement": 2,
|
| 210 |
+
"best_validation_loss": 0.8546390326817831,
|
| 211 |
"improved": false
|
| 212 |
}
|
| 213 |
},
|
| 214 |
{
|
| 215 |
+
"epoch": 13,
|
| 216 |
+
"train_soft_cross_entropy": 0.7713618384467231,
|
| 217 |
"validation": {
|
| 218 |
+
"soft_cross_entropy": 0.8543428432941437,
|
| 219 |
+
"accuracy": 0.7583333333333333,
|
| 220 |
+
"brier_soft": 0.07187043125430743,
|
| 221 |
"decisions": 600
|
| 222 |
},
|
| 223 |
"early_stopping": {
|
| 224 |
"patience": 3,
|
| 225 |
"min_epochs": 10,
|
| 226 |
"epochs_without_improvement": 0,
|
| 227 |
+
"best_validation_loss": 0.8543428432941437,
|
| 228 |
"improved": true
|
| 229 |
}
|
| 230 |
},
|
| 231 |
{
|
| 232 |
+
"epoch": 14,
|
| 233 |
+
"train_soft_cross_entropy": 0.7684376273331819,
|
| 234 |
"validation": {
|
| 235 |
+
"soft_cross_entropy": 0.8578398124376932,
|
| 236 |
+
"accuracy": 0.7483333333333333,
|
| 237 |
+
"brier_soft": 0.07348543658852577,
|
| 238 |
"decisions": 600
|
| 239 |
},
|
| 240 |
"early_stopping": {
|
| 241 |
"patience": 3,
|
| 242 |
"min_epochs": 10,
|
| 243 |
"epochs_without_improvement": 1,
|
| 244 |
+
"best_validation_loss": 0.8543428432941437,
|
| 245 |
"improved": false
|
| 246 |
}
|
| 247 |
},
|
| 248 |
{
|
| 249 |
+
"epoch": 15,
|
| 250 |
+
"train_soft_cross_entropy": 0.766730530526903,
|
| 251 |
"validation": {
|
| 252 |
+
"soft_cross_entropy": 0.860800955692927,
|
| 253 |
+
"accuracy": 0.7383333333333333,
|
| 254 |
+
"brier_soft": 0.07522662562007705,
|
| 255 |
"decisions": 600
|
| 256 |
},
|
| 257 |
"early_stopping": {
|
| 258 |
"patience": 3,
|
| 259 |
"min_epochs": 10,
|
| 260 |
"epochs_without_improvement": 2,
|
| 261 |
+
"best_validation_loss": 0.8543428432941437,
|
| 262 |
"improved": false
|
| 263 |
}
|
| 264 |
},
|
| 265 |
{
|
| 266 |
+
"epoch": 16,
|
| 267 |
+
"train_soft_cross_entropy": 0.7659876039734593,
|
| 268 |
"validation": {
|
| 269 |
+
"soft_cross_entropy": 0.8564839088916778,
|
| 270 |
+
"accuracy": 0.7533333333333333,
|
| 271 |
+
"brier_soft": 0.07288925250992179,
|
| 272 |
"decisions": 600
|
| 273 |
},
|
| 274 |
"early_stopping": {
|
| 275 |
"patience": 3,
|
| 276 |
"min_epochs": 10,
|
| 277 |
"epochs_without_improvement": 3,
|
| 278 |
+
"best_validation_loss": 0.8543428432941437,
|
| 279 |
"improved": false
|
| 280 |
}
|
| 281 |
}
|
| 282 |
],
|
| 283 |
"stopping": {
|
| 284 |
"reason": "early_stopping",
|
| 285 |
+
"epochs_completed": 16,
|
| 286 |
"patience": 3,
|
| 287 |
"min_epochs": 10,
|
| 288 |
"epochs_without_improvement": 3,
|
evidence/seed-3-history.json
CHANGED
|
@@ -9,467 +9,195 @@
|
|
| 9 |
"epochs": [
|
| 10 |
{
|
| 11 |
"epoch": 1,
|
| 12 |
-
"train_soft_cross_entropy": 1.
|
| 13 |
"validation": {
|
| 14 |
-
"soft_cross_entropy":
|
| 15 |
-
"accuracy": 0.
|
| 16 |
-
"brier_soft": 0.
|
| 17 |
"decisions": 600
|
| 18 |
},
|
| 19 |
"early_stopping": {
|
| 20 |
"patience": 3,
|
| 21 |
"min_epochs": 10,
|
| 22 |
"epochs_without_improvement": 0,
|
| 23 |
-
"best_validation_loss":
|
| 24 |
"improved": true
|
| 25 |
}
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"epoch": 2,
|
| 29 |
-
"train_soft_cross_entropy":
|
| 30 |
"validation": {
|
| 31 |
-
"soft_cross_entropy":
|
| 32 |
-
"accuracy": 0.
|
| 33 |
-
"brier_soft": 0.
|
| 34 |
"decisions": 600
|
| 35 |
},
|
| 36 |
"early_stopping": {
|
| 37 |
"patience": 3,
|
| 38 |
"min_epochs": 10,
|
| 39 |
"epochs_without_improvement": 0,
|
| 40 |
-
"best_validation_loss":
|
| 41 |
"improved": true
|
| 42 |
}
|
| 43 |
},
|
| 44 |
{
|
| 45 |
"epoch": 3,
|
| 46 |
-
"train_soft_cross_entropy":
|
| 47 |
"validation": {
|
| 48 |
-
"soft_cross_entropy":
|
| 49 |
-
"accuracy": 0.
|
| 50 |
-
"brier_soft": 0.
|
| 51 |
"decisions": 600
|
| 52 |
},
|
| 53 |
"early_stopping": {
|
| 54 |
"patience": 3,
|
| 55 |
"min_epochs": 10,
|
| 56 |
"epochs_without_improvement": 0,
|
| 57 |
-
"best_validation_loss":
|
| 58 |
"improved": true
|
| 59 |
}
|
| 60 |
},
|
| 61 |
{
|
| 62 |
"epoch": 4,
|
| 63 |
-
"train_soft_cross_entropy":
|
| 64 |
"validation": {
|
| 65 |
-
"soft_cross_entropy":
|
| 66 |
-
"accuracy": 0.
|
| 67 |
-
"brier_soft": 0.
|
| 68 |
"decisions": 600
|
| 69 |
},
|
| 70 |
"early_stopping": {
|
| 71 |
"patience": 3,
|
| 72 |
"min_epochs": 10,
|
| 73 |
"epochs_without_improvement": 0,
|
| 74 |
-
"best_validation_loss":
|
| 75 |
"improved": true
|
| 76 |
}
|
| 77 |
},
|
| 78 |
{
|
| 79 |
"epoch": 5,
|
| 80 |
-
"train_soft_cross_entropy":
|
| 81 |
"validation": {
|
| 82 |
-
"soft_cross_entropy":
|
| 83 |
-
"accuracy": 0.
|
| 84 |
-
"brier_soft": 0.
|
| 85 |
"decisions": 600
|
| 86 |
},
|
| 87 |
"early_stopping": {
|
| 88 |
"patience": 3,
|
| 89 |
"min_epochs": 10,
|
| 90 |
"epochs_without_improvement": 1,
|
| 91 |
-
"best_validation_loss":
|
| 92 |
"improved": false
|
| 93 |
}
|
| 94 |
},
|
| 95 |
{
|
| 96 |
"epoch": 6,
|
| 97 |
-
"train_soft_cross_entropy":
|
| 98 |
-
"validation": {
|
| 99 |
-
"soft_cross_entropy": 1.0705021890004476,
|
| 100 |
-
"accuracy": 0.525,
|
| 101 |
-
"brier_soft": 0.19120527582863966,
|
| 102 |
-
"decisions": 600
|
| 103 |
-
},
|
| 104 |
-
"early_stopping": {
|
| 105 |
-
"patience": 3,
|
| 106 |
-
"min_epochs": 10,
|
| 107 |
-
"epochs_without_improvement": 2,
|
| 108 |
-
"best_validation_loss": 1.0549713468551636,
|
| 109 |
-
"improved": false
|
| 110 |
-
}
|
| 111 |
-
},
|
| 112 |
-
{
|
| 113 |
-
"epoch": 7,
|
| 114 |
-
"train_soft_cross_entropy": 1.0313762316880404,
|
| 115 |
-
"validation": {
|
| 116 |
-
"soft_cross_entropy": 1.0447989773750306,
|
| 117 |
-
"accuracy": 0.5466666666666666,
|
| 118 |
-
"brier_soft": 0.17501757830381393,
|
| 119 |
-
"decisions": 600
|
| 120 |
-
},
|
| 121 |
-
"early_stopping": {
|
| 122 |
-
"patience": 3,
|
| 123 |
-
"min_epochs": 10,
|
| 124 |
-
"epochs_without_improvement": 0,
|
| 125 |
-
"best_validation_loss": 1.0447989773750306,
|
| 126 |
-
"improved": true
|
| 127 |
-
}
|
| 128 |
-
},
|
| 129 |
-
{
|
| 130 |
-
"epoch": 8,
|
| 131 |
-
"train_soft_cross_entropy": 1.017818791601393,
|
| 132 |
-
"validation": {
|
| 133 |
-
"soft_cross_entropy": 1.024161856174469,
|
| 134 |
-
"accuracy": 0.5816666666666667,
|
| 135 |
-
"brier_soft": 0.1630909529576699,
|
| 136 |
-
"decisions": 600
|
| 137 |
-
},
|
| 138 |
-
"early_stopping": {
|
| 139 |
-
"patience": 3,
|
| 140 |
-
"min_epochs": 10,
|
| 141 |
-
"epochs_without_improvement": 0,
|
| 142 |
-
"best_validation_loss": 1.024161856174469,
|
| 143 |
-
"improved": true
|
| 144 |
-
}
|
| 145 |
-
},
|
| 146 |
-
{
|
| 147 |
-
"epoch": 9,
|
| 148 |
-
"train_soft_cross_entropy": 0.9899537578335514,
|
| 149 |
-
"validation": {
|
| 150 |
-
"soft_cross_entropy": 1.0265704361597696,
|
| 151 |
-
"accuracy": 0.6083333333333333,
|
| 152 |
-
"brier_soft": 0.16808087141563496,
|
| 153 |
-
"decisions": 600
|
| 154 |
-
},
|
| 155 |
-
"early_stopping": {
|
| 156 |
-
"patience": 3,
|
| 157 |
-
"min_epochs": 10,
|
| 158 |
-
"epochs_without_improvement": 1,
|
| 159 |
-
"best_validation_loss": 1.024161856174469,
|
| 160 |
-
"improved": false
|
| 161 |
-
}
|
| 162 |
-
},
|
| 163 |
-
{
|
| 164 |
-
"epoch": 10,
|
| 165 |
-
"train_soft_cross_entropy": 0.9632394627288535,
|
| 166 |
-
"validation": {
|
| 167 |
-
"soft_cross_entropy": 1.0183856010437011,
|
| 168 |
-
"accuracy": 0.5983333333333334,
|
| 169 |
-
"brier_soft": 0.1608257148663203,
|
| 170 |
-
"decisions": 600
|
| 171 |
-
},
|
| 172 |
-
"early_stopping": {
|
| 173 |
-
"patience": 3,
|
| 174 |
-
"min_epochs": 10,
|
| 175 |
-
"epochs_without_improvement": 0,
|
| 176 |
-
"best_validation_loss": 1.0183856010437011,
|
| 177 |
-
"improved": true
|
| 178 |
-
}
|
| 179 |
-
},
|
| 180 |
-
{
|
| 181 |
-
"epoch": 11,
|
| 182 |
-
"train_soft_cross_entropy": 0.9342622081438701,
|
| 183 |
-
"validation": {
|
| 184 |
-
"soft_cross_entropy": 1.0024159566561381,
|
| 185 |
-
"accuracy": 0.6333333333333333,
|
| 186 |
-
"brier_soft": 0.14996731283764045,
|
| 187 |
-
"decisions": 600
|
| 188 |
-
},
|
| 189 |
-
"early_stopping": {
|
| 190 |
-
"patience": 3,
|
| 191 |
-
"min_epochs": 10,
|
| 192 |
-
"epochs_without_improvement": 0,
|
| 193 |
-
"best_validation_loss": 1.0024159566561381,
|
| 194 |
-
"improved": true
|
| 195 |
-
}
|
| 196 |
-
},
|
| 197 |
-
{
|
| 198 |
-
"epoch": 12,
|
| 199 |
-
"train_soft_cross_entropy": 0.9137524432606168,
|
| 200 |
-
"validation": {
|
| 201 |
-
"soft_cross_entropy": 0.9873138956228892,
|
| 202 |
-
"accuracy": 0.6516666666666666,
|
| 203 |
-
"brier_soft": 0.14277802929282188,
|
| 204 |
-
"decisions": 600
|
| 205 |
-
},
|
| 206 |
-
"early_stopping": {
|
| 207 |
-
"patience": 3,
|
| 208 |
-
"min_epochs": 10,
|
| 209 |
-
"epochs_without_improvement": 0,
|
| 210 |
-
"best_validation_loss": 0.9873138956228892,
|
| 211 |
-
"improved": true
|
| 212 |
-
}
|
| 213 |
-
},
|
| 214 |
-
{
|
| 215 |
-
"epoch": 13,
|
| 216 |
-
"train_soft_cross_entropy": 0.8881368139496556,
|
| 217 |
-
"validation": {
|
| 218 |
-
"soft_cross_entropy": 0.9626709421475729,
|
| 219 |
-
"accuracy": 0.6716666666666666,
|
| 220 |
-
"brier_soft": 0.12283751085400581,
|
| 221 |
-
"decisions": 600
|
| 222 |
-
},
|
| 223 |
-
"early_stopping": {
|
| 224 |
-
"patience": 3,
|
| 225 |
-
"min_epochs": 10,
|
| 226 |
-
"epochs_without_improvement": 0,
|
| 227 |
-
"best_validation_loss": 0.9626709421475729,
|
| 228 |
-
"improved": true
|
| 229 |
-
}
|
| 230 |
-
},
|
| 231 |
-
{
|
| 232 |
-
"epoch": 14,
|
| 233 |
-
"train_soft_cross_entropy": 0.865500467883216,
|
| 234 |
"validation": {
|
| 235 |
-
"soft_cross_entropy": 0.
|
| 236 |
-
"accuracy": 0.
|
| 237 |
-
"brier_soft": 0.
|
| 238 |
"decisions": 600
|
| 239 |
},
|
| 240 |
"early_stopping": {
|
| 241 |
"patience": 3,
|
| 242 |
"min_epochs": 10,
|
| 243 |
"epochs_without_improvement": 0,
|
| 244 |
-
"best_validation_loss": 0.
|
| 245 |
"improved": true
|
| 246 |
}
|
| 247 |
},
|
| 248 |
{
|
| 249 |
-
"epoch":
|
| 250 |
-
"train_soft_cross_entropy": 0.
|
| 251 |
-
"validation": {
|
| 252 |
-
"soft_cross_entropy": 0.9462178750832876,
|
| 253 |
-
"accuracy": 0.6916666666666667,
|
| 254 |
-
"brier_soft": 0.1187205430244406,
|
| 255 |
-
"decisions": 600
|
| 256 |
-
},
|
| 257 |
-
"early_stopping": {
|
| 258 |
-
"patience": 3,
|
| 259 |
-
"min_epochs": 10,
|
| 260 |
-
"epochs_without_improvement": 0,
|
| 261 |
-
"best_validation_loss": 0.9462178750832876,
|
| 262 |
-
"improved": true
|
| 263 |
-
}
|
| 264 |
-
},
|
| 265 |
-
{
|
| 266 |
-
"epoch": 16,
|
| 267 |
-
"train_soft_cross_entropy": 0.8394836834183446,
|
| 268 |
-
"validation": {
|
| 269 |
-
"soft_cross_entropy": 0.9463119049866994,
|
| 270 |
-
"accuracy": 0.6933333333333334,
|
| 271 |
-
"brier_soft": 0.11875337022046248,
|
| 272 |
-
"decisions": 600
|
| 273 |
-
},
|
| 274 |
-
"early_stopping": {
|
| 275 |
-
"patience": 3,
|
| 276 |
-
"min_epochs": 10,
|
| 277 |
-
"epochs_without_improvement": 1,
|
| 278 |
-
"best_validation_loss": 0.9462178750832876,
|
| 279 |
-
"improved": false
|
| 280 |
-
}
|
| 281 |
-
},
|
| 282 |
-
{
|
| 283 |
-
"epoch": 17,
|
| 284 |
-
"train_soft_cross_entropy": 0.8286281617040987,
|
| 285 |
-
"validation": {
|
| 286 |
-
"soft_cross_entropy": 0.941375896135966,
|
| 287 |
-
"accuracy": 0.685,
|
| 288 |
-
"brier_soft": 0.11455494280904531,
|
| 289 |
-
"decisions": 600
|
| 290 |
-
},
|
| 291 |
-
"early_stopping": {
|
| 292 |
-
"patience": 3,
|
| 293 |
-
"min_epochs": 10,
|
| 294 |
-
"epochs_without_improvement": 0,
|
| 295 |
-
"best_validation_loss": 0.941375896135966,
|
| 296 |
-
"improved": true
|
| 297 |
-
}
|
| 298 |
-
},
|
| 299 |
-
{
|
| 300 |
-
"epoch": 18,
|
| 301 |
-
"train_soft_cross_entropy": 0.8222652976601212,
|
| 302 |
-
"validation": {
|
| 303 |
-
"soft_cross_entropy": 0.9331588689486185,
|
| 304 |
-
"accuracy": 0.6866666666666666,
|
| 305 |
-
"brier_soft": 0.11237604923546314,
|
| 306 |
-
"decisions": 600
|
| 307 |
-
},
|
| 308 |
-
"early_stopping": {
|
| 309 |
-
"patience": 3,
|
| 310 |
-
"min_epochs": 10,
|
| 311 |
-
"epochs_without_improvement": 0,
|
| 312 |
-
"best_validation_loss": 0.9331588689486185,
|
| 313 |
-
"improved": true
|
| 314 |
-
}
|
| 315 |
-
},
|
| 316 |
-
{
|
| 317 |
-
"epoch": 19,
|
| 318 |
-
"train_soft_cross_entropy": 0.8160834203826056,
|
| 319 |
-
"validation": {
|
| 320 |
-
"soft_cross_entropy": 0.9387085942427317,
|
| 321 |
-
"accuracy": 0.6916666666666667,
|
| 322 |
-
"brier_soft": 0.11558652246991793,
|
| 323 |
-
"decisions": 600
|
| 324 |
-
},
|
| 325 |
-
"early_stopping": {
|
| 326 |
-
"patience": 3,
|
| 327 |
-
"min_epochs": 10,
|
| 328 |
-
"epochs_without_improvement": 1,
|
| 329 |
-
"best_validation_loss": 0.9331588689486185,
|
| 330 |
-
"improved": false
|
| 331 |
-
}
|
| 332 |
-
},
|
| 333 |
-
{
|
| 334 |
-
"epoch": 20,
|
| 335 |
-
"train_soft_cross_entropy": 0.8087538785846146,
|
| 336 |
-
"validation": {
|
| 337 |
-
"soft_cross_entropy": 0.9408066284656524,
|
| 338 |
-
"accuracy": 0.6783333333333333,
|
| 339 |
-
"brier_soft": 0.11930533437679211,
|
| 340 |
-
"decisions": 600
|
| 341 |
-
},
|
| 342 |
-
"early_stopping": {
|
| 343 |
-
"patience": 3,
|
| 344 |
-
"min_epochs": 10,
|
| 345 |
-
"epochs_without_improvement": 2,
|
| 346 |
-
"best_validation_loss": 0.9331588689486185,
|
| 347 |
-
"improved": false
|
| 348 |
-
}
|
| 349 |
-
},
|
| 350 |
-
{
|
| 351 |
-
"epoch": 21,
|
| 352 |
-
"train_soft_cross_entropy": 0.8048860243956248,
|
| 353 |
-
"validation": {
|
| 354 |
-
"soft_cross_entropy": 0.9236933688322703,
|
| 355 |
-
"accuracy": 0.7016666666666667,
|
| 356 |
-
"brier_soft": 0.10792215374608835,
|
| 357 |
-
"decisions": 600
|
| 358 |
-
},
|
| 359 |
-
"early_stopping": {
|
| 360 |
-
"patience": 3,
|
| 361 |
-
"min_epochs": 10,
|
| 362 |
-
"epochs_without_improvement": 0,
|
| 363 |
-
"best_validation_loss": 0.9236933688322703,
|
| 364 |
-
"improved": true
|
| 365 |
-
}
|
| 366 |
-
},
|
| 367 |
-
{
|
| 368 |
-
"epoch": 22,
|
| 369 |
-
"train_soft_cross_entropy": 0.8001213801790167,
|
| 370 |
"validation": {
|
| 371 |
-
"soft_cross_entropy": 0.
|
| 372 |
-
"accuracy": 0.
|
| 373 |
-
"brier_soft": 0.
|
| 374 |
"decisions": 600
|
| 375 |
},
|
| 376 |
"early_stopping": {
|
| 377 |
"patience": 3,
|
| 378 |
"min_epochs": 10,
|
| 379 |
"epochs_without_improvement": 1,
|
| 380 |
-
"best_validation_loss": 0.
|
| 381 |
"improved": false
|
| 382 |
}
|
| 383 |
},
|
| 384 |
{
|
| 385 |
-
"epoch":
|
| 386 |
-
"train_soft_cross_entropy": 0.
|
| 387 |
-
"validation": {
|
| 388 |
-
"soft_cross_entropy": 0.9299098292986552,
|
| 389 |
-
"accuracy": 0.7033333333333334,
|
| 390 |
-
"brier_soft": 0.11475709093113741,
|
| 391 |
-
"decisions": 600
|
| 392 |
-
},
|
| 393 |
-
"early_stopping": {
|
| 394 |
-
"patience": 3,
|
| 395 |
-
"min_epochs": 10,
|
| 396 |
-
"epochs_without_improvement": 2,
|
| 397 |
-
"best_validation_loss": 0.9236933688322703,
|
| 398 |
-
"improved": false
|
| 399 |
-
}
|
| 400 |
-
},
|
| 401 |
-
{
|
| 402 |
-
"epoch": 24,
|
| 403 |
-
"train_soft_cross_entropy": 0.7884197377717054,
|
| 404 |
"validation": {
|
| 405 |
-
"soft_cross_entropy": 0.
|
| 406 |
-
"accuracy": 0.
|
| 407 |
-
"brier_soft": 0.
|
| 408 |
"decisions": 600
|
| 409 |
},
|
| 410 |
"early_stopping": {
|
| 411 |
"patience": 3,
|
| 412 |
"min_epochs": 10,
|
| 413 |
"epochs_without_improvement": 0,
|
| 414 |
-
"best_validation_loss": 0.
|
| 415 |
"improved": true
|
| 416 |
}
|
| 417 |
},
|
| 418 |
{
|
| 419 |
-
"epoch":
|
| 420 |
-
"train_soft_cross_entropy": 0.
|
| 421 |
"validation": {
|
| 422 |
-
"soft_cross_entropy": 0.
|
| 423 |
-
"accuracy": 0.
|
| 424 |
-
"brier_soft": 0.
|
| 425 |
"decisions": 600
|
| 426 |
},
|
| 427 |
"early_stopping": {
|
| 428 |
"patience": 3,
|
| 429 |
"min_epochs": 10,
|
| 430 |
"epochs_without_improvement": 1,
|
| 431 |
-
"best_validation_loss": 0.
|
| 432 |
"improved": false
|
| 433 |
}
|
| 434 |
},
|
| 435 |
{
|
| 436 |
-
"epoch":
|
| 437 |
-
"train_soft_cross_entropy": 0.
|
| 438 |
"validation": {
|
| 439 |
-
"soft_cross_entropy": 0.
|
| 440 |
-
"accuracy": 0.
|
| 441 |
-
"brier_soft": 0.
|
| 442 |
"decisions": 600
|
| 443 |
},
|
| 444 |
"early_stopping": {
|
| 445 |
"patience": 3,
|
| 446 |
"min_epochs": 10,
|
| 447 |
"epochs_without_improvement": 2,
|
| 448 |
-
"best_validation_loss": 0.
|
| 449 |
"improved": false
|
| 450 |
}
|
| 451 |
},
|
| 452 |
{
|
| 453 |
-
"epoch":
|
| 454 |
-
"train_soft_cross_entropy": 0.
|
| 455 |
"validation": {
|
| 456 |
-
"soft_cross_entropy": 0.
|
| 457 |
-
"accuracy": 0.
|
| 458 |
-
"brier_soft": 0.
|
| 459 |
"decisions": 600
|
| 460 |
},
|
| 461 |
"early_stopping": {
|
| 462 |
"patience": 3,
|
| 463 |
"min_epochs": 10,
|
| 464 |
"epochs_without_improvement": 3,
|
| 465 |
-
"best_validation_loss": 0.
|
| 466 |
"improved": false
|
| 467 |
}
|
| 468 |
}
|
| 469 |
],
|
| 470 |
"stopping": {
|
| 471 |
"reason": "early_stopping",
|
| 472 |
-
"epochs_completed":
|
| 473 |
"patience": 3,
|
| 474 |
"min_epochs": 10,
|
| 475 |
"epochs_without_improvement": 3,
|
|
|
|
| 9 |
"epochs": [
|
| 10 |
{
|
| 11 |
"epoch": 1,
|
| 12 |
+
"train_soft_cross_entropy": 1.0557677305186237,
|
| 13 |
"validation": {
|
| 14 |
+
"soft_cross_entropy": 0.990565903186798,
|
| 15 |
+
"accuracy": 0.6166666666666667,
|
| 16 |
+
"brier_soft": 0.15243564940989018,
|
| 17 |
"decisions": 600
|
| 18 |
},
|
| 19 |
"early_stopping": {
|
| 20 |
"patience": 3,
|
| 21 |
"min_epochs": 10,
|
| 22 |
"epochs_without_improvement": 0,
|
| 23 |
+
"best_validation_loss": 0.990565903186798,
|
| 24 |
"improved": true
|
| 25 |
}
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"epoch": 2,
|
| 29 |
+
"train_soft_cross_entropy": 0.9286279639491328,
|
| 30 |
"validation": {
|
| 31 |
+
"soft_cross_entropy": 0.9022069962819418,
|
| 32 |
+
"accuracy": 0.71,
|
| 33 |
+
"brier_soft": 0.0947934572895368,
|
| 34 |
"decisions": 600
|
| 35 |
},
|
| 36 |
"early_stopping": {
|
| 37 |
"patience": 3,
|
| 38 |
"min_epochs": 10,
|
| 39 |
"epochs_without_improvement": 0,
|
| 40 |
+
"best_validation_loss": 0.9022069962819418,
|
| 41 |
"improved": true
|
| 42 |
}
|
| 43 |
},
|
| 44 |
{
|
| 45 |
"epoch": 3,
|
| 46 |
+
"train_soft_cross_entropy": 0.8668451248716424,
|
| 47 |
"validation": {
|
| 48 |
+
"soft_cross_entropy": 0.8912256610393524,
|
| 49 |
+
"accuracy": 0.725,
|
| 50 |
+
"brier_soft": 0.08988383966187637,
|
| 51 |
"decisions": 600
|
| 52 |
},
|
| 53 |
"early_stopping": {
|
| 54 |
"patience": 3,
|
| 55 |
"min_epochs": 10,
|
| 56 |
"epochs_without_improvement": 0,
|
| 57 |
+
"best_validation_loss": 0.8912256610393524,
|
| 58 |
"improved": true
|
| 59 |
}
|
| 60 |
},
|
| 61 |
{
|
| 62 |
"epoch": 4,
|
| 63 |
+
"train_soft_cross_entropy": 0.8375598757796817,
|
| 64 |
"validation": {
|
| 65 |
+
"soft_cross_entropy": 0.8716744474569956,
|
| 66 |
+
"accuracy": 0.7566666666666667,
|
| 67 |
+
"brier_soft": 0.07936330476154883,
|
| 68 |
"decisions": 600
|
| 69 |
},
|
| 70 |
"early_stopping": {
|
| 71 |
"patience": 3,
|
| 72 |
"min_epochs": 10,
|
| 73 |
"epochs_without_improvement": 0,
|
| 74 |
+
"best_validation_loss": 0.8716744474569956,
|
| 75 |
"improved": true
|
| 76 |
}
|
| 77 |
},
|
| 78 |
{
|
| 79 |
"epoch": 5,
|
| 80 |
+
"train_soft_cross_entropy": 0.8176034839064986,
|
| 81 |
"validation": {
|
| 82 |
+
"soft_cross_entropy": 0.8743925015131633,
|
| 83 |
+
"accuracy": 0.7333333333333333,
|
| 84 |
+
"brier_soft": 0.07994898026498655,
|
| 85 |
"decisions": 600
|
| 86 |
},
|
| 87 |
"early_stopping": {
|
| 88 |
"patience": 3,
|
| 89 |
"min_epochs": 10,
|
| 90 |
"epochs_without_improvement": 1,
|
| 91 |
+
"best_validation_loss": 0.8716744474569956,
|
| 92 |
"improved": false
|
| 93 |
}
|
| 94 |
},
|
| 95 |
{
|
| 96 |
"epoch": 6,
|
| 97 |
+
"train_soft_cross_entropy": 0.8031262465318044,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 98 |
"validation": {
|
| 99 |
+
"soft_cross_entropy": 0.8599367336432139,
|
| 100 |
+
"accuracy": 0.76,
|
| 101 |
+
"brier_soft": 0.07330287167181572,
|
| 102 |
"decisions": 600
|
| 103 |
},
|
| 104 |
"early_stopping": {
|
| 105 |
"patience": 3,
|
| 106 |
"min_epochs": 10,
|
| 107 |
"epochs_without_improvement": 0,
|
| 108 |
+
"best_validation_loss": 0.8599367336432139,
|
| 109 |
"improved": true
|
| 110 |
}
|
| 111 |
},
|
| 112 |
{
|
| 113 |
+
"epoch": 7,
|
| 114 |
+
"train_soft_cross_entropy": 0.7926116739820551,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 115 |
"validation": {
|
| 116 |
+
"soft_cross_entropy": 0.8638877793153127,
|
| 117 |
+
"accuracy": 0.7566666666666667,
|
| 118 |
+
"brier_soft": 0.07329062350094319,
|
| 119 |
"decisions": 600
|
| 120 |
},
|
| 121 |
"early_stopping": {
|
| 122 |
"patience": 3,
|
| 123 |
"min_epochs": 10,
|
| 124 |
"epochs_without_improvement": 1,
|
| 125 |
+
"best_validation_loss": 0.8599367336432139,
|
| 126 |
"improved": false
|
| 127 |
}
|
| 128 |
},
|
| 129 |
{
|
| 130 |
+
"epoch": 8,
|
| 131 |
+
"train_soft_cross_entropy": 0.7867316756866596,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 132 |
"validation": {
|
| 133 |
+
"soft_cross_entropy": 0.8545633804798126,
|
| 134 |
+
"accuracy": 0.77,
|
| 135 |
+
"brier_soft": 0.06962224062221746,
|
| 136 |
"decisions": 600
|
| 137 |
},
|
| 138 |
"early_stopping": {
|
| 139 |
"patience": 3,
|
| 140 |
"min_epochs": 10,
|
| 141 |
"epochs_without_improvement": 0,
|
| 142 |
+
"best_validation_loss": 0.8545633804798126,
|
| 143 |
"improved": true
|
| 144 |
}
|
| 145 |
},
|
| 146 |
{
|
| 147 |
+
"epoch": 9,
|
| 148 |
+
"train_soft_cross_entropy": 0.7820918090696688,
|
| 149 |
"validation": {
|
| 150 |
+
"soft_cross_entropy": 0.8555870747566223,
|
| 151 |
+
"accuracy": 0.7683333333333333,
|
| 152 |
+
"brier_soft": 0.06890943528773884,
|
| 153 |
"decisions": 600
|
| 154 |
},
|
| 155 |
"early_stopping": {
|
| 156 |
"patience": 3,
|
| 157 |
"min_epochs": 10,
|
| 158 |
"epochs_without_improvement": 1,
|
| 159 |
+
"best_validation_loss": 0.8545633804798126,
|
| 160 |
"improved": false
|
| 161 |
}
|
| 162 |
},
|
| 163 |
{
|
| 164 |
+
"epoch": 10,
|
| 165 |
+
"train_soft_cross_entropy": 0.779588269745862,
|
| 166 |
"validation": {
|
| 167 |
+
"soft_cross_entropy": 0.8600131020943323,
|
| 168 |
+
"accuracy": 0.76,
|
| 169 |
+
"brier_soft": 0.07248457937811811,
|
| 170 |
"decisions": 600
|
| 171 |
},
|
| 172 |
"early_stopping": {
|
| 173 |
"patience": 3,
|
| 174 |
"min_epochs": 10,
|
| 175 |
"epochs_without_improvement": 2,
|
| 176 |
+
"best_validation_loss": 0.8545633804798126,
|
| 177 |
"improved": false
|
| 178 |
}
|
| 179 |
},
|
| 180 |
{
|
| 181 |
+
"epoch": 11,
|
| 182 |
+
"train_soft_cross_entropy": 0.7761333607302772,
|
| 183 |
"validation": {
|
| 184 |
+
"soft_cross_entropy": 0.8604028668006262,
|
| 185 |
+
"accuracy": 0.7616666666666667,
|
| 186 |
+
"brier_soft": 0.07338622493979831,
|
| 187 |
"decisions": 600
|
| 188 |
},
|
| 189 |
"early_stopping": {
|
| 190 |
"patience": 3,
|
| 191 |
"min_epochs": 10,
|
| 192 |
"epochs_without_improvement": 3,
|
| 193 |
+
"best_validation_loss": 0.8545633804798126,
|
| 194 |
"improved": false
|
| 195 |
}
|
| 196 |
}
|
| 197 |
],
|
| 198 |
"stopping": {
|
| 199 |
"reason": "early_stopping",
|
| 200 |
+
"epochs_completed": 11,
|
| 201 |
"patience": 3,
|
| 202 |
"min_epochs": 10,
|
| 203 |
"epochs_without_improvement": 3,
|
evidence/seed-4-history.json
CHANGED
|
@@ -9,178 +9,195 @@
|
|
| 9 |
"epochs": [
|
| 10 |
{
|
| 11 |
"epoch": 1,
|
| 12 |
-
"train_soft_cross_entropy": 1.
|
| 13 |
"validation": {
|
| 14 |
-
"soft_cross_entropy": 0.
|
| 15 |
-
"accuracy": 0.
|
| 16 |
-
"brier_soft": 0.
|
| 17 |
"decisions": 600
|
| 18 |
},
|
| 19 |
"early_stopping": {
|
| 20 |
"patience": 3,
|
| 21 |
"min_epochs": 10,
|
| 22 |
"epochs_without_improvement": 0,
|
| 23 |
-
"best_validation_loss": 0.
|
| 24 |
"improved": true
|
| 25 |
}
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"epoch": 2,
|
| 29 |
-
"train_soft_cross_entropy": 0.
|
| 30 |
"validation": {
|
| 31 |
-
"soft_cross_entropy": 0.
|
| 32 |
-
"accuracy": 0.
|
| 33 |
-
"brier_soft": 0.
|
| 34 |
"decisions": 600
|
| 35 |
},
|
| 36 |
"early_stopping": {
|
| 37 |
"patience": 3,
|
| 38 |
"min_epochs": 10,
|
| 39 |
"epochs_without_improvement": 0,
|
| 40 |
-
"best_validation_loss": 0.
|
| 41 |
"improved": true
|
| 42 |
}
|
| 43 |
},
|
| 44 |
{
|
| 45 |
"epoch": 3,
|
| 46 |
-
"train_soft_cross_entropy": 0.
|
| 47 |
"validation": {
|
| 48 |
-
"soft_cross_entropy": 0.
|
| 49 |
-
"accuracy": 0.
|
| 50 |
-
"brier_soft": 0.
|
| 51 |
"decisions": 600
|
| 52 |
},
|
| 53 |
"early_stopping": {
|
| 54 |
"patience": 3,
|
| 55 |
"min_epochs": 10,
|
| 56 |
"epochs_without_improvement": 0,
|
| 57 |
-
"best_validation_loss": 0.
|
| 58 |
"improved": true
|
| 59 |
}
|
| 60 |
},
|
| 61 |
{
|
| 62 |
"epoch": 4,
|
| 63 |
-
"train_soft_cross_entropy": 0.
|
| 64 |
"validation": {
|
| 65 |
-
"soft_cross_entropy": 0.
|
| 66 |
-
"accuracy": 0.
|
| 67 |
-
"brier_soft": 0.
|
| 68 |
"decisions": 600
|
| 69 |
},
|
| 70 |
"early_stopping": {
|
| 71 |
"patience": 3,
|
| 72 |
"min_epochs": 10,
|
| 73 |
"epochs_without_improvement": 0,
|
| 74 |
-
"best_validation_loss": 0.
|
| 75 |
"improved": true
|
| 76 |
}
|
| 77 |
},
|
| 78 |
{
|
| 79 |
"epoch": 5,
|
| 80 |
-
"train_soft_cross_entropy": 0.
|
| 81 |
"validation": {
|
| 82 |
-
"soft_cross_entropy": 0.
|
| 83 |
-
"accuracy": 0.
|
| 84 |
-
"brier_soft": 0.
|
| 85 |
"decisions": 600
|
| 86 |
},
|
| 87 |
"early_stopping": {
|
| 88 |
"patience": 3,
|
| 89 |
"min_epochs": 10,
|
| 90 |
-
"epochs_without_improvement":
|
| 91 |
-
"best_validation_loss": 0.
|
| 92 |
-
"improved":
|
| 93 |
}
|
| 94 |
},
|
| 95 |
{
|
| 96 |
"epoch": 6,
|
| 97 |
-
"train_soft_cross_entropy": 0.
|
| 98 |
"validation": {
|
| 99 |
-
"soft_cross_entropy": 0.
|
| 100 |
-
"accuracy": 0.
|
| 101 |
-
"brier_soft": 0.
|
| 102 |
"decisions": 600
|
| 103 |
},
|
| 104 |
"early_stopping": {
|
| 105 |
"patience": 3,
|
| 106 |
"min_epochs": 10,
|
| 107 |
-
"epochs_without_improvement":
|
| 108 |
-
"best_validation_loss": 0.
|
| 109 |
-
"improved":
|
| 110 |
}
|
| 111 |
},
|
| 112 |
{
|
| 113 |
"epoch": 7,
|
| 114 |
-
"train_soft_cross_entropy": 0.
|
| 115 |
"validation": {
|
| 116 |
-
"soft_cross_entropy": 0.
|
| 117 |
-
"accuracy": 0.
|
| 118 |
-
"brier_soft": 0.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 119 |
"decisions": 600
|
| 120 |
},
|
| 121 |
"early_stopping": {
|
| 122 |
"patience": 3,
|
| 123 |
"min_epochs": 10,
|
| 124 |
"epochs_without_improvement": 0,
|
| 125 |
-
"best_validation_loss": 0.
|
| 126 |
"improved": true
|
| 127 |
}
|
| 128 |
},
|
| 129 |
{
|
| 130 |
-
"epoch":
|
| 131 |
-
"train_soft_cross_entropy": 0.
|
| 132 |
"validation": {
|
| 133 |
-
"soft_cross_entropy": 0.
|
| 134 |
-
"accuracy": 0.
|
| 135 |
-
"brier_soft": 0.
|
| 136 |
"decisions": 600
|
| 137 |
},
|
| 138 |
"early_stopping": {
|
| 139 |
"patience": 3,
|
| 140 |
"min_epochs": 10,
|
| 141 |
"epochs_without_improvement": 1,
|
| 142 |
-
"best_validation_loss": 0.
|
| 143 |
"improved": false
|
| 144 |
}
|
| 145 |
},
|
| 146 |
{
|
| 147 |
-
"epoch":
|
| 148 |
-
"train_soft_cross_entropy": 0.
|
| 149 |
"validation": {
|
| 150 |
-
"soft_cross_entropy": 0.
|
| 151 |
-
"accuracy": 0.
|
| 152 |
-
"brier_soft": 0.
|
| 153 |
"decisions": 600
|
| 154 |
},
|
| 155 |
"early_stopping": {
|
| 156 |
"patience": 3,
|
| 157 |
"min_epochs": 10,
|
| 158 |
"epochs_without_improvement": 2,
|
| 159 |
-
"best_validation_loss": 0.
|
| 160 |
"improved": false
|
| 161 |
}
|
| 162 |
},
|
| 163 |
{
|
| 164 |
-
"epoch":
|
| 165 |
-
"train_soft_cross_entropy": 0.
|
| 166 |
"validation": {
|
| 167 |
-
"soft_cross_entropy": 0.
|
| 168 |
-
"accuracy": 0.
|
| 169 |
-
"brier_soft": 0.
|
| 170 |
"decisions": 600
|
| 171 |
},
|
| 172 |
"early_stopping": {
|
| 173 |
"patience": 3,
|
| 174 |
"min_epochs": 10,
|
| 175 |
"epochs_without_improvement": 3,
|
| 176 |
-
"best_validation_loss": 0.
|
| 177 |
"improved": false
|
| 178 |
}
|
| 179 |
}
|
| 180 |
],
|
| 181 |
"stopping": {
|
| 182 |
"reason": "early_stopping",
|
| 183 |
-
"epochs_completed":
|
| 184 |
"patience": 3,
|
| 185 |
"min_epochs": 10,
|
| 186 |
"epochs_without_improvement": 3,
|
|
|
|
| 9 |
"epochs": [
|
| 10 |
{
|
| 11 |
"epoch": 1,
|
| 12 |
+
"train_soft_cross_entropy": 1.022522153324551,
|
| 13 |
"validation": {
|
| 14 |
+
"soft_cross_entropy": 0.9845864299933116,
|
| 15 |
+
"accuracy": 0.6683333333333333,
|
| 16 |
+
"brier_soft": 0.14738356913129488,
|
| 17 |
"decisions": 600
|
| 18 |
},
|
| 19 |
"early_stopping": {
|
| 20 |
"patience": 3,
|
| 21 |
"min_epochs": 10,
|
| 22 |
"epochs_without_improvement": 0,
|
| 23 |
+
"best_validation_loss": 0.9845864299933116,
|
| 24 |
"improved": true
|
| 25 |
}
|
| 26 |
},
|
| 27 |
{
|
| 28 |
"epoch": 2,
|
| 29 |
+
"train_soft_cross_entropy": 0.9041571515136295,
|
| 30 |
"validation": {
|
| 31 |
+
"soft_cross_entropy": 0.9089026947816213,
|
| 32 |
+
"accuracy": 0.7216666666666667,
|
| 33 |
+
"brier_soft": 0.10058322168886662,
|
| 34 |
"decisions": 600
|
| 35 |
},
|
| 36 |
"early_stopping": {
|
| 37 |
"patience": 3,
|
| 38 |
"min_epochs": 10,
|
| 39 |
"epochs_without_improvement": 0,
|
| 40 |
+
"best_validation_loss": 0.9089026947816213,
|
| 41 |
"improved": true
|
| 42 |
}
|
| 43 |
},
|
| 44 |
{
|
| 45 |
"epoch": 3,
|
| 46 |
+
"train_soft_cross_entropy": 0.8495574627099214,
|
| 47 |
"validation": {
|
| 48 |
+
"soft_cross_entropy": 0.8853352854649226,
|
| 49 |
+
"accuracy": 0.74,
|
| 50 |
+
"brier_soft": 0.08744065261135499,
|
| 51 |
"decisions": 600
|
| 52 |
},
|
| 53 |
"early_stopping": {
|
| 54 |
"patience": 3,
|
| 55 |
"min_epochs": 10,
|
| 56 |
"epochs_without_improvement": 0,
|
| 57 |
+
"best_validation_loss": 0.8853352854649226,
|
| 58 |
"improved": true
|
| 59 |
}
|
| 60 |
},
|
| 61 |
{
|
| 62 |
"epoch": 4,
|
| 63 |
+
"train_soft_cross_entropy": 0.8263491551522856,
|
| 64 |
"validation": {
|
| 65 |
+
"soft_cross_entropy": 0.8789768069982529,
|
| 66 |
+
"accuracy": 0.7416666666666667,
|
| 67 |
+
"brier_soft": 0.08262864720076323,
|
| 68 |
"decisions": 600
|
| 69 |
},
|
| 70 |
"early_stopping": {
|
| 71 |
"patience": 3,
|
| 72 |
"min_epochs": 10,
|
| 73 |
"epochs_without_improvement": 0,
|
| 74 |
+
"best_validation_loss": 0.8789768069982529,
|
| 75 |
"improved": true
|
| 76 |
}
|
| 77 |
},
|
| 78 |
{
|
| 79 |
"epoch": 5,
|
| 80 |
+
"train_soft_cross_entropy": 0.8096302935812209,
|
| 81 |
"validation": {
|
| 82 |
+
"soft_cross_entropy": 0.864648152589798,
|
| 83 |
+
"accuracy": 0.73,
|
| 84 |
+
"brier_soft": 0.07628292332092922,
|
| 85 |
"decisions": 600
|
| 86 |
},
|
| 87 |
"early_stopping": {
|
| 88 |
"patience": 3,
|
| 89 |
"min_epochs": 10,
|
| 90 |
+
"epochs_without_improvement": 0,
|
| 91 |
+
"best_validation_loss": 0.864648152589798,
|
| 92 |
+
"improved": true
|
| 93 |
}
|
| 94 |
},
|
| 95 |
{
|
| 96 |
"epoch": 6,
|
| 97 |
+
"train_soft_cross_entropy": 0.7991735404950601,
|
| 98 |
"validation": {
|
| 99 |
+
"soft_cross_entropy": 0.8743107922871908,
|
| 100 |
+
"accuracy": 0.74,
|
| 101 |
+
"brier_soft": 0.08211790287246307,
|
| 102 |
"decisions": 600
|
| 103 |
},
|
| 104 |
"early_stopping": {
|
| 105 |
"patience": 3,
|
| 106 |
"min_epochs": 10,
|
| 107 |
+
"epochs_without_improvement": 1,
|
| 108 |
+
"best_validation_loss": 0.864648152589798,
|
| 109 |
+
"improved": false
|
| 110 |
}
|
| 111 |
},
|
| 112 |
{
|
| 113 |
"epoch": 7,
|
| 114 |
+
"train_soft_cross_entropy": 0.7918641125714337,
|
| 115 |
"validation": {
|
| 116 |
+
"soft_cross_entropy": 0.86731658577919,
|
| 117 |
+
"accuracy": 0.7483333333333333,
|
| 118 |
+
"brier_soft": 0.07722779513647159,
|
| 119 |
+
"decisions": 600
|
| 120 |
+
},
|
| 121 |
+
"early_stopping": {
|
| 122 |
+
"patience": 3,
|
| 123 |
+
"min_epochs": 10,
|
| 124 |
+
"epochs_without_improvement": 2,
|
| 125 |
+
"best_validation_loss": 0.864648152589798,
|
| 126 |
+
"improved": false
|
| 127 |
+
}
|
| 128 |
+
},
|
| 129 |
+
{
|
| 130 |
+
"epoch": 8,
|
| 131 |
+
"train_soft_cross_entropy": 0.7861463651833711,
|
| 132 |
+
"validation": {
|
| 133 |
+
"soft_cross_entropy": 0.8540202794472377,
|
| 134 |
+
"accuracy": 0.76,
|
| 135 |
+
"brier_soft": 0.07019127607345581,
|
| 136 |
"decisions": 600
|
| 137 |
},
|
| 138 |
"early_stopping": {
|
| 139 |
"patience": 3,
|
| 140 |
"min_epochs": 10,
|
| 141 |
"epochs_without_improvement": 0,
|
| 142 |
+
"best_validation_loss": 0.8540202794472377,
|
| 143 |
"improved": true
|
| 144 |
}
|
| 145 |
},
|
| 146 |
{
|
| 147 |
+
"epoch": 9,
|
| 148 |
+
"train_soft_cross_entropy": 0.7806310887248428,
|
| 149 |
"validation": {
|
| 150 |
+
"soft_cross_entropy": 0.8646407808860143,
|
| 151 |
+
"accuracy": 0.735,
|
| 152 |
+
"brier_soft": 0.07658187456429005,
|
| 153 |
"decisions": 600
|
| 154 |
},
|
| 155 |
"early_stopping": {
|
| 156 |
"patience": 3,
|
| 157 |
"min_epochs": 10,
|
| 158 |
"epochs_without_improvement": 1,
|
| 159 |
+
"best_validation_loss": 0.8540202794472377,
|
| 160 |
"improved": false
|
| 161 |
}
|
| 162 |
},
|
| 163 |
{
|
| 164 |
+
"epoch": 10,
|
| 165 |
+
"train_soft_cross_entropy": 0.7768384672094274,
|
| 166 |
"validation": {
|
| 167 |
+
"soft_cross_entropy": 0.8618478431304296,
|
| 168 |
+
"accuracy": 0.73,
|
| 169 |
+
"brier_soft": 0.07576349244763454,
|
| 170 |
"decisions": 600
|
| 171 |
},
|
| 172 |
"early_stopping": {
|
| 173 |
"patience": 3,
|
| 174 |
"min_epochs": 10,
|
| 175 |
"epochs_without_improvement": 2,
|
| 176 |
+
"best_validation_loss": 0.8540202794472377,
|
| 177 |
"improved": false
|
| 178 |
}
|
| 179 |
},
|
| 180 |
{
|
| 181 |
+
"epoch": 11,
|
| 182 |
+
"train_soft_cross_entropy": 0.7740854684511821,
|
| 183 |
"validation": {
|
| 184 |
+
"soft_cross_entropy": 0.8627219025293986,
|
| 185 |
+
"accuracy": 0.7283333333333334,
|
| 186 |
+
"brier_soft": 0.07741292332919936,
|
| 187 |
"decisions": 600
|
| 188 |
},
|
| 189 |
"early_stopping": {
|
| 190 |
"patience": 3,
|
| 191 |
"min_epochs": 10,
|
| 192 |
"epochs_without_improvement": 3,
|
| 193 |
+
"best_validation_loss": 0.8540202794472377,
|
| 194 |
"improved": false
|
| 195 |
}
|
| 196 |
}
|
| 197 |
],
|
| 198 |
"stopping": {
|
| 199 |
"reason": "early_stopping",
|
| 200 |
+
"epochs_completed": 11,
|
| 201 |
"patience": 3,
|
| 202 |
"min_epochs": 10,
|
| 203 |
"epochs_without_improvement": 3,
|
export-verification.json
CHANGED
|
@@ -1,11 +1,11 @@
|
|
| 1 |
{
|
| 2 |
"passed": true,
|
| 3 |
-
"candidate_seed":
|
| 4 |
-
"selected_epoch":
|
| 5 |
"cases": 3516,
|
| 6 |
"decisions": 5116,
|
| 7 |
"benchmark_cases_sha256": "10fb671245cdac1ee09e9c4d2b9b09829fbf1abd06e0954fddee0b9ec8519854",
|
| 8 |
-
"checkpoint_sha256": "
|
| 9 |
"all_answer_objects_exact": true,
|
| 10 |
"all_suite_metrics_exact": true,
|
| 11 |
"bundled_code_imported_from_outside_workspace": true,
|
|
@@ -35,8 +35,8 @@
|
|
| 35 |
"portable_adapter": {
|
| 36 |
"base_model": "convaiinnovations/laya",
|
| 37 |
"base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
|
| 38 |
-
"seed":
|
| 39 |
-
"selected_epoch":
|
| 40 |
}
|
| 41 |
}
|
| 42 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"passed": true,
|
| 3 |
+
"candidate_seed": 4,
|
| 4 |
+
"selected_epoch": 8,
|
| 5 |
"cases": 3516,
|
| 6 |
"decisions": 5116,
|
| 7 |
"benchmark_cases_sha256": "10fb671245cdac1ee09e9c4d2b9b09829fbf1abd06e0954fddee0b9ec8519854",
|
| 8 |
+
"checkpoint_sha256": "f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081",
|
| 9 |
"all_answer_objects_exact": true,
|
| 10 |
"all_suite_metrics_exact": true,
|
| 11 |
"bundled_code_imported_from_outside_workspace": true,
|
|
|
|
| 35 |
"portable_adapter": {
|
| 36 |
"base_model": "convaiinnovations/laya",
|
| 37 |
"base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
|
| 38 |
+
"seed": 4,
|
| 39 |
+
"selected_epoch": 8
|
| 40 |
}
|
| 41 |
}
|
| 42 |
}
|
make_training_config.py
ADDED
|
@@ -0,0 +1,37 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Resolve the pinned base checkpoint and create a reproducible training config."""
|
| 2 |
+
|
| 3 |
+
import argparse
|
| 4 |
+
import json
|
| 5 |
+
from pathlib import Path
|
| 6 |
+
from huggingface_hub import snapshot_download
|
| 7 |
+
|
| 8 |
+
parser = argparse.ArgumentParser(description=__doc__)
|
| 9 |
+
parser.add_argument("--seed", type=int, choices=range(5), required=True)
|
| 10 |
+
parser.add_argument("--device", default="cuda:0")
|
| 11 |
+
parser.add_argument("--out", required=True)
|
| 12 |
+
args = parser.parse_args()
|
| 13 |
+
root = Path(__file__).resolve().parent
|
| 14 |
+
metadata = json.loads((root / "adapter_config.json").read_text())
|
| 15 |
+
base = snapshot_download(
|
| 16 |
+
metadata["base_model"],
|
| 17 |
+
revision=metadata["base_revision"],
|
| 18 |
+
allow_patterns=["model.safetensors", "rl_agent_config.json", "encoder/*", "tokenizer/*"],
|
| 19 |
+
)
|
| 20 |
+
config = {
|
| 21 |
+
"name": f"Deterministic frozen base + linear, seed {args.seed} (UNTRAINED)",
|
| 22 |
+
"adapter": "python",
|
| 23 |
+
"mode": "reference",
|
| 24 |
+
"factory": "ariadne_bench.frozen_input_interface:create",
|
| 25 |
+
"options": {
|
| 26 |
+
"laya_model": str(Path(base).resolve()),
|
| 27 |
+
"device": args.device,
|
| 28 |
+
"batch_size": 16,
|
| 29 |
+
"max_len": 1024,
|
| 30 |
+
"head_max_len": 256,
|
| 31 |
+
"seed": args.seed,
|
| 32 |
+
"training_scope": "bridge",
|
| 33 |
+
"add_interface": True,
|
| 34 |
+
"deterministic": True,
|
| 35 |
+
},
|
| 36 |
+
}
|
| 37 |
+
Path(args.out).write_text(json.dumps(config, indent=2) + "\n")
|
metrics.json
CHANGED
|
@@ -12,8 +12,8 @@
|
|
| 12 |
"runs": {
|
| 13 |
"0": {
|
| 14 |
"seed": 0,
|
| 15 |
-
"selected_epoch":
|
| 16 |
-
"elapsed_s":
|
| 17 |
"initial_validation": {
|
| 18 |
"soft_cross_entropy": 1.544869564374288,
|
| 19 |
"accuracy": 0.36833333333333335,
|
|
@@ -21,14 +21,14 @@
|
|
| 21 |
"decisions": 600
|
| 22 |
},
|
| 23 |
"selected_validation": {
|
| 24 |
-
"soft_cross_entropy":
|
| 25 |
-
"accuracy": 0.
|
| 26 |
-
"brier_soft": 0.
|
| 27 |
"decisions": 600
|
| 28 |
},
|
| 29 |
"stopping": {
|
| 30 |
"reason": "early_stopping",
|
| 31 |
-
"epochs_completed":
|
| 32 |
"patience": 3,
|
| 33 |
"min_epochs": 10,
|
| 34 |
"epochs_without_improvement": 3,
|
|
@@ -42,159 +42,159 @@
|
|
| 42 |
"valid": 2000,
|
| 43 |
"failed": 0,
|
| 44 |
"coverage": 1.0,
|
| 45 |
-
"accuracy_all": 0.
|
| 46 |
-
"accuracy_valid": 0.
|
| 47 |
-
"ece_top_label": 0.
|
| 48 |
-
"mean_confidence": 0.
|
| 49 |
-
"brier_hard": 0.
|
| 50 |
"brier_hard_n": 2000,
|
| 51 |
-
"nll_hard": 0.
|
| 52 |
"nll_hard_n": 2000,
|
| 53 |
"zero_probability_gold": 0.0,
|
| 54 |
"zero_probability_gold_n": 2000,
|
| 55 |
-
"soft_accuracy": 0.
|
| 56 |
"soft_accuracy_n": 2000,
|
| 57 |
-
"brier_soft": 0.
|
| 58 |
"brier_soft_n": 2000,
|
| 59 |
-
"kl_gold_to_prediction": 0.
|
| 60 |
"kl_gold_to_prediction_n": 2000,
|
| 61 |
-
"total_variation": 0.
|
| 62 |
"total_variation_n": 2000,
|
| 63 |
-
"score_mae": 0.
|
| 64 |
"score_mae_n": 800,
|
| 65 |
-
"within_one_level": 0.
|
| 66 |
"within_one_level_n": 800,
|
| 67 |
-
"macro_f1": 0.
|
| 68 |
},
|
| 69 |
"ag-news": {
|
| 70 |
"attempted": 600,
|
| 71 |
"valid": 600,
|
| 72 |
"failed": 0,
|
| 73 |
"coverage": 1.0,
|
| 74 |
-
"accuracy_all": 0.
|
| 75 |
-
"accuracy_valid": 0.
|
| 76 |
-
"ece_top_label": 0.
|
| 77 |
-
"mean_confidence": 0.
|
| 78 |
-
"brier_hard": 0.
|
| 79 |
"brier_hard_n": 600,
|
| 80 |
-
"nll_hard":
|
| 81 |
"nll_hard_n": 600,
|
| 82 |
"zero_probability_gold": 0.0,
|
| 83 |
"zero_probability_gold_n": 600,
|
| 84 |
-
"macro_f1": 0.
|
| 85 |
},
|
| 86 |
"emotion": {
|
| 87 |
"attempted": 600,
|
| 88 |
"valid": 600,
|
| 89 |
"failed": 0,
|
| 90 |
"coverage": 1.0,
|
| 91 |
-
"accuracy_all": 0.
|
| 92 |
-
"accuracy_valid": 0.
|
| 93 |
-
"ece_top_label": 0.
|
| 94 |
-
"mean_confidence": 0.
|
| 95 |
-
"brier_hard": 0.
|
| 96 |
"brier_hard_n": 600,
|
| 97 |
-
"nll_hard":
|
| 98 |
"nll_hard_n": 600,
|
| 99 |
-
"zero_probability_gold": 0.
|
| 100 |
"zero_probability_gold_n": 600,
|
| 101 |
-
"macro_f1": 0.
|
| 102 |
},
|
| 103 |
"boolq": {
|
| 104 |
"attempted": 600,
|
| 105 |
"valid": 600,
|
| 106 |
"failed": 0,
|
| 107 |
"coverage": 1.0,
|
| 108 |
-
"accuracy_all": 0.
|
| 109 |
-
"accuracy_valid": 0.
|
| 110 |
-
"ece_top_label": 0.
|
| 111 |
-
"mean_confidence": 0.
|
| 112 |
-
"brier_hard": 0.
|
| 113 |
"brier_hard_n": 600,
|
| 114 |
-
"nll_hard": 0.
|
| 115 |
"nll_hard_n": 600,
|
| 116 |
"zero_probability_gold": 0.0,
|
| 117 |
"zero_probability_gold_n": 600,
|
| 118 |
-
"macro_f1": 0.
|
| 119 |
},
|
| 120 |
"sst5": {
|
| 121 |
"attempted": 600,
|
| 122 |
"valid": 600,
|
| 123 |
"failed": 0,
|
| 124 |
"coverage": 1.0,
|
| 125 |
-
"accuracy_all": 0.
|
| 126 |
-
"accuracy_valid": 0.
|
| 127 |
-
"ece_top_label": 0.
|
| 128 |
-
"mean_confidence": 0.
|
| 129 |
-
"brier_hard": 0.
|
| 130 |
"brier_hard_n": 600,
|
| 131 |
-
"nll_hard": 1.
|
| 132 |
"nll_hard_n": 600,
|
| 133 |
"zero_probability_gold": 0.0,
|
| 134 |
"zero_probability_gold_n": 600,
|
| 135 |
-
"score_mae":
|
| 136 |
"score_mae_n": 600,
|
| 137 |
-
"within_one_level": 0.
|
| 138 |
"within_one_level_n": 600,
|
| 139 |
-
"macro_f1": 0.
|
| 140 |
},
|
| 141 |
"prompt-injections": {
|
| 142 |
"attempted": 116,
|
| 143 |
"valid": 116,
|
| 144 |
"failed": 0,
|
| 145 |
"coverage": 1.0,
|
| 146 |
-
"accuracy_all": 0.
|
| 147 |
-
"accuracy_valid": 0.
|
| 148 |
-
"ece_top_label": 0.
|
| 149 |
-
"mean_confidence": 0.
|
| 150 |
-
"brier_hard": 0.
|
| 151 |
"brier_hard_n": 116,
|
| 152 |
-
"nll_hard":
|
| 153 |
"nll_hard_n": 116,
|
| 154 |
-
"zero_probability_gold": 0.
|
| 155 |
"zero_probability_gold_n": 116,
|
| 156 |
-
"macro_f1": 0.
|
| 157 |
},
|
| 158 |
"massive-intent.en": {
|
| 159 |
"attempted": 300,
|
| 160 |
"valid": 300,
|
| 161 |
"failed": 0,
|
| 162 |
"coverage": 1.0,
|
| 163 |
-
"accuracy_all": 0.
|
| 164 |
-
"accuracy_valid": 0.
|
| 165 |
-
"ece_top_label": 0.
|
| 166 |
-
"mean_confidence": 0.
|
| 167 |
-
"brier_hard": 0.
|
| 168 |
"brier_hard_n": 300,
|
| 169 |
-
"nll_hard":
|
| 170 |
"nll_hard_n": 300,
|
| 171 |
-
"zero_probability_gold": 0.
|
| 172 |
"zero_probability_gold_n": 300,
|
| 173 |
-
"macro_f1": 0.
|
| 174 |
},
|
| 175 |
"xnli.en": {
|
| 176 |
"attempted": 300,
|
| 177 |
"valid": 300,
|
| 178 |
"failed": 0,
|
| 179 |
"coverage": 1.0,
|
| 180 |
-
"accuracy_all": 0.
|
| 181 |
-
"accuracy_valid": 0.
|
| 182 |
-
"ece_top_label": 0.
|
| 183 |
-
"mean_confidence": 0.
|
| 184 |
-
"brier_hard": 0.
|
| 185 |
"brier_hard_n": 300,
|
| 186 |
-
"nll_hard":
|
| 187 |
"nll_hard_n": 300,
|
| 188 |
"zero_probability_gold": 0.0,
|
| 189 |
"zero_probability_gold_n": 300,
|
| 190 |
-
"macro_f1": 0.
|
| 191 |
}
|
| 192 |
}
|
| 193 |
},
|
| 194 |
"1": {
|
| 195 |
"seed": 1,
|
| 196 |
-
"selected_epoch":
|
| 197 |
-
"elapsed_s":
|
| 198 |
"initial_validation": {
|
| 199 |
"soft_cross_entropy": 1.544869564374288,
|
| 200 |
"accuracy": 0.36833333333333335,
|
|
@@ -202,14 +202,14 @@
|
|
| 202 |
"decisions": 600
|
| 203 |
},
|
| 204 |
"selected_validation": {
|
| 205 |
-
"soft_cross_entropy": 0.
|
| 206 |
-
"accuracy": 0.
|
| 207 |
-
"brier_soft": 0.
|
| 208 |
"decisions": 600
|
| 209 |
},
|
| 210 |
"stopping": {
|
| 211 |
"reason": "early_stopping",
|
| 212 |
-
"epochs_completed":
|
| 213 |
"patience": 3,
|
| 214 |
"min_epochs": 10,
|
| 215 |
"epochs_without_improvement": 3,
|
|
@@ -223,159 +223,159 @@
|
|
| 223 |
"valid": 2000,
|
| 224 |
"failed": 0,
|
| 225 |
"coverage": 1.0,
|
| 226 |
-
"accuracy_all": 0.
|
| 227 |
-
"accuracy_valid": 0.
|
| 228 |
-
"ece_top_label": 0.
|
| 229 |
-
"mean_confidence": 0.
|
| 230 |
-
"brier_hard": 0.
|
| 231 |
"brier_hard_n": 2000,
|
| 232 |
-
"nll_hard": 0.
|
| 233 |
"nll_hard_n": 2000,
|
| 234 |
"zero_probability_gold": 0.0,
|
| 235 |
"zero_probability_gold_n": 2000,
|
| 236 |
-
"soft_accuracy": 0.
|
| 237 |
"soft_accuracy_n": 2000,
|
| 238 |
-
"brier_soft": 0.
|
| 239 |
"brier_soft_n": 2000,
|
| 240 |
-
"kl_gold_to_prediction": 0.
|
| 241 |
"kl_gold_to_prediction_n": 2000,
|
| 242 |
-
"total_variation": 0.
|
| 243 |
"total_variation_n": 2000,
|
| 244 |
-
"score_mae": 0.
|
| 245 |
"score_mae_n": 800,
|
| 246 |
-
"within_one_level": 0.
|
| 247 |
"within_one_level_n": 800,
|
| 248 |
-
"macro_f1": 0.
|
| 249 |
},
|
| 250 |
"ag-news": {
|
| 251 |
"attempted": 600,
|
| 252 |
"valid": 600,
|
| 253 |
"failed": 0,
|
| 254 |
"coverage": 1.0,
|
| 255 |
-
"accuracy_all": 0.
|
| 256 |
-
"accuracy_valid": 0.
|
| 257 |
-
"ece_top_label": 0.
|
| 258 |
-
"mean_confidence": 0.
|
| 259 |
-
"brier_hard": 0.
|
| 260 |
"brier_hard_n": 600,
|
| 261 |
-
"nll_hard": 0.
|
| 262 |
"nll_hard_n": 600,
|
| 263 |
"zero_probability_gold": 0.0,
|
| 264 |
"zero_probability_gold_n": 600,
|
| 265 |
-
"macro_f1": 0.
|
| 266 |
},
|
| 267 |
"emotion": {
|
| 268 |
"attempted": 600,
|
| 269 |
"valid": 600,
|
| 270 |
"failed": 0,
|
| 271 |
"coverage": 1.0,
|
| 272 |
-
"accuracy_all": 0.
|
| 273 |
-
"accuracy_valid": 0.
|
| 274 |
-
"ece_top_label": 0.
|
| 275 |
-
"mean_confidence": 0.
|
| 276 |
-
"brier_hard": 0.
|
| 277 |
"brier_hard_n": 600,
|
| 278 |
-
"nll_hard": 2.
|
| 279 |
"nll_hard_n": 600,
|
| 280 |
-
"zero_probability_gold": 0.
|
| 281 |
"zero_probability_gold_n": 600,
|
| 282 |
-
"macro_f1": 0.
|
| 283 |
},
|
| 284 |
"boolq": {
|
| 285 |
"attempted": 600,
|
| 286 |
"valid": 600,
|
| 287 |
"failed": 0,
|
| 288 |
"coverage": 1.0,
|
| 289 |
-
"accuracy_all": 0.
|
| 290 |
-
"accuracy_valid": 0.
|
| 291 |
-
"ece_top_label": 0.
|
| 292 |
-
"mean_confidence": 0.
|
| 293 |
-
"brier_hard": 0.
|
| 294 |
"brier_hard_n": 600,
|
| 295 |
-
"nll_hard": 0.
|
| 296 |
"nll_hard_n": 600,
|
| 297 |
"zero_probability_gold": 0.0,
|
| 298 |
"zero_probability_gold_n": 600,
|
| 299 |
-
"macro_f1": 0.
|
| 300 |
},
|
| 301 |
"sst5": {
|
| 302 |
"attempted": 600,
|
| 303 |
"valid": 600,
|
| 304 |
"failed": 0,
|
| 305 |
"coverage": 1.0,
|
| 306 |
-
"accuracy_all": 0.
|
| 307 |
-
"accuracy_valid": 0.
|
| 308 |
-
"ece_top_label": 0.
|
| 309 |
-
"mean_confidence": 0.
|
| 310 |
-
"brier_hard": 0.
|
| 311 |
"brier_hard_n": 600,
|
| 312 |
-
"nll_hard": 1.
|
| 313 |
"nll_hard_n": 600,
|
| 314 |
"zero_probability_gold": 0.0,
|
| 315 |
"zero_probability_gold_n": 600,
|
| 316 |
-
"score_mae": 0.
|
| 317 |
"score_mae_n": 600,
|
| 318 |
-
"within_one_level": 0.
|
| 319 |
"within_one_level_n": 600,
|
| 320 |
-
"macro_f1": 0.
|
| 321 |
},
|
| 322 |
"prompt-injections": {
|
| 323 |
"attempted": 116,
|
| 324 |
"valid": 116,
|
| 325 |
"failed": 0,
|
| 326 |
"coverage": 1.0,
|
| 327 |
-
"accuracy_all": 0.
|
| 328 |
-
"accuracy_valid": 0.
|
| 329 |
-
"ece_top_label": 0.
|
| 330 |
-
"mean_confidence": 0.
|
| 331 |
-
"brier_hard": 0.
|
| 332 |
"brier_hard_n": 116,
|
| 333 |
-
"nll_hard": 1.
|
| 334 |
"nll_hard_n": 116,
|
| 335 |
-
"zero_probability_gold": 0.
|
| 336 |
"zero_probability_gold_n": 116,
|
| 337 |
-
"macro_f1": 0.
|
| 338 |
},
|
| 339 |
"massive-intent.en": {
|
| 340 |
"attempted": 300,
|
| 341 |
"valid": 300,
|
| 342 |
"failed": 0,
|
| 343 |
"coverage": 1.0,
|
| 344 |
-
"accuracy_all": 0.
|
| 345 |
-
"accuracy_valid": 0.
|
| 346 |
-
"ece_top_label": 0.
|
| 347 |
-
"mean_confidence": 0.
|
| 348 |
-
"brier_hard": 0.
|
| 349 |
"brier_hard_n": 300,
|
| 350 |
-
"nll_hard":
|
| 351 |
"nll_hard_n": 300,
|
| 352 |
-
"zero_probability_gold": 0.
|
| 353 |
"zero_probability_gold_n": 300,
|
| 354 |
-
"macro_f1": 0.
|
| 355 |
},
|
| 356 |
"xnli.en": {
|
| 357 |
"attempted": 300,
|
| 358 |
"valid": 300,
|
| 359 |
"failed": 0,
|
| 360 |
"coverage": 1.0,
|
| 361 |
-
"accuracy_all": 0.
|
| 362 |
-
"accuracy_valid": 0.
|
| 363 |
-
"ece_top_label": 0.
|
| 364 |
-
"mean_confidence": 0.
|
| 365 |
-
"brier_hard": 0.
|
| 366 |
"brier_hard_n": 300,
|
| 367 |
-
"nll_hard": 0.
|
| 368 |
"nll_hard_n": 300,
|
| 369 |
"zero_probability_gold": 0.0,
|
| 370 |
"zero_probability_gold_n": 300,
|
| 371 |
-
"macro_f1": 0.
|
| 372 |
}
|
| 373 |
}
|
| 374 |
},
|
| 375 |
"2": {
|
| 376 |
"seed": 2,
|
| 377 |
-
"selected_epoch":
|
| 378 |
-
"elapsed_s":
|
| 379 |
"initial_validation": {
|
| 380 |
"soft_cross_entropy": 1.544869564374288,
|
| 381 |
"accuracy": 0.36833333333333335,
|
|
@@ -383,14 +383,14 @@
|
|
| 383 |
"decisions": 600
|
| 384 |
},
|
| 385 |
"selected_validation": {
|
| 386 |
-
"soft_cross_entropy": 0.
|
| 387 |
-
"accuracy": 0.
|
| 388 |
-
"brier_soft": 0.
|
| 389 |
"decisions": 600
|
| 390 |
},
|
| 391 |
"stopping": {
|
| 392 |
"reason": "early_stopping",
|
| 393 |
-
"epochs_completed":
|
| 394 |
"patience": 3,
|
| 395 |
"min_epochs": 10,
|
| 396 |
"epochs_without_improvement": 3,
|
|
@@ -404,159 +404,159 @@
|
|
| 404 |
"valid": 2000,
|
| 405 |
"failed": 0,
|
| 406 |
"coverage": 1.0,
|
| 407 |
-
"accuracy_all": 0.
|
| 408 |
-
"accuracy_valid": 0.
|
| 409 |
-
"ece_top_label": 0.
|
| 410 |
-
"mean_confidence": 0.
|
| 411 |
-
"brier_hard": 0.
|
| 412 |
"brier_hard_n": 2000,
|
| 413 |
-
"nll_hard": 0.
|
| 414 |
"nll_hard_n": 2000,
|
| 415 |
"zero_probability_gold": 0.0,
|
| 416 |
"zero_probability_gold_n": 2000,
|
| 417 |
-
"soft_accuracy": 0.
|
| 418 |
"soft_accuracy_n": 2000,
|
| 419 |
-
"brier_soft": 0.
|
| 420 |
"brier_soft_n": 2000,
|
| 421 |
-
"kl_gold_to_prediction": 0.
|
| 422 |
"kl_gold_to_prediction_n": 2000,
|
| 423 |
-
"total_variation": 0.
|
| 424 |
"total_variation_n": 2000,
|
| 425 |
-
"score_mae": 0.
|
| 426 |
"score_mae_n": 800,
|
| 427 |
-
"within_one_level": 0.
|
| 428 |
"within_one_level_n": 800,
|
| 429 |
-
"macro_f1": 0.
|
| 430 |
},
|
| 431 |
"ag-news": {
|
| 432 |
"attempted": 600,
|
| 433 |
"valid": 600,
|
| 434 |
"failed": 0,
|
| 435 |
"coverage": 1.0,
|
| 436 |
-
"accuracy_all": 0.
|
| 437 |
-
"accuracy_valid": 0.
|
| 438 |
-
"ece_top_label": 0.
|
| 439 |
-
"mean_confidence": 0.
|
| 440 |
-
"brier_hard": 0.
|
| 441 |
"brier_hard_n": 600,
|
| 442 |
-
"nll_hard":
|
| 443 |
"nll_hard_n": 600,
|
| 444 |
"zero_probability_gold": 0.0,
|
| 445 |
"zero_probability_gold_n": 600,
|
| 446 |
-
"macro_f1": 0.
|
| 447 |
},
|
| 448 |
"emotion": {
|
| 449 |
"attempted": 600,
|
| 450 |
"valid": 600,
|
| 451 |
"failed": 0,
|
| 452 |
"coverage": 1.0,
|
| 453 |
-
"accuracy_all": 0.
|
| 454 |
-
"accuracy_valid": 0.
|
| 455 |
-
"ece_top_label": 0.
|
| 456 |
-
"mean_confidence": 0.
|
| 457 |
-
"brier_hard": 0.
|
| 458 |
"brier_hard_n": 600,
|
| 459 |
-
"nll_hard": 2.
|
| 460 |
"nll_hard_n": 600,
|
| 461 |
-
"zero_probability_gold": 0.
|
| 462 |
"zero_probability_gold_n": 600,
|
| 463 |
-
"macro_f1": 0.
|
| 464 |
},
|
| 465 |
"boolq": {
|
| 466 |
"attempted": 600,
|
| 467 |
"valid": 600,
|
| 468 |
"failed": 0,
|
| 469 |
"coverage": 1.0,
|
| 470 |
-
"accuracy_all": 0.
|
| 471 |
-
"accuracy_valid": 0.
|
| 472 |
-
"ece_top_label": 0.
|
| 473 |
-
"mean_confidence": 0.
|
| 474 |
-
"brier_hard": 0.
|
| 475 |
"brier_hard_n": 600,
|
| 476 |
-
"nll_hard": 0.
|
| 477 |
"nll_hard_n": 600,
|
| 478 |
"zero_probability_gold": 0.0,
|
| 479 |
"zero_probability_gold_n": 600,
|
| 480 |
-
"macro_f1": 0.
|
| 481 |
},
|
| 482 |
"sst5": {
|
| 483 |
"attempted": 600,
|
| 484 |
"valid": 600,
|
| 485 |
"failed": 0,
|
| 486 |
"coverage": 1.0,
|
| 487 |
-
"accuracy_all": 0.
|
| 488 |
-
"accuracy_valid": 0.
|
| 489 |
-
"ece_top_label": 0.
|
| 490 |
-
"mean_confidence": 0.
|
| 491 |
-
"brier_hard": 0.
|
| 492 |
"brier_hard_n": 600,
|
| 493 |
-
"nll_hard": 1.
|
| 494 |
"nll_hard_n": 600,
|
| 495 |
"zero_probability_gold": 0.0,
|
| 496 |
"zero_probability_gold_n": 600,
|
| 497 |
-
"score_mae":
|
| 498 |
"score_mae_n": 600,
|
| 499 |
-
"within_one_level": 0.
|
| 500 |
"within_one_level_n": 600,
|
| 501 |
-
"macro_f1": 0.
|
| 502 |
},
|
| 503 |
"prompt-injections": {
|
| 504 |
"attempted": 116,
|
| 505 |
"valid": 116,
|
| 506 |
"failed": 0,
|
| 507 |
"coverage": 1.0,
|
| 508 |
-
"accuracy_all": 0.
|
| 509 |
-
"accuracy_valid": 0.
|
| 510 |
-
"ece_top_label": 0.
|
| 511 |
-
"mean_confidence": 0.
|
| 512 |
-
"brier_hard": 0.
|
| 513 |
"brier_hard_n": 116,
|
| 514 |
-
"nll_hard":
|
| 515 |
"nll_hard_n": 116,
|
| 516 |
-
"zero_probability_gold": 0.
|
| 517 |
"zero_probability_gold_n": 116,
|
| 518 |
-
"macro_f1": 0.
|
| 519 |
},
|
| 520 |
"massive-intent.en": {
|
| 521 |
"attempted": 300,
|
| 522 |
"valid": 300,
|
| 523 |
"failed": 0,
|
| 524 |
"coverage": 1.0,
|
| 525 |
-
"accuracy_all": 0.
|
| 526 |
-
"accuracy_valid": 0.
|
| 527 |
-
"ece_top_label": 0.
|
| 528 |
-
"mean_confidence": 0.
|
| 529 |
-
"brier_hard": 0.
|
| 530 |
"brier_hard_n": 300,
|
| 531 |
-
"nll_hard": 3.
|
| 532 |
"nll_hard_n": 300,
|
| 533 |
-
"zero_probability_gold": 0.
|
| 534 |
"zero_probability_gold_n": 300,
|
| 535 |
-
"macro_f1": 0.
|
| 536 |
},
|
| 537 |
"xnli.en": {
|
| 538 |
"attempted": 300,
|
| 539 |
"valid": 300,
|
| 540 |
"failed": 0,
|
| 541 |
"coverage": 1.0,
|
| 542 |
-
"accuracy_all": 0.
|
| 543 |
-
"accuracy_valid": 0.
|
| 544 |
-
"ece_top_label": 0.
|
| 545 |
-
"mean_confidence": 0.
|
| 546 |
-
"brier_hard": 0.
|
| 547 |
"brier_hard_n": 300,
|
| 548 |
-
"nll_hard":
|
| 549 |
"nll_hard_n": 300,
|
| 550 |
"zero_probability_gold": 0.0,
|
| 551 |
"zero_probability_gold_n": 300,
|
| 552 |
-
"macro_f1": 0.
|
| 553 |
}
|
| 554 |
}
|
| 555 |
},
|
| 556 |
"3": {
|
| 557 |
"seed": 3,
|
| 558 |
-
"selected_epoch":
|
| 559 |
-
"elapsed_s":
|
| 560 |
"initial_validation": {
|
| 561 |
"soft_cross_entropy": 1.544869564374288,
|
| 562 |
"accuracy": 0.36833333333333335,
|
|
@@ -564,14 +564,14 @@
|
|
| 564 |
"decisions": 600
|
| 565 |
},
|
| 566 |
"selected_validation": {
|
| 567 |
-
"soft_cross_entropy": 0.
|
| 568 |
-
"accuracy": 0.
|
| 569 |
-
"brier_soft": 0.
|
| 570 |
"decisions": 600
|
| 571 |
},
|
| 572 |
"stopping": {
|
| 573 |
"reason": "early_stopping",
|
| 574 |
-
"epochs_completed":
|
| 575 |
"patience": 3,
|
| 576 |
"min_epochs": 10,
|
| 577 |
"epochs_without_improvement": 3,
|
|
@@ -585,159 +585,159 @@
|
|
| 585 |
"valid": 2000,
|
| 586 |
"failed": 0,
|
| 587 |
"coverage": 1.0,
|
| 588 |
-
"accuracy_all": 0.
|
| 589 |
-
"accuracy_valid": 0.
|
| 590 |
-
"ece_top_label": 0.
|
| 591 |
-
"mean_confidence": 0.
|
| 592 |
-
"brier_hard": 0.
|
| 593 |
"brier_hard_n": 2000,
|
| 594 |
-
"nll_hard": 0.
|
| 595 |
"nll_hard_n": 2000,
|
| 596 |
"zero_probability_gold": 0.0,
|
| 597 |
"zero_probability_gold_n": 2000,
|
| 598 |
-
"soft_accuracy": 0.
|
| 599 |
"soft_accuracy_n": 2000,
|
| 600 |
-
"brier_soft": 0.
|
| 601 |
"brier_soft_n": 2000,
|
| 602 |
-
"kl_gold_to_prediction": 0.
|
| 603 |
"kl_gold_to_prediction_n": 2000,
|
| 604 |
-
"total_variation": 0.
|
| 605 |
"total_variation_n": 2000,
|
| 606 |
-
"score_mae": 0.
|
| 607 |
"score_mae_n": 800,
|
| 608 |
-
"within_one_level": 0.
|
| 609 |
"within_one_level_n": 800,
|
| 610 |
-
"macro_f1": 0.
|
| 611 |
},
|
| 612 |
"ag-news": {
|
| 613 |
"attempted": 600,
|
| 614 |
"valid": 600,
|
| 615 |
"failed": 0,
|
| 616 |
"coverage": 1.0,
|
| 617 |
-
"accuracy_all": 0.
|
| 618 |
-
"accuracy_valid": 0.
|
| 619 |
-
"ece_top_label": 0.
|
| 620 |
-
"mean_confidence": 0.
|
| 621 |
-
"brier_hard": 0.
|
| 622 |
"brier_hard_n": 600,
|
| 623 |
-
"nll_hard":
|
| 624 |
"nll_hard_n": 600,
|
| 625 |
"zero_probability_gold": 0.0,
|
| 626 |
"zero_probability_gold_n": 600,
|
| 627 |
-
"macro_f1": 0.
|
| 628 |
},
|
| 629 |
"emotion": {
|
| 630 |
"attempted": 600,
|
| 631 |
"valid": 600,
|
| 632 |
"failed": 0,
|
| 633 |
"coverage": 1.0,
|
| 634 |
-
"accuracy_all": 0.
|
| 635 |
-
"accuracy_valid": 0.
|
| 636 |
-
"ece_top_label": 0.
|
| 637 |
-
"mean_confidence": 0.
|
| 638 |
-
"brier_hard": 0.
|
| 639 |
"brier_hard_n": 600,
|
| 640 |
-
"nll_hard": 2.
|
| 641 |
"nll_hard_n": 600,
|
| 642 |
-
"zero_probability_gold": 0.
|
| 643 |
"zero_probability_gold_n": 600,
|
| 644 |
-
"macro_f1": 0.
|
| 645 |
},
|
| 646 |
"boolq": {
|
| 647 |
"attempted": 600,
|
| 648 |
"valid": 600,
|
| 649 |
"failed": 0,
|
| 650 |
"coverage": 1.0,
|
| 651 |
-
"accuracy_all": 0.
|
| 652 |
-
"accuracy_valid": 0.
|
| 653 |
-
"ece_top_label": 0.
|
| 654 |
-
"mean_confidence": 0.
|
| 655 |
-
"brier_hard": 0.
|
| 656 |
"brier_hard_n": 600,
|
| 657 |
-
"nll_hard": 0.
|
| 658 |
"nll_hard_n": 600,
|
| 659 |
"zero_probability_gold": 0.0,
|
| 660 |
"zero_probability_gold_n": 600,
|
| 661 |
-
"macro_f1": 0.
|
| 662 |
},
|
| 663 |
"sst5": {
|
| 664 |
"attempted": 600,
|
| 665 |
"valid": 600,
|
| 666 |
"failed": 0,
|
| 667 |
"coverage": 1.0,
|
| 668 |
-
"accuracy_all": 0.
|
| 669 |
-
"accuracy_valid": 0.
|
| 670 |
-
"ece_top_label": 0.
|
| 671 |
-
"mean_confidence": 0.
|
| 672 |
-
"brier_hard": 0.
|
| 673 |
"brier_hard_n": 600,
|
| 674 |
-
"nll_hard": 1.
|
| 675 |
"nll_hard_n": 600,
|
| 676 |
"zero_probability_gold": 0.0,
|
| 677 |
"zero_probability_gold_n": 600,
|
| 678 |
-
"score_mae":
|
| 679 |
"score_mae_n": 600,
|
| 680 |
-
"within_one_level": 0.
|
| 681 |
"within_one_level_n": 600,
|
| 682 |
-
"macro_f1": 0.
|
| 683 |
},
|
| 684 |
"prompt-injections": {
|
| 685 |
"attempted": 116,
|
| 686 |
"valid": 116,
|
| 687 |
"failed": 0,
|
| 688 |
"coverage": 1.0,
|
| 689 |
-
"accuracy_all": 0.
|
| 690 |
-
"accuracy_valid": 0.
|
| 691 |
-
"ece_top_label": 0.
|
| 692 |
-
"mean_confidence": 0.
|
| 693 |
-
"brier_hard": 0.
|
| 694 |
"brier_hard_n": 116,
|
| 695 |
-
"nll_hard":
|
| 696 |
"nll_hard_n": 116,
|
| 697 |
-
"zero_probability_gold": 0.
|
| 698 |
"zero_probability_gold_n": 116,
|
| 699 |
-
"macro_f1": 0.
|
| 700 |
},
|
| 701 |
"massive-intent.en": {
|
| 702 |
"attempted": 300,
|
| 703 |
"valid": 300,
|
| 704 |
"failed": 0,
|
| 705 |
"coverage": 1.0,
|
| 706 |
-
"accuracy_all": 0.
|
| 707 |
-
"accuracy_valid": 0.
|
| 708 |
-
"ece_top_label": 0.
|
| 709 |
-
"mean_confidence": 0.
|
| 710 |
-
"brier_hard":
|
| 711 |
"brier_hard_n": 300,
|
| 712 |
-
"nll_hard":
|
| 713 |
"nll_hard_n": 300,
|
| 714 |
-
"zero_probability_gold": 0.
|
| 715 |
"zero_probability_gold_n": 300,
|
| 716 |
-
"macro_f1": 0.
|
| 717 |
},
|
| 718 |
"xnli.en": {
|
| 719 |
"attempted": 300,
|
| 720 |
"valid": 300,
|
| 721 |
"failed": 0,
|
| 722 |
"coverage": 1.0,
|
| 723 |
-
"accuracy_all": 0.
|
| 724 |
-
"accuracy_valid": 0.
|
| 725 |
-
"ece_top_label": 0.
|
| 726 |
-
"mean_confidence": 0.
|
| 727 |
-
"brier_hard": 0.
|
| 728 |
"brier_hard_n": 300,
|
| 729 |
-
"nll_hard":
|
| 730 |
"nll_hard_n": 300,
|
| 731 |
"zero_probability_gold": 0.0,
|
| 732 |
"zero_probability_gold_n": 300,
|
| 733 |
-
"macro_f1": 0.
|
| 734 |
}
|
| 735 |
}
|
| 736 |
},
|
| 737 |
"4": {
|
| 738 |
"seed": 4,
|
| 739 |
-
"selected_epoch":
|
| 740 |
-
"elapsed_s":
|
| 741 |
"initial_validation": {
|
| 742 |
"soft_cross_entropy": 1.544869564374288,
|
| 743 |
"accuracy": 0.36833333333333335,
|
|
@@ -745,14 +745,14 @@
|
|
| 745 |
"decisions": 600
|
| 746 |
},
|
| 747 |
"selected_validation": {
|
| 748 |
-
"soft_cross_entropy": 0.
|
| 749 |
-
"accuracy": 0.
|
| 750 |
-
"brier_soft": 0.
|
| 751 |
"decisions": 600
|
| 752 |
},
|
| 753 |
"stopping": {
|
| 754 |
"reason": "early_stopping",
|
| 755 |
-
"epochs_completed":
|
| 756 |
"patience": 3,
|
| 757 |
"min_epochs": 10,
|
| 758 |
"epochs_without_improvement": 3,
|
|
@@ -766,152 +766,152 @@
|
|
| 766 |
"valid": 2000,
|
| 767 |
"failed": 0,
|
| 768 |
"coverage": 1.0,
|
| 769 |
-
"accuracy_all": 0.
|
| 770 |
-
"accuracy_valid": 0.
|
| 771 |
-
"ece_top_label": 0.
|
| 772 |
-
"mean_confidence": 0.
|
| 773 |
-
"brier_hard": 0.
|
| 774 |
"brier_hard_n": 2000,
|
| 775 |
-
"nll_hard": 0.
|
| 776 |
"nll_hard_n": 2000,
|
| 777 |
"zero_probability_gold": 0.0,
|
| 778 |
"zero_probability_gold_n": 2000,
|
| 779 |
-
"soft_accuracy": 0.
|
| 780 |
"soft_accuracy_n": 2000,
|
| 781 |
-
"brier_soft": 0.
|
| 782 |
"brier_soft_n": 2000,
|
| 783 |
-
"kl_gold_to_prediction": 0.
|
| 784 |
"kl_gold_to_prediction_n": 2000,
|
| 785 |
-
"total_variation": 0.
|
| 786 |
"total_variation_n": 2000,
|
| 787 |
-
"score_mae": 0.
|
| 788 |
"score_mae_n": 800,
|
| 789 |
-
"within_one_level": 0.
|
| 790 |
"within_one_level_n": 800,
|
| 791 |
-
"macro_f1": 0.
|
| 792 |
},
|
| 793 |
"ag-news": {
|
| 794 |
"attempted": 600,
|
| 795 |
"valid": 600,
|
| 796 |
"failed": 0,
|
| 797 |
"coverage": 1.0,
|
| 798 |
-
"accuracy_all": 0.
|
| 799 |
-
"accuracy_valid": 0.
|
| 800 |
-
"ece_top_label": 0.
|
| 801 |
-
"mean_confidence": 0.
|
| 802 |
-
"brier_hard": 0.
|
| 803 |
"brier_hard_n": 600,
|
| 804 |
-
"nll_hard": 0.
|
| 805 |
"nll_hard_n": 600,
|
| 806 |
"zero_probability_gold": 0.0,
|
| 807 |
"zero_probability_gold_n": 600,
|
| 808 |
-
"macro_f1": 0.
|
| 809 |
},
|
| 810 |
"emotion": {
|
| 811 |
"attempted": 600,
|
| 812 |
"valid": 600,
|
| 813 |
"failed": 0,
|
| 814 |
"coverage": 1.0,
|
| 815 |
-
"accuracy_all": 0.
|
| 816 |
-
"accuracy_valid": 0.
|
| 817 |
-
"ece_top_label": 0.
|
| 818 |
-
"mean_confidence": 0.
|
| 819 |
-
"brier_hard": 0.
|
| 820 |
"brier_hard_n": 600,
|
| 821 |
-
"nll_hard": 2.
|
| 822 |
"nll_hard_n": 600,
|
| 823 |
-
"zero_probability_gold": 0.
|
| 824 |
"zero_probability_gold_n": 600,
|
| 825 |
-
"macro_f1": 0.
|
| 826 |
},
|
| 827 |
"boolq": {
|
| 828 |
"attempted": 600,
|
| 829 |
"valid": 600,
|
| 830 |
"failed": 0,
|
| 831 |
"coverage": 1.0,
|
| 832 |
-
"accuracy_all": 0.
|
| 833 |
-
"accuracy_valid": 0.
|
| 834 |
-
"ece_top_label": 0.
|
| 835 |
-
"mean_confidence": 0.
|
| 836 |
-
"brier_hard": 0.
|
| 837 |
"brier_hard_n": 600,
|
| 838 |
-
"nll_hard": 0.
|
| 839 |
"nll_hard_n": 600,
|
| 840 |
-
"zero_probability_gold": 0.
|
| 841 |
"zero_probability_gold_n": 600,
|
| 842 |
-
"macro_f1": 0.
|
| 843 |
},
|
| 844 |
"sst5": {
|
| 845 |
"attempted": 600,
|
| 846 |
"valid": 600,
|
| 847 |
"failed": 0,
|
| 848 |
"coverage": 1.0,
|
| 849 |
-
"accuracy_all": 0.
|
| 850 |
-
"accuracy_valid": 0.
|
| 851 |
-
"ece_top_label": 0.
|
| 852 |
-
"mean_confidence": 0.
|
| 853 |
-
"brier_hard": 0.
|
| 854 |
"brier_hard_n": 600,
|
| 855 |
-
"nll_hard": 1.
|
| 856 |
"nll_hard_n": 600,
|
| 857 |
"zero_probability_gold": 0.0,
|
| 858 |
"zero_probability_gold_n": 600,
|
| 859 |
-
"score_mae": 0.
|
| 860 |
"score_mae_n": 600,
|
| 861 |
-
"within_one_level": 0.
|
| 862 |
"within_one_level_n": 600,
|
| 863 |
-
"macro_f1": 0.
|
| 864 |
},
|
| 865 |
"prompt-injections": {
|
| 866 |
"attempted": 116,
|
| 867 |
"valid": 116,
|
| 868 |
"failed": 0,
|
| 869 |
"coverage": 1.0,
|
| 870 |
-
"accuracy_all": 0.
|
| 871 |
-
"accuracy_valid": 0.
|
| 872 |
-
"ece_top_label": 0.
|
| 873 |
-
"mean_confidence": 0.
|
| 874 |
-
"brier_hard": 0.
|
| 875 |
"brier_hard_n": 116,
|
| 876 |
-
"nll_hard":
|
| 877 |
"nll_hard_n": 116,
|
| 878 |
-
"zero_probability_gold": 0.
|
| 879 |
"zero_probability_gold_n": 116,
|
| 880 |
-
"macro_f1": 0.
|
| 881 |
},
|
| 882 |
"massive-intent.en": {
|
| 883 |
"attempted": 300,
|
| 884 |
"valid": 300,
|
| 885 |
"failed": 0,
|
| 886 |
"coverage": 1.0,
|
| 887 |
-
"accuracy_all": 0.
|
| 888 |
-
"accuracy_valid": 0.
|
| 889 |
-
"ece_top_label": 0.
|
| 890 |
-
"mean_confidence": 0.
|
| 891 |
-
"brier_hard": 0.
|
| 892 |
"brier_hard_n": 300,
|
| 893 |
-
"nll_hard": 2.
|
| 894 |
"nll_hard_n": 300,
|
| 895 |
-
"zero_probability_gold": 0.
|
| 896 |
"zero_probability_gold_n": 300,
|
| 897 |
-
"macro_f1": 0.
|
| 898 |
},
|
| 899 |
"xnli.en": {
|
| 900 |
"attempted": 300,
|
| 901 |
"valid": 300,
|
| 902 |
"failed": 0,
|
| 903 |
"coverage": 1.0,
|
| 904 |
-
"accuracy_all": 0.
|
| 905 |
-
"accuracy_valid": 0.
|
| 906 |
-
"ece_top_label": 0.
|
| 907 |
-
"mean_confidence": 0.
|
| 908 |
-
"brier_hard": 0.
|
| 909 |
"brier_hard_n": 300,
|
| 910 |
-
"nll_hard": 0.
|
| 911 |
"nll_hard_n": 300,
|
| 912 |
"zero_probability_gold": 0.0,
|
| 913 |
"zero_probability_gold_n": 300,
|
| 914 |
-
"macro_f1": 0.
|
| 915 |
}
|
| 916 |
}
|
| 917 |
}
|
|
@@ -919,54 +919,54 @@
|
|
| 919 |
"summary": {
|
| 920 |
"typed-decisions": {
|
| 921 |
"n": 5,
|
| 922 |
-
"mean_percent":
|
| 923 |
-
"sample_variance_pp_squared":
|
| 924 |
-
"sample_std_pp":
|
| 925 |
},
|
| 926 |
"ag-news": {
|
| 927 |
"n": 5,
|
| 928 |
-
"mean_percent":
|
| 929 |
-
"sample_variance_pp_squared":
|
| 930 |
-
"sample_std_pp":
|
| 931 |
},
|
| 932 |
"boolq": {
|
| 933 |
"n": 5,
|
| 934 |
-
"mean_percent":
|
| 935 |
-
"sample_variance_pp_squared":
|
| 936 |
-
"sample_std_pp":
|
| 937 |
},
|
| 938 |
"emotion": {
|
| 939 |
"n": 5,
|
| 940 |
-
"mean_percent":
|
| 941 |
-
"sample_variance_pp_squared":
|
| 942 |
-
"sample_std_pp":
|
| 943 |
},
|
| 944 |
"prompt-injections": {
|
| 945 |
"n": 5,
|
| 946 |
-
"mean_percent":
|
| 947 |
-
"sample_variance_pp_squared":
|
| 948 |
-
"sample_std_pp":
|
| 949 |
},
|
| 950 |
"sst5": {
|
| 951 |
"n": 5,
|
| 952 |
-
"mean_percent":
|
| 953 |
-
"sample_variance_pp_squared":
|
| 954 |
-
"sample_std_pp":
|
| 955 |
},
|
| 956 |
"massive-intent.en": {
|
| 957 |
"n": 5,
|
| 958 |
-
"mean_percent":
|
| 959 |
-
"sample_variance_pp_squared":
|
| 960 |
-
"sample_std_pp":
|
| 961 |
},
|
| 962 |
"xnli.en": {
|
| 963 |
"n": 5,
|
| 964 |
-
"mean_percent":
|
| 965 |
-
"sample_variance_pp_squared":
|
| 966 |
-
"sample_std_pp":
|
| 967 |
}
|
| 968 |
},
|
| 969 |
-
"candidate_seed":
|
| 970 |
"variance_definition": "sample variance, n-1, percentage points squared",
|
| 971 |
"benchmark_manifest": {
|
| 972 |
"format_version": 1,
|
|
|
|
| 12 |
"runs": {
|
| 13 |
"0": {
|
| 14 |
"seed": 0,
|
| 15 |
+
"selected_epoch": 11,
|
| 16 |
+
"elapsed_s": 612.5042744020466,
|
| 17 |
"initial_validation": {
|
| 18 |
"soft_cross_entropy": 1.544869564374288,
|
| 19 |
"accuracy": 0.36833333333333335,
|
|
|
|
| 21 |
"decisions": 600
|
| 22 |
},
|
| 23 |
"selected_validation": {
|
| 24 |
+
"soft_cross_entropy": 0.8771124110619227,
|
| 25 |
+
"accuracy": 0.7116666666666667,
|
| 26 |
+
"brier_soft": 0.0828241604194045,
|
| 27 |
"decisions": 600
|
| 28 |
},
|
| 29 |
"stopping": {
|
| 30 |
"reason": "early_stopping",
|
| 31 |
+
"epochs_completed": 14,
|
| 32 |
"patience": 3,
|
| 33 |
"min_epochs": 10,
|
| 34 |
"epochs_without_improvement": 3,
|
|
|
|
| 42 |
"valid": 2000,
|
| 43 |
"failed": 0,
|
| 44 |
"coverage": 1.0,
|
| 45 |
+
"accuracy_all": 0.7425,
|
| 46 |
+
"accuracy_valid": 0.7425,
|
| 47 |
+
"ece_top_label": 0.19222331993705236,
|
| 48 |
+
"mean_confidence": 0.5502766800629476,
|
| 49 |
+
"brier_hard": 0.41714854154295294,
|
| 50 |
"brier_hard_n": 2000,
|
| 51 |
+
"nll_hard": 0.7346455747590726,
|
| 52 |
"nll_hard_n": 2000,
|
| 53 |
"zero_probability_gold": 0.0,
|
| 54 |
"zero_probability_gold_n": 2000,
|
| 55 |
+
"soft_accuracy": 0.4642362342348646,
|
| 56 |
"soft_accuracy_n": 2000,
|
| 57 |
+
"brier_soft": 0.07282817778125876,
|
| 58 |
"brier_soft_n": 2000,
|
| 59 |
+
"kl_gold_to_prediction": 0.13635658195747785,
|
| 60 |
"kl_gold_to_prediction_n": 2000,
|
| 61 |
+
"total_variation": 0.18626503447270873,
|
| 62 |
"total_variation_n": 2000,
|
| 63 |
+
"score_mae": 0.2670681801272688,
|
| 64 |
"score_mae_n": 800,
|
| 65 |
+
"within_one_level": 0.98125,
|
| 66 |
"within_one_level_n": 800,
|
| 67 |
+
"macro_f1": 0.6211545025911077
|
| 68 |
},
|
| 69 |
"ag-news": {
|
| 70 |
"attempted": 600,
|
| 71 |
"valid": 600,
|
| 72 |
"failed": 0,
|
| 73 |
"coverage": 1.0,
|
| 74 |
+
"accuracy_all": 0.9416666666666667,
|
| 75 |
+
"accuracy_valid": 0.9416666666666667,
|
| 76 |
+
"ece_top_label": 0.032780915974200804,
|
| 77 |
+
"mean_confidence": 0.9120336837609119,
|
| 78 |
+
"brier_hard": 0.0911117485859852,
|
| 79 |
"brier_hard_n": 600,
|
| 80 |
+
"nll_hard": 0.18384446182116462,
|
| 81 |
"nll_hard_n": 600,
|
| 82 |
"zero_probability_gold": 0.0,
|
| 83 |
"zero_probability_gold_n": 600,
|
| 84 |
+
"macro_f1": 0.9391643390047506
|
| 85 |
},
|
| 86 |
"emotion": {
|
| 87 |
"attempted": 600,
|
| 88 |
"valid": 600,
|
| 89 |
"failed": 0,
|
| 90 |
"coverage": 1.0,
|
| 91 |
+
"accuracy_all": 0.585,
|
| 92 |
+
"accuracy_valid": 0.585,
|
| 93 |
+
"ece_top_label": 0.3290210797041948,
|
| 94 |
+
"mean_confidence": 0.9044507912979094,
|
| 95 |
+
"brier_hard": 0.7352614358764159,
|
| 96 |
"brier_hard_n": 600,
|
| 97 |
+
"nll_hard": 2.325969372091789,
|
| 98 |
"nll_hard_n": 600,
|
| 99 |
+
"zero_probability_gold": 0.01,
|
| 100 |
"zero_probability_gold_n": 600,
|
| 101 |
+
"macro_f1": 0.4846329953255675
|
| 102 |
},
|
| 103 |
"boolq": {
|
| 104 |
"attempted": 600,
|
| 105 |
"valid": 600,
|
| 106 |
"failed": 0,
|
| 107 |
"coverage": 1.0,
|
| 108 |
+
"accuracy_all": 0.8083333333333333,
|
| 109 |
+
"accuracy_valid": 0.8083333333333333,
|
| 110 |
+
"ece_top_label": 0.09874549999999996,
|
| 111 |
+
"mean_confidence": 0.9070788333333333,
|
| 112 |
+
"brier_hard": 0.2864656637,
|
| 113 |
"brier_hard_n": 600,
|
| 114 |
+
"nll_hard": 0.45967158478397707,
|
| 115 |
"nll_hard_n": 600,
|
| 116 |
"zero_probability_gold": 0.0,
|
| 117 |
"zero_probability_gold_n": 600,
|
| 118 |
+
"macro_f1": 0.7946275764565816
|
| 119 |
},
|
| 120 |
"sst5": {
|
| 121 |
"attempted": 600,
|
| 122 |
"valid": 600,
|
| 123 |
"failed": 0,
|
| 124 |
"coverage": 1.0,
|
| 125 |
+
"accuracy_all": 0.38166666666666665,
|
| 126 |
+
"accuracy_valid": 0.38166666666666665,
|
| 127 |
+
"ece_top_label": 0.23314279850468647,
|
| 128 |
+
"mean_confidence": 0.614809465171353,
|
| 129 |
+
"brier_hard": 0.7997842513426261,
|
| 130 |
"brier_hard_n": 600,
|
| 131 |
+
"nll_hard": 1.7175170777887119,
|
| 132 |
"nll_hard_n": 600,
|
| 133 |
"zero_probability_gold": 0.0,
|
| 134 |
"zero_probability_gold_n": 600,
|
| 135 |
+
"score_mae": 0.9144952235712186,
|
| 136 |
"score_mae_n": 600,
|
| 137 |
+
"within_one_level": 0.6466666666666666,
|
| 138 |
"within_one_level_n": 600,
|
| 139 |
+
"macro_f1": 0.3319436236284081
|
| 140 |
},
|
| 141 |
"prompt-injections": {
|
| 142 |
"attempted": 116,
|
| 143 |
"valid": 116,
|
| 144 |
"failed": 0,
|
| 145 |
"coverage": 1.0,
|
| 146 |
+
"accuracy_all": 0.6982758620689655,
|
| 147 |
+
"accuracy_valid": 0.6982758620689655,
|
| 148 |
+
"ece_top_label": 0.25387413793103447,
|
| 149 |
+
"mean_confidence": 0.95215,
|
| 150 |
+
"brier_hard": 0.5115509806896552,
|
| 151 |
"brier_hard_n": 116,
|
| 152 |
+
"nll_hard": 2.710472238032788,
|
| 153 |
"nll_hard_n": 116,
|
| 154 |
+
"zero_probability_gold": 0.06896551724137931,
|
| 155 |
"zero_probability_gold_n": 116,
|
| 156 |
+
"macro_f1": 0.6809931641392315
|
| 157 |
},
|
| 158 |
"massive-intent.en": {
|
| 159 |
"attempted": 300,
|
| 160 |
"valid": 300,
|
| 161 |
"failed": 0,
|
| 162 |
"coverage": 1.0,
|
| 163 |
+
"accuracy_all": 0.79,
|
| 164 |
+
"accuracy_valid": 0.79,
|
| 165 |
+
"ece_top_label": 0.16259349215790048,
|
| 166 |
+
"mean_confidence": 0.9477817275562629,
|
| 167 |
+
"brier_hard": 0.3655905743529746,
|
| 168 |
"brier_hard_n": 300,
|
| 169 |
+
"nll_hard": 2.4460407722040176,
|
| 170 |
"nll_hard_n": 300,
|
| 171 |
+
"zero_probability_gold": 0.06,
|
| 172 |
"zero_probability_gold_n": 300,
|
| 173 |
+
"macro_f1": 0.5695724937908466
|
| 174 |
},
|
| 175 |
"xnli.en": {
|
| 176 |
"attempted": 300,
|
| 177 |
"valid": 300,
|
| 178 |
"failed": 0,
|
| 179 |
"coverage": 1.0,
|
| 180 |
+
"accuracy_all": 0.89,
|
| 181 |
+
"accuracy_valid": 0.89,
|
| 182 |
+
"ece_top_label": 0.05910627327283609,
|
| 183 |
+
"mean_confidence": 0.9123050519493905,
|
| 184 |
+
"brier_hard": 0.18251872229879984,
|
| 185 |
"brier_hard_n": 300,
|
| 186 |
+
"nll_hard": 0.33404125735595414,
|
| 187 |
"nll_hard_n": 300,
|
| 188 |
"zero_probability_gold": 0.0,
|
| 189 |
"zero_probability_gold_n": 300,
|
| 190 |
+
"macro_f1": 0.8905240397272149
|
| 191 |
}
|
| 192 |
}
|
| 193 |
},
|
| 194 |
"1": {
|
| 195 |
"seed": 1,
|
| 196 |
+
"selected_epoch": 9,
|
| 197 |
+
"elapsed_s": 541.4585086829611,
|
| 198 |
"initial_validation": {
|
| 199 |
"soft_cross_entropy": 1.544869564374288,
|
| 200 |
"accuracy": 0.36833333333333335,
|
|
|
|
| 202 |
"decisions": 600
|
| 203 |
},
|
| 204 |
"selected_validation": {
|
| 205 |
+
"soft_cross_entropy": 0.8583925066391627,
|
| 206 |
+
"accuracy": 0.7466666666666667,
|
| 207 |
+
"brier_soft": 0.07437195796364297,
|
| 208 |
"decisions": 600
|
| 209 |
},
|
| 210 |
"stopping": {
|
| 211 |
"reason": "early_stopping",
|
| 212 |
+
"epochs_completed": 12,
|
| 213 |
"patience": 3,
|
| 214 |
"min_epochs": 10,
|
| 215 |
"epochs_without_improvement": 3,
|
|
|
|
| 223 |
"valid": 2000,
|
| 224 |
"failed": 0,
|
| 225 |
"coverage": 1.0,
|
| 226 |
+
"accuracy_all": 0.7565,
|
| 227 |
+
"accuracy_valid": 0.7565,
|
| 228 |
+
"ece_top_label": 0.19864268503011012,
|
| 229 |
+
"mean_confidence": 0.5578573149698899,
|
| 230 |
+
"brier_hard": 0.40482086772099385,
|
| 231 |
"brier_hard_n": 2000,
|
| 232 |
+
"nll_hard": 0.7162594531200565,
|
| 233 |
"nll_hard_n": 2000,
|
| 234 |
"zero_probability_gold": 0.0,
|
| 235 |
"zero_probability_gold_n": 2000,
|
| 236 |
+
"soft_accuracy": 0.469737413364262,
|
| 237 |
"soft_accuracy_n": 2000,
|
| 238 |
+
"brier_soft": 0.0659469654317038,
|
| 239 |
"brier_soft_n": 2000,
|
| 240 |
+
"kl_gold_to_prediction": 0.12376910578634269,
|
| 241 |
"kl_gold_to_prediction_n": 2000,
|
| 242 |
+
"total_variation": 0.17750424596946363,
|
| 243 |
"total_variation_n": 2000,
|
| 244 |
+
"score_mae": 0.24687533223920302,
|
| 245 |
"score_mae_n": 800,
|
| 246 |
+
"within_one_level": 0.98625,
|
| 247 |
"within_one_level_n": 800,
|
| 248 |
+
"macro_f1": 0.6300073343403317
|
| 249 |
},
|
| 250 |
"ag-news": {
|
| 251 |
"attempted": 600,
|
| 252 |
"valid": 600,
|
| 253 |
"failed": 0,
|
| 254 |
"coverage": 1.0,
|
| 255 |
+
"accuracy_all": 0.9483333333333334,
|
| 256 |
+
"accuracy_valid": 0.9483333333333334,
|
| 257 |
+
"ece_top_label": 0.03462398162033311,
|
| 258 |
+
"mean_confidence": 0.9188607247165506,
|
| 259 |
+
"brier_hard": 0.08433672632610516,
|
| 260 |
"brier_hard_n": 600,
|
| 261 |
+
"nll_hard": 0.17417200848974151,
|
| 262 |
"nll_hard_n": 600,
|
| 263 |
"zero_probability_gold": 0.0,
|
| 264 |
"zero_probability_gold_n": 600,
|
| 265 |
+
"macro_f1": 0.9456145576151846
|
| 266 |
},
|
| 267 |
"emotion": {
|
| 268 |
"attempted": 600,
|
| 269 |
"valid": 600,
|
| 270 |
"failed": 0,
|
| 271 |
"coverage": 1.0,
|
| 272 |
+
"accuracy_all": 0.5716666666666667,
|
| 273 |
+
"accuracy_valid": 0.5716666666666667,
|
| 274 |
+
"ece_top_label": 0.33449032496578573,
|
| 275 |
+
"mean_confidence": 0.90098409027152,
|
| 276 |
+
"brier_hard": 0.7328256742449804,
|
| 277 |
"brier_hard_n": 600,
|
| 278 |
+
"nll_hard": 2.388523752936007,
|
| 279 |
"nll_hard_n": 600,
|
| 280 |
+
"zero_probability_gold": 0.013333333333333334,
|
| 281 |
"zero_probability_gold_n": 600,
|
| 282 |
+
"macro_f1": 0.4746076611364094
|
| 283 |
},
|
| 284 |
"boolq": {
|
| 285 |
"attempted": 600,
|
| 286 |
"valid": 600,
|
| 287 |
"failed": 0,
|
| 288 |
"coverage": 1.0,
|
| 289 |
+
"accuracy_all": 0.8266666666666667,
|
| 290 |
+
"accuracy_valid": 0.8266666666666667,
|
| 291 |
+
"ece_top_label": 0.08269850000000001,
|
| 292 |
+
"mean_confidence": 0.9093651666666667,
|
| 293 |
+
"brier_hard": 0.26852242536666665,
|
| 294 |
"brier_hard_n": 600,
|
| 295 |
+
"nll_hard": 0.44571515698724307,
|
| 296 |
"nll_hard_n": 600,
|
| 297 |
"zero_probability_gold": 0.0,
|
| 298 |
"zero_probability_gold_n": 600,
|
| 299 |
+
"macro_f1": 0.8140998140998141
|
| 300 |
},
|
| 301 |
"sst5": {
|
| 302 |
"attempted": 600,
|
| 303 |
"valid": 600,
|
| 304 |
"failed": 0,
|
| 305 |
"coverage": 1.0,
|
| 306 |
+
"accuracy_all": 0.4483333333333333,
|
| 307 |
+
"accuracy_valid": 0.4483333333333333,
|
| 308 |
+
"ece_top_label": 0.1552169576555402,
|
| 309 |
+
"mean_confidence": 0.5963372951766569,
|
| 310 |
+
"brier_hard": 0.715830700927659,
|
| 311 |
"brier_hard_n": 600,
|
| 312 |
+
"nll_hard": 1.4312799777899081,
|
| 313 |
"nll_hard_n": 600,
|
| 314 |
"zero_probability_gold": 0.0,
|
| 315 |
"zero_probability_gold_n": 600,
|
| 316 |
+
"score_mae": 0.763809024216731,
|
| 317 |
"score_mae_n": 600,
|
| 318 |
+
"within_one_level": 0.7216666666666667,
|
| 319 |
"within_one_level_n": 600,
|
| 320 |
+
"macro_f1": 0.4241661418376961
|
| 321 |
},
|
| 322 |
"prompt-injections": {
|
| 323 |
"attempted": 116,
|
| 324 |
"valid": 116,
|
| 325 |
"failed": 0,
|
| 326 |
"coverage": 1.0,
|
| 327 |
+
"accuracy_all": 0.7844827586206896,
|
| 328 |
+
"accuracy_valid": 0.7844827586206896,
|
| 329 |
+
"ece_top_label": 0.17151379310344822,
|
| 330 |
+
"mean_confidence": 0.9531103448275862,
|
| 331 |
+
"brier_hard": 0.36615182827586207,
|
| 332 |
"brier_hard_n": 116,
|
| 333 |
+
"nll_hard": 1.9254142815070245,
|
| 334 |
"nll_hard_n": 116,
|
| 335 |
+
"zero_probability_gold": 0.05172413793103448,
|
| 336 |
"zero_probability_gold_n": 116,
|
| 337 |
+
"macro_f1": 0.7785414280259642
|
| 338 |
},
|
| 339 |
"massive-intent.en": {
|
| 340 |
"attempted": 300,
|
| 341 |
"valid": 300,
|
| 342 |
"failed": 0,
|
| 343 |
"coverage": 1.0,
|
| 344 |
+
"accuracy_all": 0.7566666666666667,
|
| 345 |
+
"accuracy_valid": 0.7566666666666667,
|
| 346 |
+
"ece_top_label": 0.19351909840865464,
|
| 347 |
+
"mean_confidence": 0.9501857650753213,
|
| 348 |
+
"brier_hard": 0.4290677996735415,
|
| 349 |
"brier_hard_n": 300,
|
| 350 |
+
"nll_hard": 3.2153508464750185,
|
| 351 |
"nll_hard_n": 300,
|
| 352 |
+
"zero_probability_gold": 0.08666666666666667,
|
| 353 |
"zero_probability_gold_n": 300,
|
| 354 |
+
"macro_f1": 0.4943833232448793
|
| 355 |
},
|
| 356 |
"xnli.en": {
|
| 357 |
"attempted": 300,
|
| 358 |
"valid": 300,
|
| 359 |
"failed": 0,
|
| 360 |
"coverage": 1.0,
|
| 361 |
+
"accuracy_all": 0.87,
|
| 362 |
+
"accuracy_valid": 0.87,
|
| 363 |
+
"ece_top_label": 0.07153730747712309,
|
| 364 |
+
"mean_confidence": 0.9147377091848204,
|
| 365 |
+
"brier_hard": 0.22017864577608134,
|
| 366 |
"brier_hard_n": 300,
|
| 367 |
+
"nll_hard": 0.4059236764304257,
|
| 368 |
"nll_hard_n": 300,
|
| 369 |
"zero_probability_gold": 0.0,
|
| 370 |
"zero_probability_gold_n": 300,
|
| 371 |
+
"macro_f1": 0.8710695955086495
|
| 372 |
}
|
| 373 |
}
|
| 374 |
},
|
| 375 |
"2": {
|
| 376 |
"seed": 2,
|
| 377 |
+
"selected_epoch": 13,
|
| 378 |
+
"elapsed_s": 699.7859417049913,
|
| 379 |
"initial_validation": {
|
| 380 |
"soft_cross_entropy": 1.544869564374288,
|
| 381 |
"accuracy": 0.36833333333333335,
|
|
|
|
| 383 |
"decisions": 600
|
| 384 |
},
|
| 385 |
"selected_validation": {
|
| 386 |
+
"soft_cross_entropy": 0.8543428432941437,
|
| 387 |
+
"accuracy": 0.7583333333333333,
|
| 388 |
+
"brier_soft": 0.07187043125430743,
|
| 389 |
"decisions": 600
|
| 390 |
},
|
| 391 |
"stopping": {
|
| 392 |
"reason": "early_stopping",
|
| 393 |
+
"epochs_completed": 16,
|
| 394 |
"patience": 3,
|
| 395 |
"min_epochs": 10,
|
| 396 |
"epochs_without_improvement": 3,
|
|
|
|
| 404 |
"valid": 2000,
|
| 405 |
"failed": 0,
|
| 406 |
"coverage": 1.0,
|
| 407 |
+
"accuracy_all": 0.745,
|
| 408 |
+
"accuracy_valid": 0.745,
|
| 409 |
+
"ece_top_label": 0.18594273050266366,
|
| 410 |
+
"mean_confidence": 0.5591527012331047,
|
| 411 |
+
"brier_hard": 0.40870440390133495,
|
| 412 |
"brier_hard_n": 2000,
|
| 413 |
+
"nll_hard": 0.7205417712302097,
|
| 414 |
"nll_hard_n": 2000,
|
| 415 |
"zero_probability_gold": 0.0,
|
| 416 |
"zero_probability_gold_n": 2000,
|
| 417 |
+
"soft_accuracy": 0.46875735649306255,
|
| 418 |
"soft_accuracy_n": 2000,
|
| 419 |
+
"brier_soft": 0.06782387317668824,
|
| 420 |
"brier_soft_n": 2000,
|
| 421 |
+
"kl_gold_to_prediction": 0.12687181845569698,
|
| 422 |
"kl_gold_to_prediction_n": 2000,
|
| 423 |
+
"total_variation": 0.17935579359201664,
|
| 424 |
"total_variation_n": 2000,
|
| 425 |
+
"score_mae": 0.25587944352517583,
|
| 426 |
"score_mae_n": 800,
|
| 427 |
+
"within_one_level": 0.99,
|
| 428 |
"within_one_level_n": 800,
|
| 429 |
+
"macro_f1": 0.608344198168488
|
| 430 |
},
|
| 431 |
"ag-news": {
|
| 432 |
"attempted": 600,
|
| 433 |
"valid": 600,
|
| 434 |
"failed": 0,
|
| 435 |
"coverage": 1.0,
|
| 436 |
+
"accuracy_all": 0.945,
|
| 437 |
+
"accuracy_valid": 0.945,
|
| 438 |
+
"ece_top_label": 0.03201179295141958,
|
| 439 |
+
"mean_confidence": 0.9180146184255912,
|
| 440 |
+
"brier_hard": 0.08699114920264774,
|
| 441 |
"brier_hard_n": 600,
|
| 442 |
+
"nll_hard": 0.17458965637122142,
|
| 443 |
"nll_hard_n": 600,
|
| 444 |
"zero_probability_gold": 0.0,
|
| 445 |
"zero_probability_gold_n": 600,
|
| 446 |
+
"macro_f1": 0.9421288699146882
|
| 447 |
},
|
| 448 |
"emotion": {
|
| 449 |
"attempted": 600,
|
| 450 |
"valid": 600,
|
| 451 |
"failed": 0,
|
| 452 |
"coverage": 1.0,
|
| 453 |
+
"accuracy_all": 0.56,
|
| 454 |
+
"accuracy_valid": 0.56,
|
| 455 |
+
"ece_top_label": 0.3501647005433899,
|
| 456 |
+
"mean_confidence": 0.9077857005433899,
|
| 457 |
+
"brier_hard": 0.7581475668473461,
|
| 458 |
"brier_hard_n": 600,
|
| 459 |
+
"nll_hard": 2.427279492608694,
|
| 460 |
"nll_hard_n": 600,
|
| 461 |
+
"zero_probability_gold": 0.013333333333333334,
|
| 462 |
"zero_probability_gold_n": 600,
|
| 463 |
+
"macro_f1": 0.4667304743466239
|
| 464 |
},
|
| 465 |
"boolq": {
|
| 466 |
"attempted": 600,
|
| 467 |
"valid": 600,
|
| 468 |
"failed": 0,
|
| 469 |
"coverage": 1.0,
|
| 470 |
+
"accuracy_all": 0.8183333333333334,
|
| 471 |
+
"accuracy_valid": 0.8183333333333334,
|
| 472 |
+
"ece_top_label": 0.08817149999999999,
|
| 473 |
+
"mean_confidence": 0.9065048333333333,
|
| 474 |
+
"brier_hard": 0.2762812369666667,
|
| 475 |
"brier_hard_n": 600,
|
| 476 |
+
"nll_hard": 0.453869636876017,
|
| 477 |
"nll_hard_n": 600,
|
| 478 |
"zero_probability_gold": 0.0,
|
| 479 |
"zero_probability_gold_n": 600,
|
| 480 |
+
"macro_f1": 0.8046122269724754
|
| 481 |
},
|
| 482 |
"sst5": {
|
| 483 |
"attempted": 600,
|
| 484 |
"valid": 600,
|
| 485 |
"failed": 0,
|
| 486 |
"coverage": 1.0,
|
| 487 |
+
"accuracy_all": 0.3983333333333333,
|
| 488 |
+
"accuracy_valid": 0.3983333333333333,
|
| 489 |
+
"ece_top_label": 0.21926419101950614,
|
| 490 |
+
"mean_confidence": 0.6138282869290971,
|
| 491 |
+
"brier_hard": 0.7663539453668811,
|
| 492 |
"brier_hard_n": 600,
|
| 493 |
+
"nll_hard": 1.5341982352412973,
|
| 494 |
"nll_hard_n": 600,
|
| 495 |
"zero_probability_gold": 0.0,
|
| 496 |
"zero_probability_gold_n": 600,
|
| 497 |
+
"score_mae": 0.7954480085445932,
|
| 498 |
"score_mae_n": 600,
|
| 499 |
+
"within_one_level": 0.695,
|
| 500 |
"within_one_level_n": 600,
|
| 501 |
+
"macro_f1": 0.3618869575527743
|
| 502 |
},
|
| 503 |
"prompt-injections": {
|
| 504 |
"attempted": 116,
|
| 505 |
"valid": 116,
|
| 506 |
"failed": 0,
|
| 507 |
"coverage": 1.0,
|
| 508 |
+
"accuracy_all": 0.7155172413793104,
|
| 509 |
+
"accuracy_valid": 0.7155172413793104,
|
| 510 |
+
"ece_top_label": 0.2621448275862069,
|
| 511 |
+
"mean_confidence": 0.9488327586206896,
|
| 512 |
+
"brier_hard": 0.4988274203448276,
|
| 513 |
"brier_hard_n": 116,
|
| 514 |
+
"nll_hard": 2.61023697496577,
|
| 515 |
"nll_hard_n": 116,
|
| 516 |
+
"zero_probability_gold": 0.06896551724137931,
|
| 517 |
"zero_probability_gold_n": 116,
|
| 518 |
+
"macro_f1": 0.6992221261884184
|
| 519 |
},
|
| 520 |
"massive-intent.en": {
|
| 521 |
"attempted": 300,
|
| 522 |
"valid": 300,
|
| 523 |
"failed": 0,
|
| 524 |
"coverage": 1.0,
|
| 525 |
+
"accuracy_all": 0.75,
|
| 526 |
+
"accuracy_valid": 0.75,
|
| 527 |
+
"ece_top_label": 0.20336797564979028,
|
| 528 |
+
"mean_confidence": 0.9533679756497903,
|
| 529 |
+
"brier_hard": 0.4358421604675745,
|
| 530 |
"brier_hard_n": 300,
|
| 531 |
+
"nll_hard": 3.1322756175288426,
|
| 532 |
"nll_hard_n": 300,
|
| 533 |
+
"zero_probability_gold": 0.08,
|
| 534 |
"zero_probability_gold_n": 300,
|
| 535 |
+
"macro_f1": 0.5034639202079357
|
| 536 |
},
|
| 537 |
"xnli.en": {
|
| 538 |
"attempted": 300,
|
| 539 |
"valid": 300,
|
| 540 |
"failed": 0,
|
| 541 |
"coverage": 1.0,
|
| 542 |
+
"accuracy_all": 0.8733333333333333,
|
| 543 |
+
"accuracy_valid": 0.8733333333333333,
|
| 544 |
+
"ece_top_label": 0.05108680216338136,
|
| 545 |
+
"mean_confidence": 0.914837261771426,
|
| 546 |
+
"brier_hard": 0.20561652153986584,
|
| 547 |
"brier_hard_n": 300,
|
| 548 |
+
"nll_hard": 0.3742216981370551,
|
| 549 |
"nll_hard_n": 300,
|
| 550 |
"zero_probability_gold": 0.0,
|
| 551 |
"zero_probability_gold_n": 300,
|
| 552 |
+
"macro_f1": 0.8742255731346232
|
| 553 |
}
|
| 554 |
}
|
| 555 |
},
|
| 556 |
"3": {
|
| 557 |
"seed": 3,
|
| 558 |
+
"selected_epoch": 8,
|
| 559 |
+
"elapsed_s": 495.011001034989,
|
| 560 |
"initial_validation": {
|
| 561 |
"soft_cross_entropy": 1.544869564374288,
|
| 562 |
"accuracy": 0.36833333333333335,
|
|
|
|
| 564 |
"decisions": 600
|
| 565 |
},
|
| 566 |
"selected_validation": {
|
| 567 |
+
"soft_cross_entropy": 0.8545633804798126,
|
| 568 |
+
"accuracy": 0.77,
|
| 569 |
+
"brier_soft": 0.06962224062221746,
|
| 570 |
"decisions": 600
|
| 571 |
},
|
| 572 |
"stopping": {
|
| 573 |
"reason": "early_stopping",
|
| 574 |
+
"epochs_completed": 11,
|
| 575 |
"patience": 3,
|
| 576 |
"min_epochs": 10,
|
| 577 |
"epochs_without_improvement": 3,
|
|
|
|
| 585 |
"valid": 2000,
|
| 586 |
"failed": 0,
|
| 587 |
"coverage": 1.0,
|
| 588 |
+
"accuracy_all": 0.7655,
|
| 589 |
+
"accuracy_valid": 0.7655,
|
| 590 |
+
"ece_top_label": 0.2096452515650509,
|
| 591 |
+
"mean_confidence": 0.5558547484349491,
|
| 592 |
+
"brier_hard": 0.4011842616732685,
|
| 593 |
"brier_hard_n": 2000,
|
| 594 |
+
"nll_hard": 0.7113399953872087,
|
| 595 |
"nll_hard_n": 2000,
|
| 596 |
"zero_probability_gold": 0.0,
|
| 597 |
"zero_probability_gold_n": 2000,
|
| 598 |
+
"soft_accuracy": 0.4683026193430771,
|
| 599 |
"soft_accuracy_n": 2000,
|
| 600 |
+
"brier_soft": 0.06585425455571205,
|
| 601 |
"brier_soft_n": 2000,
|
| 602 |
+
"kl_gold_to_prediction": 0.12507067261764912,
|
| 603 |
"kl_gold_to_prediction_n": 2000,
|
| 604 |
+
"total_variation": 0.17848467876854548,
|
| 605 |
"total_variation_n": 2000,
|
| 606 |
+
"score_mae": 0.24395106365305969,
|
| 607 |
"score_mae_n": 800,
|
| 608 |
+
"within_one_level": 0.98375,
|
| 609 |
"within_one_level_n": 800,
|
| 610 |
+
"macro_f1": 0.6399493527359326
|
| 611 |
},
|
| 612 |
"ag-news": {
|
| 613 |
"attempted": 600,
|
| 614 |
"valid": 600,
|
| 615 |
"failed": 0,
|
| 616 |
"coverage": 1.0,
|
| 617 |
+
"accuracy_all": 0.9366666666666666,
|
| 618 |
+
"accuracy_valid": 0.9366666666666666,
|
| 619 |
+
"ece_top_label": 0.0312034834128748,
|
| 620 |
+
"mean_confidence": 0.914535932360556,
|
| 621 |
+
"brier_hard": 0.09381807669545053,
|
| 622 |
"brier_hard_n": 600,
|
| 623 |
+
"nll_hard": 0.1866731647377733,
|
| 624 |
"nll_hard_n": 600,
|
| 625 |
"zero_probability_gold": 0.0,
|
| 626 |
"zero_probability_gold_n": 600,
|
| 627 |
+
"macro_f1": 0.9336084571946038
|
| 628 |
},
|
| 629 |
"emotion": {
|
| 630 |
"attempted": 600,
|
| 631 |
"valid": 600,
|
| 632 |
"failed": 0,
|
| 633 |
"coverage": 1.0,
|
| 634 |
+
"accuracy_all": 0.565,
|
| 635 |
+
"accuracy_valid": 0.565,
|
| 636 |
+
"ece_top_label": 0.3370553794856075,
|
| 637 |
+
"mean_confidence": 0.8987647229086076,
|
| 638 |
+
"brier_hard": 0.7489678137408704,
|
| 639 |
"brier_hard_n": 600,
|
| 640 |
+
"nll_hard": 2.3318777367931305,
|
| 641 |
"nll_hard_n": 600,
|
| 642 |
+
"zero_probability_gold": 0.011666666666666667,
|
| 643 |
"zero_probability_gold_n": 600,
|
| 644 |
+
"macro_f1": 0.4760859469034663
|
| 645 |
},
|
| 646 |
"boolq": {
|
| 647 |
"attempted": 600,
|
| 648 |
"valid": 600,
|
| 649 |
"failed": 0,
|
| 650 |
"coverage": 1.0,
|
| 651 |
+
"accuracy_all": 0.82,
|
| 652 |
+
"accuracy_valid": 0.82,
|
| 653 |
+
"ece_top_label": 0.09564866666666666,
|
| 654 |
+
"mean_confidence": 0.9069513333333333,
|
| 655 |
+
"brier_hard": 0.2824305872,
|
| 656 |
"brier_hard_n": 600,
|
| 657 |
+
"nll_hard": 0.4619444556248629,
|
| 658 |
"nll_hard_n": 600,
|
| 659 |
"zero_probability_gold": 0.0,
|
| 660 |
"zero_probability_gold_n": 600,
|
| 661 |
+
"macro_f1": 0.8069498069498069
|
| 662 |
},
|
| 663 |
"sst5": {
|
| 664 |
"attempted": 600,
|
| 665 |
"valid": 600,
|
| 666 |
"failed": 0,
|
| 667 |
"coverage": 1.0,
|
| 668 |
+
"accuracy_all": 0.4116666666666667,
|
| 669 |
+
"accuracy_valid": 0.4116666666666667,
|
| 670 |
+
"ece_top_label": 0.19022023187751014,
|
| 671 |
+
"mean_confidence": 0.6018868985441768,
|
| 672 |
+
"brier_hard": 0.7571439809814903,
|
| 673 |
"brier_hard_n": 600,
|
| 674 |
+
"nll_hard": 1.5683213482322946,
|
| 675 |
"nll_hard_n": 600,
|
| 676 |
"zero_probability_gold": 0.0,
|
| 677 |
"zero_probability_gold_n": 600,
|
| 678 |
+
"score_mae": 0.8321864596740475,
|
| 679 |
"score_mae_n": 600,
|
| 680 |
+
"within_one_level": 0.6783333333333333,
|
| 681 |
"within_one_level_n": 600,
|
| 682 |
+
"macro_f1": 0.3816445513227996
|
| 683 |
},
|
| 684 |
"prompt-injections": {
|
| 685 |
"attempted": 116,
|
| 686 |
"valid": 116,
|
| 687 |
"failed": 0,
|
| 688 |
"coverage": 1.0,
|
| 689 |
+
"accuracy_all": 0.7844827586206896,
|
| 690 |
+
"accuracy_valid": 0.7844827586206896,
|
| 691 |
+
"ece_top_label": 0.17475344827586206,
|
| 692 |
+
"mean_confidence": 0.9523568965517242,
|
| 693 |
+
"brier_hard": 0.3675763103448276,
|
| 694 |
"brier_hard_n": 116,
|
| 695 |
+
"nll_hard": 1.8418113090371786,
|
| 696 |
"nll_hard_n": 116,
|
| 697 |
+
"zero_probability_gold": 0.04310344827586207,
|
| 698 |
"zero_probability_gold_n": 116,
|
| 699 |
+
"macro_f1": 0.7797524113313588
|
| 700 |
},
|
| 701 |
"massive-intent.en": {
|
| 702 |
"attempted": 300,
|
| 703 |
"valid": 300,
|
| 704 |
"failed": 0,
|
| 705 |
"coverage": 1.0,
|
| 706 |
+
"accuracy_all": 0.7733333333333333,
|
| 707 |
+
"accuracy_valid": 0.7733333333333333,
|
| 708 |
+
"ece_top_label": 0.17472328670404882,
|
| 709 |
+
"mean_confidence": 0.9480566200373822,
|
| 710 |
+
"brier_hard": 0.38383190196466177,
|
| 711 |
"brier_hard_n": 300,
|
| 712 |
+
"nll_hard": 2.4090375945333213,
|
| 713 |
"nll_hard_n": 300,
|
| 714 |
+
"zero_probability_gold": 0.05,
|
| 715 |
"zero_probability_gold_n": 300,
|
| 716 |
+
"macro_f1": 0.502438924453748
|
| 717 |
},
|
| 718 |
"xnli.en": {
|
| 719 |
"attempted": 300,
|
| 720 |
"valid": 300,
|
| 721 |
"failed": 0,
|
| 722 |
"coverage": 1.0,
|
| 723 |
+
"accuracy_all": 0.8666666666666667,
|
| 724 |
+
"accuracy_valid": 0.8666666666666667,
|
| 725 |
+
"ece_top_label": 0.054425126999604626,
|
| 726 |
+
"mean_confidence": 0.9150511806609052,
|
| 727 |
+
"brier_hard": 0.2024701804692983,
|
| 728 |
"brier_hard_n": 300,
|
| 729 |
+
"nll_hard": 0.3723722500840691,
|
| 730 |
"nll_hard_n": 300,
|
| 731 |
"zero_probability_gold": 0.0,
|
| 732 |
"zero_probability_gold_n": 300,
|
| 733 |
+
"macro_f1": 0.8674417053037536
|
| 734 |
}
|
| 735 |
}
|
| 736 |
},
|
| 737 |
"4": {
|
| 738 |
"seed": 4,
|
| 739 |
+
"selected_epoch": 8,
|
| 740 |
+
"elapsed_s": 496.1691781419795,
|
| 741 |
"initial_validation": {
|
| 742 |
"soft_cross_entropy": 1.544869564374288,
|
| 743 |
"accuracy": 0.36833333333333335,
|
|
|
|
| 745 |
"decisions": 600
|
| 746 |
},
|
| 747 |
"selected_validation": {
|
| 748 |
+
"soft_cross_entropy": 0.8540202794472377,
|
| 749 |
+
"accuracy": 0.76,
|
| 750 |
+
"brier_soft": 0.07019127607345581,
|
| 751 |
"decisions": 600
|
| 752 |
},
|
| 753 |
"stopping": {
|
| 754 |
"reason": "early_stopping",
|
| 755 |
+
"epochs_completed": 11,
|
| 756 |
"patience": 3,
|
| 757 |
"min_epochs": 10,
|
| 758 |
"epochs_without_improvement": 3,
|
|
|
|
| 766 |
"valid": 2000,
|
| 767 |
"failed": 0,
|
| 768 |
"coverage": 1.0,
|
| 769 |
+
"accuracy_all": 0.7545,
|
| 770 |
+
"accuracy_valid": 0.7545,
|
| 771 |
+
"ece_top_label": 0.2042592508573846,
|
| 772 |
+
"mean_confidence": 0.5504920240151281,
|
| 773 |
+
"brier_hard": 0.4113014025244486,
|
| 774 |
"brier_hard_n": 2000,
|
| 775 |
+
"nll_hard": 0.7247566141198507,
|
| 776 |
"nll_hard_n": 2000,
|
| 777 |
"zero_probability_gold": 0.0,
|
| 778 |
"zero_probability_gold_n": 2000,
|
| 779 |
+
"soft_accuracy": 0.46590212671846004,
|
| 780 |
"soft_accuracy_n": 2000,
|
| 781 |
+
"brier_soft": 0.06812301484891332,
|
| 782 |
"brier_soft_n": 2000,
|
| 783 |
+
"kl_gold_to_prediction": 0.12790407031921358,
|
| 784 |
"kl_gold_to_prediction_n": 2000,
|
| 785 |
+
"total_variation": 0.18054231568172208,
|
| 786 |
"total_variation_n": 2000,
|
| 787 |
+
"score_mae": 0.25377940453843617,
|
| 788 |
"score_mae_n": 800,
|
| 789 |
+
"within_one_level": 0.98375,
|
| 790 |
"within_one_level_n": 800,
|
| 791 |
+
"macro_f1": 0.6337422972118567
|
| 792 |
},
|
| 793 |
"ag-news": {
|
| 794 |
"attempted": 600,
|
| 795 |
"valid": 600,
|
| 796 |
"failed": 0,
|
| 797 |
"coverage": 1.0,
|
| 798 |
+
"accuracy_all": 0.9516666666666667,
|
| 799 |
+
"accuracy_valid": 0.9516666666666667,
|
| 800 |
+
"ece_top_label": 0.03233758800344756,
|
| 801 |
+
"mean_confidence": 0.9199325508493339,
|
| 802 |
+
"brier_hard": 0.08178925980745573,
|
| 803 |
"brier_hard_n": 600,
|
| 804 |
+
"nll_hard": 0.16974235809511878,
|
| 805 |
"nll_hard_n": 600,
|
| 806 |
"zero_probability_gold": 0.0,
|
| 807 |
"zero_probability_gold_n": 600,
|
| 808 |
+
"macro_f1": 0.9489828512364638
|
| 809 |
},
|
| 810 |
"emotion": {
|
| 811 |
"attempted": 600,
|
| 812 |
"valid": 600,
|
| 813 |
"failed": 0,
|
| 814 |
"coverage": 1.0,
|
| 815 |
+
"accuracy_all": 0.5783333333333334,
|
| 816 |
+
"accuracy_valid": 0.5783333333333334,
|
| 817 |
+
"ece_top_label": 0.3295382068772924,
|
| 818 |
+
"mean_confidence": 0.9035819218093762,
|
| 819 |
+
"brier_hard": 0.7425345488860924,
|
| 820 |
"brier_hard_n": 600,
|
| 821 |
+
"nll_hard": 2.3910245621280453,
|
| 822 |
"nll_hard_n": 600,
|
| 823 |
+
"zero_probability_gold": 0.013333333333333334,
|
| 824 |
"zero_probability_gold_n": 600,
|
| 825 |
+
"macro_f1": 0.48527808236443093
|
| 826 |
},
|
| 827 |
"boolq": {
|
| 828 |
"attempted": 600,
|
| 829 |
"valid": 600,
|
| 830 |
"failed": 0,
|
| 831 |
"coverage": 1.0,
|
| 832 |
+
"accuracy_all": 0.8316666666666667,
|
| 833 |
+
"accuracy_valid": 0.8316666666666667,
|
| 834 |
+
"ece_top_label": 0.09521133333333336,
|
| 835 |
+
"mean_confidence": 0.9097233333333333,
|
| 836 |
+
"brier_hard": 0.27490432886666666,
|
| 837 |
"brier_hard_n": 600,
|
| 838 |
+
"nll_hard": 0.49707214589381826,
|
| 839 |
"nll_hard_n": 600,
|
| 840 |
+
"zero_probability_gold": 0.0016666666666666668,
|
| 841 |
"zero_probability_gold_n": 600,
|
| 842 |
+
"macro_f1": 0.8196294367140413
|
| 843 |
},
|
| 844 |
"sst5": {
|
| 845 |
"attempted": 600,
|
| 846 |
"valid": 600,
|
| 847 |
"failed": 0,
|
| 848 |
"coverage": 1.0,
|
| 849 |
+
"accuracy_all": 0.41,
|
| 850 |
+
"accuracy_valid": 0.41,
|
| 851 |
+
"ece_top_label": 0.20145470242332547,
|
| 852 |
+
"mean_confidence": 0.6098530357566588,
|
| 853 |
+
"brier_hard": 0.7512849781669847,
|
| 854 |
"brier_hard_n": 600,
|
| 855 |
+
"nll_hard": 1.5195838995022448,
|
| 856 |
"nll_hard_n": 600,
|
| 857 |
"zero_probability_gold": 0.0,
|
| 858 |
"zero_probability_gold_n": 600,
|
| 859 |
+
"score_mae": 0.8012926973404637,
|
| 860 |
"score_mae_n": 600,
|
| 861 |
+
"within_one_level": 0.7016666666666667,
|
| 862 |
"within_one_level_n": 600,
|
| 863 |
+
"macro_f1": 0.37608916605464326
|
| 864 |
},
|
| 865 |
"prompt-injections": {
|
| 866 |
"attempted": 116,
|
| 867 |
"valid": 116,
|
| 868 |
"failed": 0,
|
| 869 |
"coverage": 1.0,
|
| 870 |
+
"accuracy_all": 0.7413793103448276,
|
| 871 |
+
"accuracy_valid": 0.7413793103448276,
|
| 872 |
+
"ece_top_label": 0.21426120689655168,
|
| 873 |
+
"mean_confidence": 0.9556405172413793,
|
| 874 |
+
"brier_hard": 0.43549763362068966,
|
| 875 |
"brier_hard_n": 116,
|
| 876 |
+
"nll_hard": 2.2934410633534186,
|
| 877 |
"nll_hard_n": 116,
|
| 878 |
+
"zero_probability_gold": 0.0603448275862069,
|
| 879 |
"zero_probability_gold_n": 116,
|
| 880 |
+
"macro_f1": 0.7276995305164319
|
| 881 |
},
|
| 882 |
"massive-intent.en": {
|
| 883 |
"attempted": 300,
|
| 884 |
"valid": 300,
|
| 885 |
"failed": 0,
|
| 886 |
"coverage": 1.0,
|
| 887 |
+
"accuracy_all": 0.77,
|
| 888 |
+
"accuracy_valid": 0.77,
|
| 889 |
+
"ece_top_label": 0.17971198952007902,
|
| 890 |
+
"mean_confidence": 0.9447345124535086,
|
| 891 |
+
"brier_hard": 0.4058348152163361,
|
| 892 |
"brier_hard_n": 300,
|
| 893 |
+
"nll_hard": 2.9106532352305194,
|
| 894 |
"nll_hard_n": 300,
|
| 895 |
+
"zero_probability_gold": 0.07333333333333333,
|
| 896 |
"zero_probability_gold_n": 300,
|
| 897 |
+
"macro_f1": 0.5030980723143226
|
| 898 |
},
|
| 899 |
"xnli.en": {
|
| 900 |
"attempted": 300,
|
| 901 |
"valid": 300,
|
| 902 |
"failed": 0,
|
| 903 |
"coverage": 1.0,
|
| 904 |
+
"accuracy_all": 0.8633333333333333,
|
| 905 |
+
"accuracy_valid": 0.8633333333333333,
|
| 906 |
+
"ece_top_label": 0.0751956003124493,
|
| 907 |
+
"mean_confidence": 0.9162548569668819,
|
| 908 |
+
"brier_hard": 0.22384958324541668,
|
| 909 |
"brier_hard_n": 300,
|
| 910 |
+
"nll_hard": 0.4127381776210696,
|
| 911 |
"nll_hard_n": 300,
|
| 912 |
"zero_probability_gold": 0.0,
|
| 913 |
"zero_probability_gold_n": 300,
|
| 914 |
+
"macro_f1": 0.864378848443158
|
| 915 |
}
|
| 916 |
}
|
| 917 |
}
|
|
|
|
| 919 |
"summary": {
|
| 920 |
"typed-decisions": {
|
| 921 |
"n": 5,
|
| 922 |
+
"mean_percent": 75.28,
|
| 923 |
+
"sample_variance_pp_squared": 0.8619999999999957,
|
| 924 |
+
"sample_std_pp": 0.9284395510748105
|
| 925 |
},
|
| 926 |
"ag-news": {
|
| 927 |
"n": 5,
|
| 928 |
+
"mean_percent": 94.46666666666667,
|
| 929 |
+
"sample_variance_pp_squared": 0.3388888888888897,
|
| 930 |
+
"sample_std_pp": 0.5821416398857667
|
| 931 |
},
|
| 932 |
"boolq": {
|
| 933 |
"n": 5,
|
| 934 |
+
"mean_percent": 82.10000000000001,
|
| 935 |
+
"sample_variance_pp_squared": 0.7861111111111168,
|
| 936 |
+
"sample_std_pp": 0.8866290718846956
|
| 937 |
},
|
| 938 |
"emotion": {
|
| 939 |
"n": 5,
|
| 940 |
+
"mean_percent": 57.2,
|
| 941 |
+
"sample_variance_pp_squared": 1.0055555555555546,
|
| 942 |
+
"sample_std_pp": 1.0027739304327543
|
| 943 |
},
|
| 944 |
"prompt-injections": {
|
| 945 |
"n": 5,
|
| 946 |
+
"mean_percent": 74.48275862068965,
|
| 947 |
+
"sample_variance_pp_squared": 15.457788347205712,
|
| 948 |
+
"sample_std_pp": 3.93163939689358
|
| 949 |
},
|
| 950 |
"sst5": {
|
| 951 |
"n": 5,
|
| 952 |
+
"mean_percent": 41.0,
|
| 953 |
+
"sample_variance_pp_squared": 6.027777777777775,
|
| 954 |
+
"sample_std_pp": 2.4551533104427055
|
| 955 |
},
|
| 956 |
"massive-intent.en": {
|
| 957 |
"n": 5,
|
| 958 |
+
"mean_percent": 76.8,
|
| 959 |
+
"sample_variance_pp_squared": 2.422222222222218,
|
| 960 |
+
"sample_std_pp": 1.556349003990499
|
| 961 |
},
|
| 962 |
"xnli.en": {
|
| 963 |
"n": 5,
|
| 964 |
+
"mean_percent": 87.26666666666667,
|
| 965 |
+
"sample_variance_pp_squared": 1.0777777777777784,
|
| 966 |
+
"sample_std_pp": 1.038160766826496
|
| 967 |
}
|
| 968 |
},
|
| 969 |
+
"candidate_seed": 4,
|
| 970 |
"variance_definition": "sample variance, n-1, percentage points squared",
|
| 971 |
"benchmark_manifest": {
|
| 972 |
"format_version": 1,
|
release_manifest.json
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"published": false,
|
| 3 |
+
"review_status": "export verified; package prepared for user review and upload",
|
| 4 |
+
"files_sha256": {
|
| 5 |
+
"LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
|
| 6 |
+
"README.md": "db6038098fae3817f20fdabc66afb13d7e7cd1306d0b611eea7f70ef84ad08bc",
|
| 7 |
+
"THIRD_PARTY.md": "3dd2683344d45ce64c15f45c723f4b0db720deee84446365fdf4084a73ebeceb",
|
| 8 |
+
"USAGE.md": "038e1d21dca7d2e2da5b086a34e5b446edb8901406609c8c8b0e3407f31ab7d9",
|
| 9 |
+
"adapter.safetensors": "f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081",
|
| 10 |
+
"adapter_config.json": "b26f0cb6ea6c6f1e66f254b832d5000b245d309fc05880f234d0f7489cadb54a",
|
| 11 |
+
"ariadne_bench/__init__.py": "a73c9fdf21413382cb93b523bfb3925aa3fd354e18faeead4650e86a30645411",
|
| 12 |
+
"ariadne_bench/__main__.py": "935a1c1166b0c1ea35a82256345000bf2c73ded718d77773bc27a71ecce28f7d",
|
| 13 |
+
"ariadne_bench/adapters.py": "3bd4820165be077b3b98fe83ac23e649e1ad69186ad5f2a4f5ee7addf1427cf5",
|
| 14 |
+
"ariadne_bench/cli.py": "de267c0502489eb4bb37e4b8f6faa8da8be37a56ced270f754424a40404d1af6",
|
| 15 |
+
"ariadne_bench/datasets.py": "31c722c64895519a4635b501da88862b446384682e3e8ee884495941d6990912",
|
| 16 |
+
"ariadne_bench/experiments/__init__.py": "95ad0a7f30c4e0a194d0a6e2ae8905abae5d622edbb3d4e5d791b4bc460aee5d",
|
| 17 |
+
"ariadne_bench/experiments/align.py": "31e2733dcc223a1e1a6250d41c4158f84091f1cb4067d61388e06b5b01ae34c4",
|
| 18 |
+
"ariadne_bench/experiments/prepare_alignment.py": "c4f39c10ce46ee13c0d89128d18a85ec4a86729e81e070f5834ed01fc1c405a9",
|
| 19 |
+
"ariadne_bench/experiments/spectrum.py": "22d4801b585c539d0d579a9d2a00076e1002fd337bfe31d1906d8ae3fb8d4320",
|
| 20 |
+
"ariadne_bench/frozen_input_interface.py": "253f0186dac6e4a7f4b0fcf3c33449c3b3882cf8be538dca6eeca8d9e4ce9b83",
|
| 21 |
+
"ariadne_bench/full_finetune.py": "6067be8cf837d254b919c841c9a9851e1546a413dbf6324ff242e5d2f56c9734",
|
| 22 |
+
"ariadne_bench/full_input_interface.py": "a8f440a2185085c64896300a243ac52a9619c4d05a1b4e43c50c2c32aa1da996",
|
| 23 |
+
"ariadne_bench/hybrid.py": "90fbbdab9b8157a32854d27c2d93089be6dc88dfe9320426f102eae4e03d9ffd",
|
| 24 |
+
"ariadne_bench/hybrid_control.py": "1b53678bea73fe7608570b78d7215d1dff2e38d6ad3627615f26c06ae74d8b4c",
|
| 25 |
+
"ariadne_bench/interfaces.py": "19700fb8170b2460401900cb96ebdff22bd7bf778ca5132e98d56ba68f344187",
|
| 26 |
+
"ariadne_bench/metrics.py": "fc6fde20c95d5052de52d05d6fb9b6eb69b780c46966f5bf1e517738b276a4b4",
|
| 27 |
+
"ariadne_bench/portable.py": "259f633a8bb6663d0594d9639ea8d80ef493f5a63d4b11444a4003f2f9a318e8",
|
| 28 |
+
"ariadne_bench/reproducibility.py": "5b0c161c596f25277b3326c645df53d5d1f2dc0272da9e77bc4a4c0d56188b6e",
|
| 29 |
+
"ariadne_bench/runner.py": "b20e2af4ba254e2d48a07563356f95728c4807018c3c989b429e9922674c64cd",
|
| 30 |
+
"ariadne_bench/schema.py": "67794f1172c1ef9d9625b84c8096fd48b8fa7cd10e3f88fc13ba3a98ef7c7a76",
|
| 31 |
+
"benchmarks/sources.lock.json": "6e6eaab2778092c8127257fce202c9fa8d4625dd0d029490dd9f137da91d134c",
|
| 32 |
+
"environment.json": "c5b9c796d97f3dec6d242caf79aa5d2ac69c91717f2be6ed5a51627bf6d58c17",
|
| 33 |
+
"evidence/adapter-equivalence.json": "7ffdc6216702ab5de02a1745217b369de9f2f8934e16e5c80dd090a3343e9e84",
|
| 34 |
+
"evidence/audit.json": "7a2954a470b3f94626ca678aef7f9bf1f13551be20e642c1a16a0a90b751bd4d",
|
| 35 |
+
"evidence/candidate-uncertainty.json": "9e97410249531de151864285412ab3bd337b117b3573f6439cf74ae17926cae3",
|
| 36 |
+
"evidence/identity-check.json": "0cc8904dc58e1e0ed49496c7ae01980c52e3a1395d61ef91286cec2c9dcdfa25",
|
| 37 |
+
"evidence/lr-comparison.json": "724a08155ff3f39a335132cffb0248fdb08f06874e84ce6a582c8af9199ac69d",
|
| 38 |
+
"evidence/replay-check.json": "d362a58603077feddc4ad2513c19f8caea25da1a9ab03eec25c5be5c14c3eec0",
|
| 39 |
+
"evidence/seed-0-history.json": "1090a62e8f1256e9510f8c2c913de14954a5e54e14396a962578ebef83b4a6f3",
|
| 40 |
+
"evidence/seed-1-history.json": "6aee05fe087ec5d970329743a06c1cbff3e2a0e8209cdcb2db41649b66b75432",
|
| 41 |
+
"evidence/seed-2-history.json": "71afe93da2e39ef81417074bc1a4559b9d6f070f677e76f74818129633cf3c33",
|
| 42 |
+
"evidence/seed-3-history.json": "4bf723c1659ffc8adeb57d06d51132f4f5af74c811c98feef71de4f6fa3bd66f",
|
| 43 |
+
"evidence/seed-4-history.json": "2609c807813f55ffe5aef7e6c96aafefe1647de1ac589568d7fdc56e46f5c9fe",
|
| 44 |
+
"export-verification.json": "65e46623219d80979a7108e69c0ed89d0b1d5c23a82c4255fee64647694aac3f",
|
| 45 |
+
"licenses/jev-benchmarks-LICENSE": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
|
| 46 |
+
"licenses/laya-LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
|
| 47 |
+
"load_adapter.py": "2f305ef95fef2b2482e9214836fc0071a59cfd6c6d3aaf42cebdbfb131272081",
|
| 48 |
+
"make_training_config.py": "23e9d4acb90141a3d690bc607c861df92137b9457304ae19b8a925398b55536d",
|
| 49 |
+
"metrics.json": "e3cd35993afc309f7ccd99b1a3d96316944e1ab42fd489c7f15465f43ec4a78b",
|
| 50 |
+
"requirements.txt": "285a8ae6529abeefd058f2e9679e0e314b19fda38169e685ea16aec59b053bd9",
|
| 51 |
+
"training_protocol.json": "00b2db7abbd07ba3220a130fe36ed64958f1dfac15773d183f1ae8dc9d8334b3",
|
| 52 |
+
"upload-manifest.json": "613b92d535a5c8fca8dfe1d6535600f6da68ba0bf110b53829d13638614830b6"
|
| 53 |
+
}
|
| 54 |
+
}
|
training_protocol.json
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
{
|
| 2 |
-
"experiment": "
|
| 3 |
"seeds": [
|
| 4 |
0,
|
| 5 |
1,
|
|
@@ -7,7 +7,7 @@
|
|
| 7 |
3,
|
| 8 |
4
|
| 9 |
],
|
| 10 |
-
"created_at": "2026-09-
|
| 11 |
"data": {
|
| 12 |
"train": {
|
| 13 |
"format_version": 1,
|
|
@@ -67,7 +67,7 @@
|
|
| 67 |
"microbatch": 8,
|
| 68 |
"accumulation": 4,
|
| 69 |
"effective_batch": 32,
|
| 70 |
-
"interface_lr": 0.
|
| 71 |
"weight_decay": 0.01,
|
| 72 |
"gradient_clip": 1,
|
| 73 |
"training_scope": "bridge",
|
|
@@ -77,8 +77,23 @@
|
|
| 77 |
"device": "cuda:1",
|
| 78 |
"deterministic": true,
|
| 79 |
"native_sdk_calibration": true,
|
| 80 |
-
"publish_policy": "
|
| 81 |
"baseline_predictions_sha256": "37dff666a00a340cd6e37f775dc60df33a3c6eb78f76bfee7f094443c225b6e8",
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
"base_model": "convaiinnovations/laya",
|
| 83 |
"base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
|
| 84 |
"training_sources_sha256": {
|
|
|
|
| 1 |
{
|
| 2 |
+
"experiment": "E0j five-seed confirmation",
|
| 3 |
"seeds": [
|
| 4 |
0,
|
| 5 |
1,
|
|
|
|
| 7 |
3,
|
| 8 |
4
|
| 9 |
],
|
| 10 |
+
"created_at": "2026-09-27T15:29:06.749690+00:00",
|
| 11 |
"data": {
|
| 12 |
"train": {
|
| 13 |
"format_version": 1,
|
|
|
|
| 67 |
"microbatch": 8,
|
| 68 |
"accumulation": 4,
|
| 69 |
"effective_batch": 32,
|
| 70 |
+
"interface_lr": 0.0001,
|
| 71 |
"weight_decay": 0.01,
|
| 72 |
"gradient_clip": 1,
|
| 73 |
"training_scope": "bridge",
|
|
|
|
| 77 |
"device": "cuda:1",
|
| 78 |
"deterministic": true,
|
| 79 |
"native_sdk_calibration": true,
|
| 80 |
+
"publish_policy": "provide user-requested local bundle; preserve existing card; no assistant upload",
|
| 81 |
"baseline_predictions_sha256": "37dff666a00a340cd6e37f775dc60df33a3c6eb78f76bfee7f094443c225b6e8",
|
| 82 |
+
"run_order": [
|
| 83 |
+
0,
|
| 84 |
+
2,
|
| 85 |
+
3,
|
| 86 |
+
1,
|
| 87 |
+
4
|
| 88 |
+
],
|
| 89 |
+
"seed_0_reference": "completed controlled pilot, unchanged",
|
| 90 |
+
"comparison": "LR is the only training change from E0i; min10/patience3 and all other settings fixed",
|
| 91 |
+
"scope": "adaptive experimental follow-up selected after the seed-0 pilot; not a fresh blind test",
|
| 92 |
+
"replay_evidence": "Two fresh seed-0 LR 1e-4 processes matched losses, metrics and tensors for two epochs; reported seed-0 metric prefix matched. Verified after the sweep.",
|
| 93 |
+
"driver_sha256": {
|
| 94 |
+
"/mnt/storage/ariadne/scripts/run_frozen_linear_lr_confirmation.py": "f7c3841277682cbae43f4eb51fa3cb5fc87fd53068607206240ce137b7e94a9b"
|
| 95 |
+
},
|
| 96 |
+
"release_replay_sha256": "dd3479114e8af11d06708febc1e86b887cc6e3cc5ce1523a249e54b9432412ec",
|
| 97 |
"base_model": "convaiinnovations/laya",
|
| 98 |
"base_revision": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
|
| 99 |
"training_sources_sha256": {
|
upload-manifest.json
CHANGED
|
@@ -1,17 +1,17 @@
|
|
| 1 |
{
|
| 2 |
"published_by_assistant": false,
|
| 3 |
-
"candidate_seed":
|
| 4 |
-
"selected_epoch":
|
| 5 |
"format": "compact affine adapter requiring pinned base Laya",
|
| 6 |
-
"
|
| 7 |
-
"
|
| 8 |
-
"removed_only_other_seed_loading_entries": true,
|
| 9 |
"files_sha256": {
|
| 10 |
"LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
|
|
|
|
| 11 |
"THIRD_PARTY.md": "3dd2683344d45ce64c15f45c723f4b0db720deee84446365fdf4084a73ebeceb",
|
| 12 |
-
"USAGE.md": "
|
| 13 |
-
"adapter.safetensors": "
|
| 14 |
-
"adapter_config.json": "
|
| 15 |
"ariadne_bench/__init__.py": "a73c9fdf21413382cb93b523bfb3925aa3fd354e18faeead4650e86a30645411",
|
| 16 |
"ariadne_bench/__main__.py": "935a1c1166b0c1ea35a82256345000bf2c73ded718d77773bc27a71ecce28f7d",
|
| 17 |
"ariadne_bench/adapters.py": "3bd4820165be077b3b98fe83ac23e649e1ad69186ad5f2a4f5ee7addf1427cf5",
|
|
@@ -34,22 +34,24 @@
|
|
| 34 |
"ariadne_bench/schema.py": "67794f1172c1ef9d9625b84c8096fd48b8fa7cd10e3f88fc13ba3a98ef7c7a76",
|
| 35 |
"benchmarks/sources.lock.json": "6e6eaab2778092c8127257fce202c9fa8d4625dd0d029490dd9f137da91d134c",
|
| 36 |
"environment.json": "c5b9c796d97f3dec6d242caf79aa5d2ac69c91717f2be6ed5a51627bf6d58c17",
|
| 37 |
-
"evidence/adapter-equivalence.json": "
|
| 38 |
-
"evidence/audit.json": "
|
| 39 |
-
"evidence/candidate-uncertainty.json": "
|
| 40 |
"evidence/identity-check.json": "0cc8904dc58e1e0ed49496c7ae01980c52e3a1395d61ef91286cec2c9dcdfa25",
|
| 41 |
-
"evidence/
|
| 42 |
-
"evidence/
|
| 43 |
-
"evidence/seed-
|
| 44 |
-
"evidence/seed-
|
| 45 |
-
"evidence/seed-
|
| 46 |
-
"evidence/seed-
|
| 47 |
-
"
|
|
|
|
| 48 |
"licenses/jev-benchmarks-LICENSE": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
|
| 49 |
"licenses/laya-LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
|
| 50 |
"load_adapter.py": "2f305ef95fef2b2482e9214836fc0071a59cfd6c6d3aaf42cebdbfb131272081",
|
| 51 |
-
"
|
|
|
|
| 52 |
"requirements.txt": "285a8ae6529abeefd058f2e9679e0e314b19fda38169e685ea16aec59b053bd9",
|
| 53 |
-
"training_protocol.json": "
|
| 54 |
}
|
| 55 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"published_by_assistant": false,
|
| 3 |
+
"candidate_seed": 4,
|
| 4 |
+
"selected_epoch": 8,
|
| 5 |
"format": "compact affine adapter requiring pinned base Laya",
|
| 6 |
+
"includes_model_card": true,
|
| 7 |
+
"all_5116_export_answers_exact": true,
|
|
|
|
| 8 |
"files_sha256": {
|
| 9 |
"LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
|
| 10 |
+
"README.md": "db6038098fae3817f20fdabc66afb13d7e7cd1306d0b611eea7f70ef84ad08bc",
|
| 11 |
"THIRD_PARTY.md": "3dd2683344d45ce64c15f45c723f4b0db720deee84446365fdf4084a73ebeceb",
|
| 12 |
+
"USAGE.md": "038e1d21dca7d2e2da5b086a34e5b446edb8901406609c8c8b0e3407f31ab7d9",
|
| 13 |
+
"adapter.safetensors": "f23ad342a969773bbbcb18619ce8a8843bbb105d94fe64c5a837d68f26829081",
|
| 14 |
+
"adapter_config.json": "b26f0cb6ea6c6f1e66f254b832d5000b245d309fc05880f234d0f7489cadb54a",
|
| 15 |
"ariadne_bench/__init__.py": "a73c9fdf21413382cb93b523bfb3925aa3fd354e18faeead4650e86a30645411",
|
| 16 |
"ariadne_bench/__main__.py": "935a1c1166b0c1ea35a82256345000bf2c73ded718d77773bc27a71ecce28f7d",
|
| 17 |
"ariadne_bench/adapters.py": "3bd4820165be077b3b98fe83ac23e649e1ad69186ad5f2a4f5ee7addf1427cf5",
|
|
|
|
| 34 |
"ariadne_bench/schema.py": "67794f1172c1ef9d9625b84c8096fd48b8fa7cd10e3f88fc13ba3a98ef7c7a76",
|
| 35 |
"benchmarks/sources.lock.json": "6e6eaab2778092c8127257fce202c9fa8d4625dd0d029490dd9f137da91d134c",
|
| 36 |
"environment.json": "c5b9c796d97f3dec6d242caf79aa5d2ac69c91717f2be6ed5a51627bf6d58c17",
|
| 37 |
+
"evidence/adapter-equivalence.json": "7ffdc6216702ab5de02a1745217b369de9f2f8934e16e5c80dd090a3343e9e84",
|
| 38 |
+
"evidence/audit.json": "7a2954a470b3f94626ca678aef7f9bf1f13551be20e642c1a16a0a90b751bd4d",
|
| 39 |
+
"evidence/candidate-uncertainty.json": "9e97410249531de151864285412ab3bd337b117b3573f6439cf74ae17926cae3",
|
| 40 |
"evidence/identity-check.json": "0cc8904dc58e1e0ed49496c7ae01980c52e3a1395d61ef91286cec2c9dcdfa25",
|
| 41 |
+
"evidence/lr-comparison.json": "724a08155ff3f39a335132cffb0248fdb08f06874e84ce6a582c8af9199ac69d",
|
| 42 |
+
"evidence/replay-check.json": "d362a58603077feddc4ad2513c19f8caea25da1a9ab03eec25c5be5c14c3eec0",
|
| 43 |
+
"evidence/seed-0-history.json": "1090a62e8f1256e9510f8c2c913de14954a5e54e14396a962578ebef83b4a6f3",
|
| 44 |
+
"evidence/seed-1-history.json": "6aee05fe087ec5d970329743a06c1cbff3e2a0e8209cdcb2db41649b66b75432",
|
| 45 |
+
"evidence/seed-2-history.json": "71afe93da2e39ef81417074bc1a4559b9d6f070f677e76f74818129633cf3c33",
|
| 46 |
+
"evidence/seed-3-history.json": "4bf723c1659ffc8adeb57d06d51132f4f5af74c811c98feef71de4f6fa3bd66f",
|
| 47 |
+
"evidence/seed-4-history.json": "2609c807813f55ffe5aef7e6c96aafefe1647de1ac589568d7fdc56e46f5c9fe",
|
| 48 |
+
"export-verification.json": "65e46623219d80979a7108e69c0ed89d0b1d5c23a82c4255fee64647694aac3f",
|
| 49 |
"licenses/jev-benchmarks-LICENSE": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
|
| 50 |
"licenses/laya-LICENSE": "a6cba85bc92e0cff7a450b1d873c0eaa2e9fc96bf472df0247a26bec77bf3ff9",
|
| 51 |
"load_adapter.py": "2f305ef95fef2b2482e9214836fc0071a59cfd6c6d3aaf42cebdbfb131272081",
|
| 52 |
+
"make_training_config.py": "23e9d4acb90141a3d690bc607c861df92137b9457304ae19b8a925398b55536d",
|
| 53 |
+
"metrics.json": "e3cd35993afc309f7ccd99b1a3d96316944e1ab42fd489c7f15465f43ec4a78b",
|
| 54 |
"requirements.txt": "285a8ae6529abeefd058f2e9679e0e314b19fda38169e685ea16aec59b053bd9",
|
| 55 |
+
"training_protocol.json": "00b2db7abbd07ba3220a130fe36ed64958f1dfac15773d183f1ae8dc9d8334b3"
|
| 56 |
}
|
| 57 |
}
|