yusufdxb commited on
Commit
2f68bc9
·
1 Parent(s): 56c6849

v0.2.0 reconciliation: banner, v3 result, shallow-baseline framing, checksums, model-index, provenance

Browse files

- Add 'research benchmark, not for deployment' banner.
- Headline AP 0.9926 placed alongside all-frames LR AP 0.8762.
- Add v3 section: xor_order_color defeats v2 (LR 0.5486, Transformer flatlines AP 0.5068).
- Fix pipeline_tag tabular-classification -> time-series-classification.
- Add last_updated, model-index (verified v2 metrics only).
- Add SHA256 CHECKSUMS for all 10 shipped non-README artifacts.
- Add PROVENANCE.md mapping each model-index value to its source JSON.
- Add license clarification (synthetic data, no third-party dataset license).

Files changed (3) hide show
  1. CHECKSUMS.sha256 +10 -0
  2. PROVENANCE.md +70 -0
  3. README.md +100 -5
CHECKSUMS.sha256 ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ 7fb1523a96a0a7f23c4cf343fc6f5a27d7bd5fa4a844f92304b19afcca4d396e artifact_index.json
2
+ 9e20b26b5cec972be2882f4cb1333bbc18d15335c0199c78ba682113704bfb03 claims_ledger.md
3
+ 8ac6b51e1449f19eac5edf1d4e5ce5d74c0b0f3a259c3ef8629c0d141999bbc3 assets/release_visual.svg
4
+ 29979a6d89071cf05e1b7ffa5ec27ae967430f24ef47b5e8d7601aef54625a93 baselines.json
5
+ 051c5426158bf0957cac905844222f38ef18ea3634ba78b47e7a6e9036adf1f8 baselines.md
6
+ 428c87622bcb6191c4ce2081d0543da0b9a3e4c9425c10168d9e83ba4a2cbb22 benchmark_v2_summary.json
7
+ 85c704597307aa2484bda87fd5215f35cc3217cc05d1dc8a6a978d8f1792a5a9 data_diagnostics.json
8
+ 361f8077142e9704fb6d45edf14c8d5c0b0c35e926a99487d78f6c2722b2ed67 data_diagnostics.md
9
+ 016dd558e2e08fa70cbb9b6f78bdd332e69d2416be78b13dad38affa7aa1574c transformer_v2_35e_panel.json
10
+ a85a02fe164e1493ccddc55d4275951a10422957e6ec62c15587be5a781d1a90 transformer_v2_minrecall080_fpr020.json
PROVENANCE.md ADDED
@@ -0,0 +1,70 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PROVENANCE — UI-PreFail-v2 reported metrics
2
+
3
+ Every numeric value in the `model-index` block of `README.md` traces to a JSON
4
+ file shipped in this repo or in the local PreFailureNet repo. The mapping is
5
+ recorded here rather than inline in `model-index` so the YAML stays
6
+ schema-compliant for Hugging Face's parser.
7
+
8
+ Tag: `v0.2.0-reconciliation` (2026-05-25).
9
+
10
+ ## v2 Transformer (held-out test, n=256, 128 failure / 128 stable)
11
+
12
+ Source file: `transformer_v2_35e_panel.json` (shipped in this repo)
13
+
14
+ | Metric name on card | Field in JSON | Raw value |
15
+ | --- | --- | --- |
16
+ | AP (failure vs stable) | `ap_failure_vs_stable` | `0.9926430180293419` |
17
+ | Macro F1 (used classes) | `macro_f1_used_only` | `0.9910648756727424` |
18
+ | Pair-rank accuracy | `pair_rank_acc` | `1.0` |
19
+ | Pair intervention success | `pair_intervention_success` | `0.984375` |
20
+ | FPR at predicted class | `fpr_at_predicted_class` | `0.015625` |
21
+ | Recall at predicted class | `recall_at_predicted_class` | `1.0` |
22
+
23
+ ## Gate calibration (risk_only, min_recall=0.80, max_FPR=0.20)
24
+
25
+ Source file: `transformer_v2_minrecall080_fpr020.json` (shipped in this repo)
26
+
27
+ | Metric name on card | Field in JSON | Raw value |
28
+ | --- | --- | --- |
29
+ | Gate test recall | `test_metrics.recall` | `1.0` |
30
+ | Gate test FPR (false intervention rate) | `test_metrics.false_intervention_rate` | `0.0078125` |
31
+ | Gate test precision | `test_metrics.precision` | `0.9922480620155039` |
32
+
33
+ For reference (validation split, not on card):
34
+
35
+ | Validation metric | Field | Raw value |
36
+ | --- | --- | --- |
37
+ | Recall | `validation_metrics.recall` | `0.9921875` |
38
+ | FPR | `validation_metrics.false_intervention_rate` | `0.0` |
39
+ | Precision | `validation_metrics.precision` | `1.0` |
40
+ | Threshold | `validation_metrics.threshold` | `0.9435477256774902` |
41
+
42
+ ## Shallow baseline (all-frames LR) on v2
43
+
44
+ Source file: `baselines.json` (shipped in this repo), entry `all_frames`.
45
+
46
+ | Field | Raw value |
47
+ | --- | --- |
48
+ | `ap_failure_vs_stable` | `0.8761761121534359` |
49
+ | `pair_rank_acc` | `0.8828125` |
50
+ | `pair_intervention_success` | `0.6015625` |
51
+
52
+ ## v3 results (not shipped in this repo; cited in README v3 section)
53
+
54
+ Sources in the local PreFailureNet working tree at
55
+ `~/Projects/PreFailureNet/runs/benchmark_v3/`:
56
+
57
+ | Card claim | Source file | Field | Raw value |
58
+ | --- | --- | --- | --- |
59
+ | v3 all-frames LR AP (rounded 0.5486 / 0.549 in docs) | `baselines_seed2920/baselines.json` (entry `all_frames`) | `ap_failure_vs_stable` | regenerable from local repo |
60
+ | v3 Transformer precursor_detection_ap | `transformer_v3/metrics.json` | `precursor_detection_ap` | `0.5068114175400058` |
61
+ | v3 Transformer intervention_auroc | `transformer_v3/metrics.json` | `intervention_auroc` | `0.500152587890625` |
62
+ | v3 generator commit / implementation | local commit `3a323b1`, `src/prefailurenet/data/synthetic.py`, function `_temporal_order_v3_pair` | n/a | n/a |
63
+ | v3 cross-reference doc | `docs/v3_split_evaluation.md` (rounds AP to `0.549`) | n/a | n/a |
64
+
65
+ ## Phase 1 audit reference
66
+
67
+ Full ground-truth audit: `/tmp/hf_audit/AUDIT.md` (claim-by-claim verification
68
+ against shipped JSONs). Phase 1 verified the numbers above as artifact-match,
69
+ not as re-execution; no metric was re-trained or re-fit during the
70
+ reconciliation.
README.md CHANGED
@@ -2,7 +2,8 @@
2
  language: en
3
  license: apache-2.0
4
  library_name: pytorch
5
- pipeline_tag: tabular-classification
 
6
  tags:
7
  - benchmark
8
  - temporal-modeling
@@ -13,19 +14,88 @@ tags:
13
  - evaluation
14
  - browser-agents
15
  - ai-agents
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
16
  ---
17
 
 
 
18
  # PreFailureNet / UI-PreFail-v2
19
 
20
  ![PreFailureNet UI-PreFail-v2 release visual](assets/release_visual.svg)
21
 
22
  PreFailureNet / UI-PreFail-v2 is a synthetic ordered temporal counterfactual benchmark for pre-failure detection in browser-agent-like workflows. It includes shortcut probes, baseline ladders, a reference Transformer, and calibrated gate diagnostics. It is not a production browser-agent guardrail.
23
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
24
  ## Exact Positioning
25
 
26
  This package is a benchmark and research artifact. The correct claim is:
27
 
28
- > Ordered temporal evidence is required under synthetic counterfactual controls.
29
 
30
  Do not read this release as evidence of production browser-agent safety, real-world website generalization, or a deployable intervention gate.
31
 
@@ -39,6 +109,7 @@ Do not read this release as evidence of production browser-agent safety, real-wo
39
  - Reference Transformer v2 held-out test metrics.
40
  - Risk-only gate calibration diagnostics.
41
  - Reproduction commands.
 
42
 
43
  Raw NPZ arrays and model checkpoint weights are not included in this compact HF package. They are reproducible from the commands below.
44
 
@@ -82,7 +153,7 @@ UI-PreFail-v2 / `temporal_order_v2` changes the label rule:
82
  - Safe and unsafe samples have the same unordered frame bag.
83
  - The label depends on the order of two event-bearing frames.
84
 
85
- This makes static, final-frame, non-vision, single mid-frame, and unordered-bag probes weak by construction. Ordered frame sequences and temporal deltas still contain useful signal.
86
 
87
  ## Held-Out v2 Pair Controls
88
 
@@ -197,6 +268,24 @@ PYTHONPATH=src python3 scripts/check_hf_package.py
197
  | `data_diagnostics.md` | Human-readable diagnostics. |
198
  | `transformer_v2_35e_panel.json` | Reference Transformer held-out test metrics. |
199
  | `transformer_v2_minrecall080_fpr020.json` | Gate calibration diagnostics. |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
200
 
201
  ## Verified
202
 
@@ -207,6 +296,7 @@ PYTHONPATH=src python3 scripts/check_hf_package.py
207
  - Transformer v2 result exists.
208
  - Gate calibration result exists.
209
  - v1 flaw is documented.
 
210
 
211
  ## Unverified / Not Claimed
212
 
@@ -222,14 +312,19 @@ PYTHONPATH=src python3 scripts/check_hf_package.py
222
  - The benchmark is synthetic.
223
  - The UI events are rendered primitives, not production browser traces.
224
  - The reference result is on generated data only.
225
- - A flattened all-frames LR is strong because it sees ordered frames; v2 tests ordered temporal evidence, not uniquely deep temporal reasoning.
226
  - Gate calibration is a benchmark diagnostic, not a safety policy.
227
  - The class schema contains unused classes in this split.
228
 
 
 
 
 
229
  ## Next Research Steps
230
 
231
  - Add external browser-agent traces with consent and clear provenance.
232
- - Add stronger adversarial temporal controls where full-frame linear probes degrade.
233
  - Add held-out UI event grammars and templates.
234
  - Evaluate under real browser screenshots and DOM/action logs.
235
  - Separate benchmark scoring from intervention-policy deployment work.
 
 
2
  language: en
3
  license: apache-2.0
4
  library_name: pytorch
5
+ pipeline_tag: time-series-classification
6
+ last_updated: 2026-05-25
7
  tags:
8
  - benchmark
9
  - temporal-modeling
 
14
  - evaluation
15
  - browser-agents
16
  - ai-agents
17
+ model-index:
18
+ - name: PreFailureNet UI-PreFail-v2 (Transformer, 35 epochs)
19
+ results:
20
+ - task:
21
+ type: time-series-classification
22
+ name: Pre-failure detection (synthetic ordered counterfactual, v2)
23
+ dataset:
24
+ type: synthetic
25
+ name: UI-PreFail-v2 temporal_order_v2 (seed 2910 test split, n=256)
26
+ metrics:
27
+ - type: average_precision
28
+ name: AP (failure vs stable)
29
+ value: 0.9926
30
+ - type: f1
31
+ name: Macro F1 (used classes)
32
+ value: 0.9911
33
+ - type: accuracy
34
+ name: Pair-rank accuracy
35
+ value: 1.0
36
+ - type: accuracy
37
+ name: Pair intervention success
38
+ value: 0.9844
39
+ - type: false_positive_rate
40
+ name: FPR at predicted class
41
+ value: 0.0156
42
+ - type: recall
43
+ name: Recall at predicted class
44
+ value: 1.0
45
+ - type: recall
46
+ name: Gate test recall
47
+ value: 1.0
48
+ - type: false_positive_rate
49
+ name: Gate test FPR (false intervention rate)
50
+ value: 0.0078
51
+ - type: precision
52
+ name: Gate test precision
53
+ value: 0.9922
54
  ---
55
 
56
+ > **RESEARCH BENCHMARK, NOT FOR DEPLOYMENT.** This package is a synthetic counterfactual benchmark for studying ordered-temporal evidence in pre-failure detection. It is not a browser-agent guardrail and has no real-world validation.
57
+
58
  # PreFailureNet / UI-PreFail-v2
59
 
60
  ![PreFailureNet UI-PreFail-v2 release visual](assets/release_visual.svg)
61
 
62
  PreFailureNet / UI-PreFail-v2 is a synthetic ordered temporal counterfactual benchmark for pre-failure detection in browser-agent-like workflows. It includes shortcut probes, baseline ladders, a reference Transformer, and calibrated gate diagnostics. It is not a production browser-agent guardrail.
63
 
64
+ ## Headline Result vs Shallow Baseline
65
+
66
+ Sources: `transformer_v2_35e_panel.json`, `baselines.json` (entry `all_frames`).
67
+
68
+ | Method | AP (failure vs stable) | n_test |
69
+ | --- | ---: | ---: |
70
+ | Transformer v2, 35 epochs | `0.9926` | `256` |
71
+ | All-frames logistic regression (flattened frames) | `0.8762` | `256` |
72
+
73
+ A shallow logistic regression over flattened frames already reaches AP 0.8762 on this split, so the v2 headline reflects a benchmark that does not require deep temporal modeling. See v3 below for a split where this gap closes.
74
+
75
+ ## v3: xor_order_color defeats v2
76
+
77
+ v3 introduces the `xor_order_color` family (local commit `3a323b1`, file `src/prefailurenet/data/synthetic.py`, function `_temporal_order_v3_pair`). The label is the XOR of an event-order bit and a per-pair color bit painted identically into every frame, so ordered-frame linear probes lose their handle while the pair structure stays intact.
78
+
79
+ On v3 (test seed 2920, n_test=256, 128/128 pairs), the all-frames ordered LR drops from AP `0.8762` on v2 to AP `0.5486` on v3 (rounded to `0.549` in `docs/v3_split_evaluation.md`), which is below the project's 0.60 hard gate and within ~0.05 of chance. Source: `~/Projects/PreFailureNet/runs/benchmark_v3/baselines_seed2920/baselines.json` (entry `all_frames`, field `ap_failure_vs_stable`), cross-referenced in `docs/v3_split_evaluation.md`.
80
+
81
+ The v2 Transformer recipe retrained on v3 flatlines:
82
+
83
+ | Field | Value |
84
+ | --- | ---: |
85
+ | `precursor_detection_ap` | `0.5068` |
86
+ | `intervention_auroc` | `0.5002` |
87
+
88
+ Source: `~/Projects/PreFailureNet/runs/benchmark_v3/transformer_v3/metrics.json` (raw values `0.5068114175400058` and `0.500152587890625`).
89
+
90
+ v3 inverts the v2 narrative. v2 results should be read as a defeated benchmark, not as evidence that the v2 Transformer architecture learns temporal order.
91
+
92
+ v3 will be published as a separate HF repo or addendum after a v3-targeted model exists. Currently no model beats v3 baseline.
93
+
94
  ## Exact Positioning
95
 
96
  This package is a benchmark and research artifact. The correct claim is:
97
 
98
+ > Ordered temporal evidence is required under synthetic counterfactual controls (v2). v3 shows this benchmark is itself defeated by an XOR-color counterfactual; the v2 headline is therefore a benchmark, not a capability claim.
99
 
100
  Do not read this release as evidence of production browser-agent safety, real-world website generalization, or a deployable intervention gate.
101
 
 
109
  - Reference Transformer v2 held-out test metrics.
110
  - Risk-only gate calibration diagnostics.
111
  - Reproduction commands.
112
+ - SHA256 checksums for all shipped artifacts (`CHECKSUMS.sha256`).
113
 
114
  Raw NPZ arrays and model checkpoint weights are not included in this compact HF package. They are reproducible from the commands below.
115
 
 
153
  - Safe and unsafe samples have the same unordered frame bag.
154
  - The label depends on the order of two event-bearing frames.
155
 
156
+ This makes static, final-frame, non-vision, single mid-frame, and unordered-bag probes weak by construction. Ordered frame sequences and temporal deltas still contain useful signal. (v3 closes the remaining handle. See above.)
157
 
158
  ## Held-Out v2 Pair Controls
159
 
 
268
  | `data_diagnostics.md` | Human-readable diagnostics. |
269
  | `transformer_v2_35e_panel.json` | Reference Transformer held-out test metrics. |
270
  | `transformer_v2_minrecall080_fpr020.json` | Gate calibration diagnostics. |
271
+ | `CHECKSUMS.sha256` | SHA256 checksums for all shipped non-README artifacts. |
272
+
273
+ ## Checksums
274
+
275
+ Verify with `sha256sum -c CHECKSUMS.sha256` from the repo root.
276
+
277
+ | File | SHA256 |
278
+ | --- | --- |
279
+ | `artifact_index.json` | `7fb1523a96a0a7f23c4cf343fc6f5a27d7bd5fa4a844f92304b19afcca4d396e` |
280
+ | `claims_ledger.md` | `9e20b26b5cec972be2882f4cb1333bbc18d15335c0199c78ba682113704bfb03` |
281
+ | `assets/release_visual.svg` | `8ac6b51e1449f19eac5edf1d4e5ce5d74c0b0f3a259c3ef8629c0d141999bbc3` |
282
+ | `baselines.json` | `29979a6d89071cf05e1b7ffa5ec27ae967430f24ef47b5e8d7601aef54625a93` |
283
+ | `baselines.md` | `051c5426158bf0957cac905844222f38ef18ea3634ba78b47e7a6e9036adf1f8` |
284
+ | `benchmark_v2_summary.json` | `428c87622bcb6191c4ce2081d0543da0b9a3e4c9425c10168d9e83ba4a2cbb22` |
285
+ | `data_diagnostics.json` | `85c704597307aa2484bda87fd5215f35cc3217cc05d1dc8a6a978d8f1792a5a9` |
286
+ | `data_diagnostics.md` | `361f8077142e9704fb6d45edf14c8d5c0b0c35e926a99487d78f6c2722b2ed67` |
287
+ | `transformer_v2_35e_panel.json` | `016dd558e2e08fa70cbb9b6f78bdd332e69d2416be78b13dad38affa7aa1574c` |
288
+ | `transformer_v2_minrecall080_fpr020.json` | `a85a02fe164e1493ccddc55d4275951a10422957e6ec62c15587be5a781d1a90` |
289
 
290
  ## Verified
291
 
 
296
  - Transformer v2 result exists.
297
  - Gate calibration result exists.
298
  - v1 flaw is documented.
299
+ - v3 XOR counterfactual split exists locally and defeats both the ordered-LR baseline and the v2 Transformer recipe (see v3 section above).
300
 
301
  ## Unverified / Not Claimed
302
 
 
312
  - The benchmark is synthetic.
313
  - The UI events are rendered primitives, not production browser traces.
314
  - The reference result is on generated data only.
315
+ - A flattened all-frames LR is strong because it sees ordered frames; v2 tests ordered temporal evidence, not uniquely deep temporal reasoning. The v3 XOR split removes even this handle and currently has no model that beats baseline.
316
  - Gate calibration is a benchmark diagnostic, not a safety policy.
317
  - The class schema contains unused classes in this split.
318
 
319
+ ## License
320
+
321
+ Apache-2.0 for the package code, artifacts, and model card. Training and test data are synthetic and generated by the reproduction commands in this repo. No third-party dataset license applies.
322
+
323
  ## Next Research Steps
324
 
325
  - Add external browser-agent traces with consent and clear provenance.
326
+ - Add stronger adversarial temporal controls where full-frame linear probes degrade (v3 is the first such step).
327
  - Add held-out UI event grammars and templates.
328
  - Evaluate under real browser screenshots and DOM/action logs.
329
  - Separate benchmark scoring from intervention-policy deployment work.
330
+ - Train a v3-targeted model that actually beats v3 baseline before publishing a v3 HF repo.