PyTorch
Transformers
English
benchmark
temporal-modeling
sequence-learning
synthetic-data
counterfactuals
evaluation
browser-agents
ai-agents
time-series-classification
Eval Results (legacy)
Instructions to use yusufdxb/prefailurenet-ui-prefail-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yusufdxb/prefailurenet-ui-prefail-v2 with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("yusufdxb/prefailurenet-ui-prefail-v2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
v0.2.0 reconciliation: banner, v3 result, shallow-baseline framing, checksums, model-index, provenance
Browse files- Add 'research benchmark, not for deployment' banner.
- Headline AP 0.9926 placed alongside all-frames LR AP 0.8762.
- Add v3 section: xor_order_color defeats v2 (LR 0.5486, Transformer flatlines AP 0.5068).
- Fix pipeline_tag tabular-classification -> time-series-classification.
- Add last_updated, model-index (verified v2 metrics only).
- Add SHA256 CHECKSUMS for all 10 shipped non-README artifacts.
- Add PROVENANCE.md mapping each model-index value to its source JSON.
- Add license clarification (synthetic data, no third-party dataset license).
- CHECKSUMS.sha256 +10 -0
- PROVENANCE.md +70 -0
- README.md +100 -5
CHECKSUMS.sha256
ADDED
|
@@ -0,0 +1,10 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
7fb1523a96a0a7f23c4cf343fc6f5a27d7bd5fa4a844f92304b19afcca4d396e artifact_index.json
|
| 2 |
+
9e20b26b5cec972be2882f4cb1333bbc18d15335c0199c78ba682113704bfb03 claims_ledger.md
|
| 3 |
+
8ac6b51e1449f19eac5edf1d4e5ce5d74c0b0f3a259c3ef8629c0d141999bbc3 assets/release_visual.svg
|
| 4 |
+
29979a6d89071cf05e1b7ffa5ec27ae967430f24ef47b5e8d7601aef54625a93 baselines.json
|
| 5 |
+
051c5426158bf0957cac905844222f38ef18ea3634ba78b47e7a6e9036adf1f8 baselines.md
|
| 6 |
+
428c87622bcb6191c4ce2081d0543da0b9a3e4c9425c10168d9e83ba4a2cbb22 benchmark_v2_summary.json
|
| 7 |
+
85c704597307aa2484bda87fd5215f35cc3217cc05d1dc8a6a978d8f1792a5a9 data_diagnostics.json
|
| 8 |
+
361f8077142e9704fb6d45edf14c8d5c0b0c35e926a99487d78f6c2722b2ed67 data_diagnostics.md
|
| 9 |
+
016dd558e2e08fa70cbb9b6f78bdd332e69d2416be78b13dad38affa7aa1574c transformer_v2_35e_panel.json
|
| 10 |
+
a85a02fe164e1493ccddc55d4275951a10422957e6ec62c15587be5a781d1a90 transformer_v2_minrecall080_fpr020.json
|
PROVENANCE.md
ADDED
|
@@ -0,0 +1,70 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# PROVENANCE — UI-PreFail-v2 reported metrics
|
| 2 |
+
|
| 3 |
+
Every numeric value in the `model-index` block of `README.md` traces to a JSON
|
| 4 |
+
file shipped in this repo or in the local PreFailureNet repo. The mapping is
|
| 5 |
+
recorded here rather than inline in `model-index` so the YAML stays
|
| 6 |
+
schema-compliant for Hugging Face's parser.
|
| 7 |
+
|
| 8 |
+
Tag: `v0.2.0-reconciliation` (2026-05-25).
|
| 9 |
+
|
| 10 |
+
## v2 Transformer (held-out test, n=256, 128 failure / 128 stable)
|
| 11 |
+
|
| 12 |
+
Source file: `transformer_v2_35e_panel.json` (shipped in this repo)
|
| 13 |
+
|
| 14 |
+
| Metric name on card | Field in JSON | Raw value |
|
| 15 |
+
| --- | --- | --- |
|
| 16 |
+
| AP (failure vs stable) | `ap_failure_vs_stable` | `0.9926430180293419` |
|
| 17 |
+
| Macro F1 (used classes) | `macro_f1_used_only` | `0.9910648756727424` |
|
| 18 |
+
| Pair-rank accuracy | `pair_rank_acc` | `1.0` |
|
| 19 |
+
| Pair intervention success | `pair_intervention_success` | `0.984375` |
|
| 20 |
+
| FPR at predicted class | `fpr_at_predicted_class` | `0.015625` |
|
| 21 |
+
| Recall at predicted class | `recall_at_predicted_class` | `1.0` |
|
| 22 |
+
|
| 23 |
+
## Gate calibration (risk_only, min_recall=0.80, max_FPR=0.20)
|
| 24 |
+
|
| 25 |
+
Source file: `transformer_v2_minrecall080_fpr020.json` (shipped in this repo)
|
| 26 |
+
|
| 27 |
+
| Metric name on card | Field in JSON | Raw value |
|
| 28 |
+
| --- | --- | --- |
|
| 29 |
+
| Gate test recall | `test_metrics.recall` | `1.0` |
|
| 30 |
+
| Gate test FPR (false intervention rate) | `test_metrics.false_intervention_rate` | `0.0078125` |
|
| 31 |
+
| Gate test precision | `test_metrics.precision` | `0.9922480620155039` |
|
| 32 |
+
|
| 33 |
+
For reference (validation split, not on card):
|
| 34 |
+
|
| 35 |
+
| Validation metric | Field | Raw value |
|
| 36 |
+
| --- | --- | --- |
|
| 37 |
+
| Recall | `validation_metrics.recall` | `0.9921875` |
|
| 38 |
+
| FPR | `validation_metrics.false_intervention_rate` | `0.0` |
|
| 39 |
+
| Precision | `validation_metrics.precision` | `1.0` |
|
| 40 |
+
| Threshold | `validation_metrics.threshold` | `0.9435477256774902` |
|
| 41 |
+
|
| 42 |
+
## Shallow baseline (all-frames LR) on v2
|
| 43 |
+
|
| 44 |
+
Source file: `baselines.json` (shipped in this repo), entry `all_frames`.
|
| 45 |
+
|
| 46 |
+
| Field | Raw value |
|
| 47 |
+
| --- | --- |
|
| 48 |
+
| `ap_failure_vs_stable` | `0.8761761121534359` |
|
| 49 |
+
| `pair_rank_acc` | `0.8828125` |
|
| 50 |
+
| `pair_intervention_success` | `0.6015625` |
|
| 51 |
+
|
| 52 |
+
## v3 results (not shipped in this repo; cited in README v3 section)
|
| 53 |
+
|
| 54 |
+
Sources in the local PreFailureNet working tree at
|
| 55 |
+
`~/Projects/PreFailureNet/runs/benchmark_v3/`:
|
| 56 |
+
|
| 57 |
+
| Card claim | Source file | Field | Raw value |
|
| 58 |
+
| --- | --- | --- | --- |
|
| 59 |
+
| v3 all-frames LR AP (rounded 0.5486 / 0.549 in docs) | `baselines_seed2920/baselines.json` (entry `all_frames`) | `ap_failure_vs_stable` | regenerable from local repo |
|
| 60 |
+
| v3 Transformer precursor_detection_ap | `transformer_v3/metrics.json` | `precursor_detection_ap` | `0.5068114175400058` |
|
| 61 |
+
| v3 Transformer intervention_auroc | `transformer_v3/metrics.json` | `intervention_auroc` | `0.500152587890625` |
|
| 62 |
+
| v3 generator commit / implementation | local commit `3a323b1`, `src/prefailurenet/data/synthetic.py`, function `_temporal_order_v3_pair` | n/a | n/a |
|
| 63 |
+
| v3 cross-reference doc | `docs/v3_split_evaluation.md` (rounds AP to `0.549`) | n/a | n/a |
|
| 64 |
+
|
| 65 |
+
## Phase 1 audit reference
|
| 66 |
+
|
| 67 |
+
Full ground-truth audit: `/tmp/hf_audit/AUDIT.md` (claim-by-claim verification
|
| 68 |
+
against shipped JSONs). Phase 1 verified the numbers above as artifact-match,
|
| 69 |
+
not as re-execution; no metric was re-trained or re-fit during the
|
| 70 |
+
reconciliation.
|
README.md
CHANGED
|
@@ -2,7 +2,8 @@
|
|
| 2 |
language: en
|
| 3 |
license: apache-2.0
|
| 4 |
library_name: pytorch
|
| 5 |
-
pipeline_tag:
|
|
|
|
| 6 |
tags:
|
| 7 |
- benchmark
|
| 8 |
- temporal-modeling
|
|
@@ -13,19 +14,88 @@ tags:
|
|
| 13 |
- evaluation
|
| 14 |
- browser-agents
|
| 15 |
- ai-agents
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
---
|
| 17 |
|
|
|
|
|
|
|
| 18 |
# PreFailureNet / UI-PreFail-v2
|
| 19 |
|
| 20 |

|
| 21 |
|
| 22 |
PreFailureNet / UI-PreFail-v2 is a synthetic ordered temporal counterfactual benchmark for pre-failure detection in browser-agent-like workflows. It includes shortcut probes, baseline ladders, a reference Transformer, and calibrated gate diagnostics. It is not a production browser-agent guardrail.
|
| 23 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 24 |
## Exact Positioning
|
| 25 |
|
| 26 |
This package is a benchmark and research artifact. The correct claim is:
|
| 27 |
|
| 28 |
-
> Ordered temporal evidence is required under synthetic counterfactual controls.
|
| 29 |
|
| 30 |
Do not read this release as evidence of production browser-agent safety, real-world website generalization, or a deployable intervention gate.
|
| 31 |
|
|
@@ -39,6 +109,7 @@ Do not read this release as evidence of production browser-agent safety, real-wo
|
|
| 39 |
- Reference Transformer v2 held-out test metrics.
|
| 40 |
- Risk-only gate calibration diagnostics.
|
| 41 |
- Reproduction commands.
|
|
|
|
| 42 |
|
| 43 |
Raw NPZ arrays and model checkpoint weights are not included in this compact HF package. They are reproducible from the commands below.
|
| 44 |
|
|
@@ -82,7 +153,7 @@ UI-PreFail-v2 / `temporal_order_v2` changes the label rule:
|
|
| 82 |
- Safe and unsafe samples have the same unordered frame bag.
|
| 83 |
- The label depends on the order of two event-bearing frames.
|
| 84 |
|
| 85 |
-
This makes static, final-frame, non-vision, single mid-frame, and unordered-bag probes weak by construction. Ordered frame sequences and temporal deltas still contain useful signal.
|
| 86 |
|
| 87 |
## Held-Out v2 Pair Controls
|
| 88 |
|
|
@@ -197,6 +268,24 @@ PYTHONPATH=src python3 scripts/check_hf_package.py
|
|
| 197 |
| `data_diagnostics.md` | Human-readable diagnostics. |
|
| 198 |
| `transformer_v2_35e_panel.json` | Reference Transformer held-out test metrics. |
|
| 199 |
| `transformer_v2_minrecall080_fpr020.json` | Gate calibration diagnostics. |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 200 |
|
| 201 |
## Verified
|
| 202 |
|
|
@@ -207,6 +296,7 @@ PYTHONPATH=src python3 scripts/check_hf_package.py
|
|
| 207 |
- Transformer v2 result exists.
|
| 208 |
- Gate calibration result exists.
|
| 209 |
- v1 flaw is documented.
|
|
|
|
| 210 |
|
| 211 |
## Unverified / Not Claimed
|
| 212 |
|
|
@@ -222,14 +312,19 @@ PYTHONPATH=src python3 scripts/check_hf_package.py
|
|
| 222 |
- The benchmark is synthetic.
|
| 223 |
- The UI events are rendered primitives, not production browser traces.
|
| 224 |
- The reference result is on generated data only.
|
| 225 |
-
- A flattened all-frames LR is strong because it sees ordered frames; v2 tests ordered temporal evidence, not uniquely deep temporal reasoning.
|
| 226 |
- Gate calibration is a benchmark diagnostic, not a safety policy.
|
| 227 |
- The class schema contains unused classes in this split.
|
| 228 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 229 |
## Next Research Steps
|
| 230 |
|
| 231 |
- Add external browser-agent traces with consent and clear provenance.
|
| 232 |
-
- Add stronger adversarial temporal controls where full-frame linear probes degrade.
|
| 233 |
- Add held-out UI event grammars and templates.
|
| 234 |
- Evaluate under real browser screenshots and DOM/action logs.
|
| 235 |
- Separate benchmark scoring from intervention-policy deployment work.
|
|
|
|
|
|
| 2 |
language: en
|
| 3 |
license: apache-2.0
|
| 4 |
library_name: pytorch
|
| 5 |
+
pipeline_tag: time-series-classification
|
| 6 |
+
last_updated: 2026-05-25
|
| 7 |
tags:
|
| 8 |
- benchmark
|
| 9 |
- temporal-modeling
|
|
|
|
| 14 |
- evaluation
|
| 15 |
- browser-agents
|
| 16 |
- ai-agents
|
| 17 |
+
model-index:
|
| 18 |
+
- name: PreFailureNet UI-PreFail-v2 (Transformer, 35 epochs)
|
| 19 |
+
results:
|
| 20 |
+
- task:
|
| 21 |
+
type: time-series-classification
|
| 22 |
+
name: Pre-failure detection (synthetic ordered counterfactual, v2)
|
| 23 |
+
dataset:
|
| 24 |
+
type: synthetic
|
| 25 |
+
name: UI-PreFail-v2 temporal_order_v2 (seed 2910 test split, n=256)
|
| 26 |
+
metrics:
|
| 27 |
+
- type: average_precision
|
| 28 |
+
name: AP (failure vs stable)
|
| 29 |
+
value: 0.9926
|
| 30 |
+
- type: f1
|
| 31 |
+
name: Macro F1 (used classes)
|
| 32 |
+
value: 0.9911
|
| 33 |
+
- type: accuracy
|
| 34 |
+
name: Pair-rank accuracy
|
| 35 |
+
value: 1.0
|
| 36 |
+
- type: accuracy
|
| 37 |
+
name: Pair intervention success
|
| 38 |
+
value: 0.9844
|
| 39 |
+
- type: false_positive_rate
|
| 40 |
+
name: FPR at predicted class
|
| 41 |
+
value: 0.0156
|
| 42 |
+
- type: recall
|
| 43 |
+
name: Recall at predicted class
|
| 44 |
+
value: 1.0
|
| 45 |
+
- type: recall
|
| 46 |
+
name: Gate test recall
|
| 47 |
+
value: 1.0
|
| 48 |
+
- type: false_positive_rate
|
| 49 |
+
name: Gate test FPR (false intervention rate)
|
| 50 |
+
value: 0.0078
|
| 51 |
+
- type: precision
|
| 52 |
+
name: Gate test precision
|
| 53 |
+
value: 0.9922
|
| 54 |
---
|
| 55 |
|
| 56 |
+
> **RESEARCH BENCHMARK, NOT FOR DEPLOYMENT.** This package is a synthetic counterfactual benchmark for studying ordered-temporal evidence in pre-failure detection. It is not a browser-agent guardrail and has no real-world validation.
|
| 57 |
+
|
| 58 |
# PreFailureNet / UI-PreFail-v2
|
| 59 |
|
| 60 |

|
| 61 |
|
| 62 |
PreFailureNet / UI-PreFail-v2 is a synthetic ordered temporal counterfactual benchmark for pre-failure detection in browser-agent-like workflows. It includes shortcut probes, baseline ladders, a reference Transformer, and calibrated gate diagnostics. It is not a production browser-agent guardrail.
|
| 63 |
|
| 64 |
+
## Headline Result vs Shallow Baseline
|
| 65 |
+
|
| 66 |
+
Sources: `transformer_v2_35e_panel.json`, `baselines.json` (entry `all_frames`).
|
| 67 |
+
|
| 68 |
+
| Method | AP (failure vs stable) | n_test |
|
| 69 |
+
| --- | ---: | ---: |
|
| 70 |
+
| Transformer v2, 35 epochs | `0.9926` | `256` |
|
| 71 |
+
| All-frames logistic regression (flattened frames) | `0.8762` | `256` |
|
| 72 |
+
|
| 73 |
+
A shallow logistic regression over flattened frames already reaches AP 0.8762 on this split, so the v2 headline reflects a benchmark that does not require deep temporal modeling. See v3 below for a split where this gap closes.
|
| 74 |
+
|
| 75 |
+
## v3: xor_order_color defeats v2
|
| 76 |
+
|
| 77 |
+
v3 introduces the `xor_order_color` family (local commit `3a323b1`, file `src/prefailurenet/data/synthetic.py`, function `_temporal_order_v3_pair`). The label is the XOR of an event-order bit and a per-pair color bit painted identically into every frame, so ordered-frame linear probes lose their handle while the pair structure stays intact.
|
| 78 |
+
|
| 79 |
+
On v3 (test seed 2920, n_test=256, 128/128 pairs), the all-frames ordered LR drops from AP `0.8762` on v2 to AP `0.5486` on v3 (rounded to `0.549` in `docs/v3_split_evaluation.md`), which is below the project's 0.60 hard gate and within ~0.05 of chance. Source: `~/Projects/PreFailureNet/runs/benchmark_v3/baselines_seed2920/baselines.json` (entry `all_frames`, field `ap_failure_vs_stable`), cross-referenced in `docs/v3_split_evaluation.md`.
|
| 80 |
+
|
| 81 |
+
The v2 Transformer recipe retrained on v3 flatlines:
|
| 82 |
+
|
| 83 |
+
| Field | Value |
|
| 84 |
+
| --- | ---: |
|
| 85 |
+
| `precursor_detection_ap` | `0.5068` |
|
| 86 |
+
| `intervention_auroc` | `0.5002` |
|
| 87 |
+
|
| 88 |
+
Source: `~/Projects/PreFailureNet/runs/benchmark_v3/transformer_v3/metrics.json` (raw values `0.5068114175400058` and `0.500152587890625`).
|
| 89 |
+
|
| 90 |
+
v3 inverts the v2 narrative. v2 results should be read as a defeated benchmark, not as evidence that the v2 Transformer architecture learns temporal order.
|
| 91 |
+
|
| 92 |
+
v3 will be published as a separate HF repo or addendum after a v3-targeted model exists. Currently no model beats v3 baseline.
|
| 93 |
+
|
| 94 |
## Exact Positioning
|
| 95 |
|
| 96 |
This package is a benchmark and research artifact. The correct claim is:
|
| 97 |
|
| 98 |
+
> Ordered temporal evidence is required under synthetic counterfactual controls (v2). v3 shows this benchmark is itself defeated by an XOR-color counterfactual; the v2 headline is therefore a benchmark, not a capability claim.
|
| 99 |
|
| 100 |
Do not read this release as evidence of production browser-agent safety, real-world website generalization, or a deployable intervention gate.
|
| 101 |
|
|
|
|
| 109 |
- Reference Transformer v2 held-out test metrics.
|
| 110 |
- Risk-only gate calibration diagnostics.
|
| 111 |
- Reproduction commands.
|
| 112 |
+
- SHA256 checksums for all shipped artifacts (`CHECKSUMS.sha256`).
|
| 113 |
|
| 114 |
Raw NPZ arrays and model checkpoint weights are not included in this compact HF package. They are reproducible from the commands below.
|
| 115 |
|
|
|
|
| 153 |
- Safe and unsafe samples have the same unordered frame bag.
|
| 154 |
- The label depends on the order of two event-bearing frames.
|
| 155 |
|
| 156 |
+
This makes static, final-frame, non-vision, single mid-frame, and unordered-bag probes weak by construction. Ordered frame sequences and temporal deltas still contain useful signal. (v3 closes the remaining handle. See above.)
|
| 157 |
|
| 158 |
## Held-Out v2 Pair Controls
|
| 159 |
|
|
|
|
| 268 |
| `data_diagnostics.md` | Human-readable diagnostics. |
|
| 269 |
| `transformer_v2_35e_panel.json` | Reference Transformer held-out test metrics. |
|
| 270 |
| `transformer_v2_minrecall080_fpr020.json` | Gate calibration diagnostics. |
|
| 271 |
+
| `CHECKSUMS.sha256` | SHA256 checksums for all shipped non-README artifacts. |
|
| 272 |
+
|
| 273 |
+
## Checksums
|
| 274 |
+
|
| 275 |
+
Verify with `sha256sum -c CHECKSUMS.sha256` from the repo root.
|
| 276 |
+
|
| 277 |
+
| File | SHA256 |
|
| 278 |
+
| --- | --- |
|
| 279 |
+
| `artifact_index.json` | `7fb1523a96a0a7f23c4cf343fc6f5a27d7bd5fa4a844f92304b19afcca4d396e` |
|
| 280 |
+
| `claims_ledger.md` | `9e20b26b5cec972be2882f4cb1333bbc18d15335c0199c78ba682113704bfb03` |
|
| 281 |
+
| `assets/release_visual.svg` | `8ac6b51e1449f19eac5edf1d4e5ce5d74c0b0f3a259c3ef8629c0d141999bbc3` |
|
| 282 |
+
| `baselines.json` | `29979a6d89071cf05e1b7ffa5ec27ae967430f24ef47b5e8d7601aef54625a93` |
|
| 283 |
+
| `baselines.md` | `051c5426158bf0957cac905844222f38ef18ea3634ba78b47e7a6e9036adf1f8` |
|
| 284 |
+
| `benchmark_v2_summary.json` | `428c87622bcb6191c4ce2081d0543da0b9a3e4c9425c10168d9e83ba4a2cbb22` |
|
| 285 |
+
| `data_diagnostics.json` | `85c704597307aa2484bda87fd5215f35cc3217cc05d1dc8a6a978d8f1792a5a9` |
|
| 286 |
+
| `data_diagnostics.md` | `361f8077142e9704fb6d45edf14c8d5c0b0c35e926a99487d78f6c2722b2ed67` |
|
| 287 |
+
| `transformer_v2_35e_panel.json` | `016dd558e2e08fa70cbb9b6f78bdd332e69d2416be78b13dad38affa7aa1574c` |
|
| 288 |
+
| `transformer_v2_minrecall080_fpr020.json` | `a85a02fe164e1493ccddc55d4275951a10422957e6ec62c15587be5a781d1a90` |
|
| 289 |
|
| 290 |
## Verified
|
| 291 |
|
|
|
|
| 296 |
- Transformer v2 result exists.
|
| 297 |
- Gate calibration result exists.
|
| 298 |
- v1 flaw is documented.
|
| 299 |
+
- v3 XOR counterfactual split exists locally and defeats both the ordered-LR baseline and the v2 Transformer recipe (see v3 section above).
|
| 300 |
|
| 301 |
## Unverified / Not Claimed
|
| 302 |
|
|
|
|
| 312 |
- The benchmark is synthetic.
|
| 313 |
- The UI events are rendered primitives, not production browser traces.
|
| 314 |
- The reference result is on generated data only.
|
| 315 |
+
- A flattened all-frames LR is strong because it sees ordered frames; v2 tests ordered temporal evidence, not uniquely deep temporal reasoning. The v3 XOR split removes even this handle and currently has no model that beats baseline.
|
| 316 |
- Gate calibration is a benchmark diagnostic, not a safety policy.
|
| 317 |
- The class schema contains unused classes in this split.
|
| 318 |
|
| 319 |
+
## License
|
| 320 |
+
|
| 321 |
+
Apache-2.0 for the package code, artifacts, and model card. Training and test data are synthetic and generated by the reproduction commands in this repo. No third-party dataset license applies.
|
| 322 |
+
|
| 323 |
## Next Research Steps
|
| 324 |
|
| 325 |
- Add external browser-agent traces with consent and clear provenance.
|
| 326 |
+
- Add stronger adversarial temporal controls where full-frame linear probes degrade (v3 is the first such step).
|
| 327 |
- Add held-out UI event grammars and templates.
|
| 328 |
- Evaluate under real browser screenshots and DOM/action logs.
|
| 329 |
- Separate benchmark scoring from intervention-policy deployment work.
|
| 330 |
+
- Train a v3-targeted model that actually beats v3 baseline before publishing a v3 HF repo.
|