Link GitHub pathx folder (code + math docs); update to N=3 result 92.71 +/- 0.89
Browse files
README.md
CHANGED
|
@@ -21,21 +21,29 @@ LRU/S4D-style recurrence — combined with a **non-competing complex
|
|
| 21 |
screening attention** module, trained on **Path-X** (Long Range Arena),
|
| 22 |
the 16,384-token binary sequence-connectivity task.
|
| 23 |
|
|
|
|
|
|
|
|
|
|
| 24 |
- **Task**: raw 1D token sequence in, single binary label out. No 2D
|
| 25 |
structure, no auxiliary supervision, no handcrafted features — the same
|
| 26 |
rule-compliant setting as S4/S5/LRU/MEGA on the LRA leaderboard.
|
| 27 |
-
- **Test accuracy**: **0.
|
| 28 |
-
|
|
|
|
| 29 |
|
| 30 |
| Model | Path-X (test) |
|
| 31 |
|---|---|
|
|
|
|
| 32 |
| S4D-Real (θ=0, no phase) | chance |
|
| 33 |
-
|
|
|
|
|
|
|
|
| 34 |
| **PCR (this repo, screening ablated)** | **92.54 ± 0.28** |
|
| 35 |
-
|
|
|
|
|
| 36 |
| MEGA-chunk | 93.81 |
|
| 37 |
-
|
|
| 38 |
-
| S4 | 96.35 |
|
| 39 |
| MEGA | 97.98 |
|
| 40 |
| S5 | 98.58 |
|
| 41 |
|
|
@@ -70,8 +78,9 @@ paths; the ~6% real-valued mass is the input-dependent gates, norms, and
|
|
| 70 |
readout — kept real by design, not by omission).
|
| 71 |
|
| 72 |
Full experimental record, ablations (phase-necessity via a real-eigenvalue
|
| 73 |
-
control, phase-bandwidth-vs-generalization sweep), and the
|
| 74 |
-
harness are
|
|
|
|
| 75 |
|
| 76 |
## Files
|
| 77 |
|
|
@@ -81,12 +90,21 @@ harness are described in the source repository (see below).
|
|
| 81 |
|
| 82 |
## Usage
|
| 83 |
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 90 |
|
| 91 |
## Training details
|
| 92 |
|
|
@@ -102,7 +120,8 @@ hyperparameters to reconstruct the module before calling
|
|
| 102 |
|
| 103 |
## Caveats
|
| 104 |
|
| 105 |
-
-
|
|
|
|
| 106 |
ablation number (92.54 ± 0.28) is averaged over 2 seeds.
|
| 107 |
- Not benchmarked beyond Path-X, LRA Text, and LRA Image; no
|
| 108 |
task-specific hyperparameter tuning was performed for those two
|
|
|
|
| 21 |
screening attention** module, trained on **Path-X** (Long Range Arena),
|
| 22 |
the 16,384-token binary sequence-connectivity task.
|
| 23 |
|
| 24 |
+
**➡️ Code, mathematical documentation, and the paper section:
|
| 25 |
+
[github.com/leohio/phase-coherent-transformer-r-d/tree/main/pathx](https://github.com/leohio/phase-coherent-transformer-r-d/tree/main/pathx)**
|
| 26 |
+
|
| 27 |
- **Task**: raw 1D token sequence in, single binary label out. No 2D
|
| 28 |
structure, no auxiliary supervision, no handcrafted features — the same
|
| 29 |
rule-compliant setting as S4/S5/LRU/MEGA on the LRA leaderboard.
|
| 30 |
+
- **Test accuracy**: **92.71 ± 0.89** over 3 seeds (n=20,000, full
|
| 31 |
+
deterministic sweep); the checkpoint in this repo is the best seed, **93.50**
|
| 32 |
+
- **PCR-only ablation** (no screening attention): 92.54 ± 0.28 (N=2 seeds)
|
| 33 |
|
| 34 |
| Model | Path-X (test) |
|
| 35 |
|---|---|
|
| 36 |
+
| Transformer / Reformer / Performer / Linformer / BigBird / Luna-256 | chance (≈50) |
|
| 37 |
| S4D-Real (θ=0, no phase) | chance |
|
| 38 |
+
| S4-v1 | 88.10 |
|
| 39 |
+
| DSS | 89.72 |
|
| 40 |
+
| S4D-LegS | 91.95 |
|
| 41 |
| **PCR (this repo, screening ablated)** | **92.54 ± 0.28** |
|
| 42 |
+
| S4D-Inv | 92.80 |
|
| 43 |
+
| **PCR + screening hybrid (this repo)** | **92.71 ± 0.89** (best 93.50) |
|
| 44 |
| MEGA-chunk | 93.81 |
|
| 45 |
+
| LRU | 94.20 |
|
| 46 |
+
| S4 (S4-LegS) | 96.35 |
|
| 47 |
| MEGA | 97.98 |
|
| 48 |
| S5 | 98.58 |
|
| 49 |
|
|
|
|
| 78 |
readout — kept real by design, not by omission).
|
| 79 |
|
| 80 |
Full experimental record, ablations (phase-necessity via a real-eigenvalue
|
| 81 |
+
control, phase-bandwidth-vs-generalization sweep), the derivations, and the
|
| 82 |
+
training/eval harness are in the companion repository:
|
| 83 |
+
[phase-coherent-transformer-r-d/pathx](https://github.com/leohio/phase-coherent-transformer-r-d/tree/main/pathx).
|
| 84 |
|
| 85 |
## Files
|
| 86 |
|
|
|
|
| 90 |
|
| 91 |
## Usage
|
| 92 |
|
| 93 |
+
This repo ships raw weights, not a packaged Python module. The model code
|
| 94 |
+
(`PCRClassifier` / `PCRBlock` / `ComplexScreenBlock`, self-contained, torch
|
| 95 |
+
only) and a ready-made loading example are here:
|
| 96 |
+
|
| 97 |
+
**https://github.com/leohio/phase-coherent-transformer-r-d/tree/main/pathx**
|
| 98 |
+
|
| 99 |
+
```python
|
| 100 |
+
import json, torch
|
| 101 |
+
from pcr_screening import build_pcr_classifier # pathx/code/pcr_screening.py
|
| 102 |
+
|
| 103 |
+
cfg = json.load(open("config.json"))["pcr_config"]
|
| 104 |
+
model = build_pcr_classifier(seq_len=16384, vocab=256, **cfg)
|
| 105 |
+
model.load_state_dict(torch.load("pytorch_model.pt", weights_only=True), strict=True)
|
| 106 |
+
model.eval()
|
| 107 |
+
```
|
| 108 |
|
| 109 |
## Training details
|
| 110 |
|
|
|
|
| 120 |
|
| 121 |
## Caveats
|
| 122 |
|
| 123 |
+
- The hybrid result is 92.71 ± 0.89 over 3 seeds (93.50 / 92.89 / 91.75);
|
| 124 |
+
the checkpoint released here is the best of the three. The PCR-only
|
| 125 |
ablation number (92.54 ± 0.28) is averaged over 2 seeds.
|
| 126 |
- Not benchmarked beyond Path-X, LRA Text, and LRA Image; no
|
| 127 |
task-specific hyperparameter tuning was performed for those two
|