pcr-screening-pathx / README.md
complexedleo's picture
Upload PCR + complex screening hybrid, Path-X test acc 0.9350
33b7380 verified
|
Raw History Blame
3.96 kB
---
license: mit
tags:
- pytorch
- long-range-arena
- path-x
- state-space-model
- linear-recurrence
- complex-valued-neural-network
- sequence-classification
datasets:
- long-range-arena
metrics:
- accuracy
---
# PCR + Complex Screening Hybrid — Path-X (Long Range Arena)
A **phase-coherent linear recurrence (PCR)** model — a complex-diagonal
LRU/S4D-style recurrence — combined with a **non-competing complex
screening attention** module, trained on **Path-X** (Long Range Arena),
the 16,384-token binary sequence-connectivity task.
- **Task**: raw 1D token sequence in, single binary label out. No 2D
structure, no auxiliary supervision, no handcrafted features — the same
rule-compliant setting as S4/S5/LRU/MEGA on the LRA leaderboard.
- **Test accuracy**: **0.9350** (n=20,000, full deterministic sweep)
- **PCR-only ablation** (no screening attention): 0.9254 ± 0.0028 (N=2 seeds)
| Model | Path-X (test) |
|---|---|
| S4D-Real (θ=0, no phase) | chance |
| S4D-LegS | 91.9 |
| **PCR (this repo, screening ablated)** | **92.54 ± 0.28** |
| LRU | 94.2 |
| MEGA-chunk | 93.81 |
| **PCR + screening hybrid (this repo)** | **93.50** |
| S4 | 96.35 |
| MEGA | 97.98 |
| S5 | 98.58 |
## Architecture
```
tokens (B, 16384)
-> linear encoder (scalar pixel -> d_model)
-> 6 x PCRBlock:
[BatchNorm -> PCRLayer (complex diagonal LTI, bidirectional, FFT-conv)
-> half-GLU -> residual]
with ComplexScreenBlock inserted after layers 2 and 4:
[chunked (1024) non-competing complex screening attention:
L2-normalized complex q,k -> trim-and-square gate
(no softmax, no row-normalization) -> TanhNorm -> modReLU gate
-> complex Hadamard -> residual]
-> LayerNorm -> mean-pool -> linear head -> 2-class logits
```
**Design principle (Phase-Coherent Transformer / PCT)**: complex
eigenvalues implement input-independent phase rotation as coherent
long-range transport (a continuous analogue of RoPE); all input-dependent
gating, normalization, and readout stay real-valued. ~94% of parameters
are complex-valued (100% within the recurrence and attention score/value
paths; the ~6% real-valued mass is the input-dependent gates, norms, and
readout — kept real by design, not by omission).
Full experimental record, ablations (phase-necessity via a real-eigenvalue
control, phase-bandwidth-vs-generalization sweep), and the training/eval
harness are described in the source repository (see below).
## Files
- `pytorch_model.pt` — `state_dict` only (2,013,716 tensor elements across
116 parameter tensors)
- `config.json` — architecture + optimizer config used for this run
## Usage
Load with the `PCRClassifier` / `PCRBlock` / `ComplexScreenBlock` definitions
from the source training script (`train_with_checkpoint.py`, `cell="pcr"`
with `pcr_config` matching `config.json`'s `pcr_config` field). This repo
ships raw weights, not a packaged Python module — see the config for exact
hyperparameters to reconstruct the module before calling
`model.load_state_dict(torch.load("pytorch_model.pt"))`.
## Training details
- Optimizer: AdamW, base lr 4.5e-4, recurrence/B/C params at 1/3 lr with no
weight decay, cosine-hold-then-linear-decay schedule (decay starts at
step 200,000), 250,000 steps total, batch size 32.
- Eigenvalue init: ring `|λ| ∈ [0.999, 0.9999]`, phase restricted to
`θ ∈ [0, π/10]` — the phase bandwidth was found necessary for
generalization (a narrower `[0, π/50]` band memorizes train perfectly
but fails to generalize; a real-only ablation, θ=0, fails to learn at
all).
- No dropout, weight decay 0.05, gradient clip 1.0.
## Caveats
- Single seed for the hybrid checkpoint in this repo (N=1); the PCR-only
ablation number (92.54 ± 0.28) is averaged over 2 seeds.
- Not benchmarked beyond Path-X, LRA Text, and LRA Image; no
task-specific hyperparameter tuning was performed for those two
auxiliary benchmarks.