pcr-screening-pathx / README.md
complexedleo's picture
Upload PCR + complex screening hybrid, Path-X test acc 0.9350
33b7380 verified
|
Raw History Blame
3.96 kB
metadata
license: mit
tags:
  - pytorch
  - long-range-arena
  - path-x
  - state-space-model
  - linear-recurrence
  - complex-valued-neural-network
  - sequence-classification
datasets:
  - long-range-arena
metrics:
  - accuracy

PCR + Complex Screening Hybrid — Path-X (Long Range Arena)

A phase-coherent linear recurrence (PCR) model — a complex-diagonal LRU/S4D-style recurrence — combined with a non-competing complex screening attention module, trained on Path-X (Long Range Arena), the 16,384-token binary sequence-connectivity task.

  • Task: raw 1D token sequence in, single binary label out. No 2D structure, no auxiliary supervision, no handcrafted features — the same rule-compliant setting as S4/S5/LRU/MEGA on the LRA leaderboard.
  • Test accuracy: 0.9350 (n=20,000, full deterministic sweep)
  • PCR-only ablation (no screening attention): 0.9254 ± 0.0028 (N=2 seeds)
Model Path-X (test)
S4D-Real (θ=0, no phase) chance
S4D-LegS 91.9
PCR (this repo, screening ablated) 92.54 ± 0.28
LRU 94.2
MEGA-chunk 93.81
PCR + screening hybrid (this repo) 93.50
S4 96.35
MEGA 97.98
S5 98.58

Architecture

tokens (B, 16384)
  -> linear encoder (scalar pixel -> d_model)
  -> 6 x PCRBlock:
       [BatchNorm -> PCRLayer (complex diagonal LTI, bidirectional, FFT-conv)
        -> half-GLU -> residual]
     with ComplexScreenBlock inserted after layers 2 and 4:
       [chunked (1024) non-competing complex screening attention:
        L2-normalized complex q,k -> trim-and-square gate
        (no softmax, no row-normalization) -> TanhNorm -> modReLU gate
        -> complex Hadamard -> residual]
  -> LayerNorm -> mean-pool -> linear head -> 2-class logits

Design principle (Phase-Coherent Transformer / PCT): complex eigenvalues implement input-independent phase rotation as coherent long-range transport (a continuous analogue of RoPE); all input-dependent gating, normalization, and readout stay real-valued. ~94% of parameters are complex-valued (100% within the recurrence and attention score/value paths; the ~6% real-valued mass is the input-dependent gates, norms, and readout — kept real by design, not by omission).

Full experimental record, ablations (phase-necessity via a real-eigenvalue control, phase-bandwidth-vs-generalization sweep), and the training/eval harness are described in the source repository (see below).

Files

  • pytorch_model.pt — state_dict only (2,013,716 tensor elements across 116 parameter tensors)
  • config.json — architecture + optimizer config used for this run

Usage

Load with the PCRClassifier / PCRBlock / ComplexScreenBlock definitions from the source training script (train_with_checkpoint.py, cell="pcr" with pcr_config matching config.json's pcr_config field). This repo ships raw weights, not a packaged Python module — see the config for exact hyperparameters to reconstruct the module before calling model.load_state_dict(torch.load("pytorch_model.pt")).

Training details

  • Optimizer: AdamW, base lr 4.5e-4, recurrence/B/C params at 1/3 lr with no weight decay, cosine-hold-then-linear-decay schedule (decay starts at step 200,000), 250,000 steps total, batch size 32.
  • Eigenvalue init: ring |λ| ∈ [0.999, 0.9999], phase restricted to θ ∈ [0, π/10] — the phase bandwidth was found necessary for generalization (a narrower [0, π/50] band memorizes train perfectly but fails to generalize; a real-only ablation, θ=0, fails to learn at all).
  • No dropout, weight decay 0.05, gradient clip 1.0.

Caveats

  • Single seed for the hybrid checkpoint in this repo (N=1); the PCR-only ablation number (92.54 ± 0.28) is averaged over 2 seeds.
  • Not benchmarked beyond Path-X, LRA Text, and LRA Image; no task-specific hyperparameter tuning was performed for those two auxiliary benchmarks.