YMRohit's picture
Update logbook: Repro - Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
d371ed1 verified
|
Raw History Blame Contribute Delete
5.81 kB

Conclusion


The core limit theorem reproduces exactly; the two quantitative claims reproduce only in their literal form, and we identify precisely why. CPU/numpy only, ~25 min, $0.

The scaling convention is the crux, and it is unusual. The paper's logits are the raw bilinear form <Qz, Kz'> with no 1/sqrt(d), no 1/d, no temperature and no prompt-length scaling β€” the 1/L lives only in the empirical measure and cancels inside the softmax ratio. So the 'infinite-prompt limit' is a plain law-of-large-numbers limit at fixed U, V, d. Because the logits are unnormalised, every concentration constant carries a factor exp(4 sigma^2 ||Uz||^2), and that single factor drives both of our partial verdicts. Identifying this was the heart of the reproduction: many 'softmax becomes linear' claims hold only under a specific normalisation.

C1 (softmax -> linear operator) β€” REPRODUCED. Lemma 2.1's closed form is an exact identity, not merely asymptotic: independent Gauss-Hermite quadrature matches it to 2.9e-15 (criterion 1e-8). Convergence in prompt length holds with slope -0.982 to -1.027 across d in {2,5,10} x ||U|| in {0.1,0.3,0.5}, monotone in 9/9 settings; fitting the paper's own form gives an exponent c2 in [1.15,1.20] where the paper only asserts c2>0. Three negative controls confirm the assumptions are load-bearing: Student-t3 inputs make the error grow by 6.3e3x, scaling U with sqrt(L) makes it grow 27.6x, and non-Gaussian tokens plateau at a bias of 2.177e-2 against the exact predicted 2.158e-2 while still converging to their own different limit.

C2 (non-asymptotic concentration + gradient stability) β€” PARTIALLY REPRODUCED. The explicit bounds are never violated (0/236 tail tests over 147,000 prompts; 0/36 cells for the gradient lemma; gradient rates -0.899 to -1.043) β€” but they are never approached either: median looseness 1470x and 2.0e6x respectively, with 2 of 36 cells outright vacuous. Trajectory stability fails our predeclared band: the gap at fixed parameters decays as L^-0.526, but taken uniformly over training it degrades to L^-0.342 and is non-monotone. We derived and confirmed the mechanism rather than leaving it as a failure: at the Bayes optimum the query-averaged constant is (1 - 2 tr(Sigma^-2)^-1/2)^-d/2, which is finite if and only if tr(Sigma^-2) > 4. Sweeping that quantity from 0.75 to 14.81 moves the uniform exponent monotonically from +0.000 +/- 0.029 (no convergence at all) to -0.386, while the fixed-parameter exponent stays pinned near -1/2.

C3 (inheriting linear-attention structure) β€” PARTIALLY REPRODUCED. The paper's literal statement passes: held-out risk converges to the stated optimum with slope -0.722/-0.757, reaching 0.95%/0.64% of initial risk, for both isotropic and anisotropic covariance, and the predicted block structure is preserved to ~1e-17 (verified, not imposed). The colloquial reading β€” that softmax and linear trajectories converge to each other at fixed horizon β€” fails: parameter gaps decay only as L^-0.134/L^-0.176 and remain 15.7%/11.1% of the limit norm at L=4096. Since the paper explicitly declines to claim fixed-horizon convergence, this falsifies the folk reading, not the theorem.

Scope & cost

This reproduction Full replication
Scope Numerical verification of Lemma 2.1, the L-limit, concentration/gradient bounds, and training-trajectory transfer, on synthetic Gaussian data at the paper's own scale Adds larger d, real data, deeper architectures
Hardware CPU only (numpy; no torch/scipy/GPU) same class β€” theory paper
Compute time ~25 min hours
Cost $0 (local) minimal
Outcome C1 reproduced; C2, C3 partially reproduced β€”

Honest limitations. Propositions 3.1/3.4 carry existential constants c1,c2 the paper never quantifies; we recorded before running that these are not numerically falsifiable, so 'bound held' is reported as consistent with, never as confirmation. Three deviations are disclosed rather than hidden: one control was rerun at a larger ||U|| because the predeclared setting produced a bias too small to resolve (flagged post-hoc); the C2 threshold study is entirely post-hoc, written to explain a predeclared FAIL; and one training run was redone at a smaller learning rate after the original made the anisotropic risk transiently increase (an unfaithful discretisation of the flow), with the superseded artifact deleted and a provenance note kept. One small-L control is reported INCONCLUSIVE (2.8x/4.0x against a predeclared 5x) rather than counted as a pass. C3 used only 2 seeds β€” the non-monotone bump exceeds those error bars, but 2 seeds is thin.


🎯 Trackio dashboard icml5886-softmaxlinear

https://huggingface.co/spaces/YMRohit/icml5886-softmaxlinear


πŸ“¦ Artifact icml5886-softmaxlinear/icml5886-softmaxlinear-bundle:v0 Β· dataset

https://huggingface.co/buckets/YMRohit/icml2026-5886-softmax-linear-attention-logbook-artifacts#icml5886-softmaxlinear/icml5886-softmaxlinear-bundle:v0