Fix stale docstring: identity-scan prediction was pre-tower; scan decorative at LM inference (internal + Watson independent replication) — see paper §3.2
Browse files- model_v2.py +8 -3
model_v2.py
CHANGED
|
@@ -14,9 +14,14 @@ Operator-native attention (per block = per head):
|
|
| 14 |
|
| 15 |
Because rotors preserve η, score(t,s) = ⟨s_t, (R_q⁻¹R_k)·s_s⟩_η — attention is
|
| 16 |
the alignment of the current state with a LEARNED-RELATIVELY-ROTATED past state.
|
| 17 |
-
Everything is operator-derived → operator-only bottleneck holds
|
| 18 |
-
|
| 19 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 20 |
|
| 21 |
Scan + attention in fp32 (bf16 proven-unsafe on the recurrence).
|
| 22 |
"""
|
|
|
|
| 14 |
|
| 15 |
Because rotors preserve η, score(t,s) = ⟨s_t, (R_q⁻¹R_k)·s_s⟩_η — attention is
|
| 16 |
the alignment of the current state with a LEARNED-RELATIVELY-ROTATED past state.
|
| 17 |
+
Everything is operator-derived → operator-only bottleneck holds. STALE-PREDICTION
|
| 18 |
+
NOTE (2026-09-10, kept for the record): "the R_state→identity bypass collapses to
|
| 19 |
+
unigram" was written for the PRE-TOWER readout and was true of it; with the grade
|
| 20 |
+
tower (shipped config) the readout has a direct route to B_t..B_{t-2} and identity-
|
| 21 |
+
scan costs only ~1.2x (a LayerNorm distribution artifact — time-shuffling S costs
|
| 22 |
+
~1.0x; measured internally and independently replicated by N. Watson 2026-09-10).
|
| 23 |
+
At LM inference the transported state is causally decorative; in sole-channel/tape
|
| 24 |
+
configs it is load-bearing. See paper §3.2.
|
| 25 |
|
| 26 |
Scan + attention in fp32 (bf16 proven-unsafe on the recurrence).
|
| 27 |
"""
|