NYS β€” Sparse Kuramoto EqProp (public lineage)

NYS is not a transformer. It is a coupled-oscillator organism: phases (\theta_i), amplitudes, natural frequencies (\omega_i), and a sparse coupling graph (K). Language, compiler bits, and speech-frame IDs live on one graph ((N = 5{,}000{,}000) nodes, (k = 512) edges per node). Training is Equilibrium Propagation (free RK4, then nudge RK4, then a contrastive (K) update). There is no backprop and no cross-entropy loss.

Dakuwon Moody (YNSScarSaiyan) Β· Saiyan Corp


Read this first β€” 2026-09-11: the training rule was measured to be broken, and why

Every accuracy number below step ~1,350,000 in this card was produced under a contrastive update that we have since measured to be dominated by noise unrelated to the training signal, on every slot. This is not a data problem, a wiring problem, or a readout problem β€” those were real and were fixed first (see the changelog below) β€” it is a problem in how this specific kernel computed the EqProp contrast. Recording the failure honestly, and the fix, in public:

The mechanism. Each training step: 5 RK4 steps free, then 5 more RK4 steps with the nudge force on, then (\Delta K_{ij} \propto \cos\Delta\theta_{\mathrm{nudge}} - \cos\Delta\theta_{\mathrm{free}}) on gated edges. That contrasts the state 5 steps later against the state 5 steps earlier β€” two different points in time, not two conditions at the same moment. Natural frequencies (\omega_i) span 50–1000 rad/s (host_init_graph, uniform), so in the 0.1 s (5 Γ— dt=0.02) between the two snapshots a typical pair of nodes counter-rotates by ~36 rad from (\omega) alone. The nudge itself moves a target by ~0.003 rad in the same window. We measured this directly by replaying real training steps offline against a live checkpoint, once with the nudge force on and once with it zeroed (same amplitudes, same gated edges), and comparing:

value
Phase drift between snapshots from (\omega) alone ~36 rad
Phase shift a gold target gets from the nudge ~0.003 rad
Size of (\Delta K) that has nothing to do with the nudge (std) ~1.0
Size of the part the nudge actually contributes (mean, std) βˆ’0.00016, 0.0035
Signal-to-noise ratio on the update ~1 : 6,000

At that SNR, (K) on any edge is a random walk that happens to be centered near the target most of the time β€” not a trained value. This explains a pattern that was otherwise puzzling: the speech clamped-pair accuracy went from 16.1% at 930k to 4.4% at 1,350,000 β€” worse, after 420,000 more steps, on a slot that was correctly wired the whole time. A noise-dominated update drifts; it does not reliably improve.

A second, smaller effect compounded this on speech specifically: the target phase ((\bar\theta), what nudged nodes are pulled toward) was computed as the mean over the entire prompt, which for a typical speech step is 32 neighbor-fill tokens spliced in for exploration plus ~2.3 real text tokens. The fill tokens were 91.5% of that mean. Gold frames were being pulled toward the phase of unrelated neighbor tokens, not the phase of the text that was supposed to teach them.

The fix (deployed 2026-09-11, verified offline, live verification in progress β€” check the changelog / latest checkpoint metadata for confirmed results before citing numbers past this point):

  1. Same-moment contrast. After the free phase, two branches run from the identical state: one continues free, one is nudged, both for the same number of steps. (\Delta K) contrasts those two β€” the (\omega) rotation is common to both branches and cancels exactly. Costs one extra 5-step integration per training step.

  2. Target from the real prompt only, not the neighbor-fill tokens.

  3. (\omega) scaled by 0.01 inside the integrator (--omega-scale, the stored (\omega) in the checkpoint is untouched and this is reversible by restarting with a different value). Comparing branches removes the drift, but at the original (\omega) magnitude no pair can phase-lock within a 5-step nudge window regardless β€” the coupling term is too slow relative to the free rotation. Offline replay on 40 real speech records, same edges, before vs. after:

    Rule gold-edge (\Delta K) negative-edge (\Delta K)
    Old (time-shifted, (\omega) full) βˆ’0.00012 (βˆ’2.1 SE) β€” sign noise +0.0001 (+1.6 SE)
    Same-moment, text-only mean, (\omega) full βˆ’0.00012 (βˆ’2.1 SE) β€” still noise +0.0001
    Same-moment, text-only mean, (\omega \times 0.01) +0.0044 (+23 SE) βˆ’0.0048 (βˆ’22 SE)

    Both signs correct, both far outside noise, only once (\omega) is slowed. A HIP self-test confirmed the new kernel runs correctly on-GPU at both settings before it went into the training loop.

What this means for the numbers in the rest of this card. They are real measurements of a real (and, until now, undiagnosed) failure mode β€” useful for exactly that reason β€” but they are not evidence about what this graph and this coupling rule can or cannot learn. Compiler and speech being stuck near chance was previously attributed to topology and shared-register conflicts (both real, both fixed at 925k and 2026-09-10 respectively); it now looks like the update itself never had the SNR to learn regardless. That question is open again, correctly this time, and is what the current run is testing.


Two artifacts on this repo β€” do not confuse them

Lineage Files Size Graph Trainer Can tlc-infer load it?
Current β€” sparse GPU EqProp gpu_eqprop_<step>_<utc>.bin 20,640,000,016 B exactly (N=5\times10^6), CSR (k=512) HIP eqprop_gpu (no PyTorch) No
Legacy β€” dense ASM SFT nys_sft_final.bin ~32–34 GiB (N=65{,}536) dense (K) x86-64 NASM / AVX-512 Yes (old path)

The model card you are reading describes the sparse GPU lineage. The legacy file is kept for history. tlc-infer / sampler.cpp still assume the 65,536 dense (K). Dropping a gpu_eqprop_*.bin into that binary will not work.

A complete CPU sparse base also exists on YNSScarSaiyan/nys-checkpoints (base_sparse_*). This public GPU run did not resume that file. It was --init (random ring + stubs) at step 0, then trained on-card. Same (N), same (k) the whole way. No second --init, ever, on this lineage.


Changelog (most recent first)

  • 2026-09-11 β€” contrastive-update SNR fix. Same-moment branch contrast (cancels (\omega) drift exactly), target phase from real prompt tokens only, (\omega \times 0.01) inside the integrator. See the section above. Offline-verified; live results pending.
  • 2026-09-10 β€” speech data/training fixes, in order:
    • Vocoder codebook rebuilt: the old k-means fit left 235 of 256 atoms at their identity-DC initialization (empty-cluster failure β€” a centroid with no assigned frames never moves). One code covered 79% of every frame in the training corpus. Refit with data-seeded k-means++ and empty-cluster reseeding from the worst-fit frame: 1 of 256 atoms left flat, all 256 codes in use, reconstruction SNR 2.9β†’9.2 dB.
    • Silence frames (the code that maps to near-zero-variance PCM, ~65–75% of real speech audio) excluded from both training targets and negative examples, on both the HIP trainer and the Python wiring pass. They were previously being pulled toward every record's own phase simultaneously β€” the same "shared register, many writers" conflict described below for the compiler band.
    • Speech negative-example generation fixed: the old rule added code+1 and code+127 neighbors of every gold frame as negatives appended to the prompt (so they dominated the target-phase mean; see above), and did not check whether a negative for one record was gold for another. Measured on the live corpus: 39.9% of positive pushes were being fought by a negative push on the identical physical node. Fixed: one negative per gold frame, checked against a corpus-wide gold set (load_global_gold_frames) so a negative can never be a duplicate of someone else's target.
    • (K)-value clamp added to the kernel ((\lvert K \rvert \le 3)): contested nodes had walked to (\lvert K \rvert = 16)–18 (vs. a wiring-pass seed of (K=0.2)) before this was caught; a stale hard reset of the speech band's edges (--reset-speech) was run once to clear the accumulated damage.
  • 2026-09-09 (925k) β€” CSR wire. Random ring+stubs never connected vocab tokens to the compiler band (3,800,000+) or the speech band (4,300,000+). Wired in place on existing (k=512) slots (weakest non-protected edge replaced, seed (K=0.2)). (N), (k) unchanged.
  • Earlier: neighbor-fill (32), speech (\beta) floor (0.25), compiler signed 0-bit nudge β€” see prior card revisions in this repo's commit history.

Status at 1,350,000 β€” under the pre-2026-09-11 rule (see caveat above)

These are the last numbers produced before the SNR fix. Treat them as a record of the failure mode, not a capability grade.

Slot 930k 1,350,000 Trend
Compiler free bits (clamped-pair) 49.2% 52.8% Flat-ish, chance β‰ˆ 50% either way (38 free bits, 11 shared gadgets)
Speech codes (clamped-pair, in-pool) 16.1% (in-pool 58%) 4.4% (in-pool 41%) Degraded β€” random-walk signature
Language headline 0.625% 0.0% Noise-level throughout

Compiler additionally has its own, separate confound even with a correct update: 11 gold gadgets share the same 64 bit-oscillators. A settle-based probe (clamp the spec, run free RK4, read where the 64 nodes land β€” bypassing the old "read whatever theta a node was last left at" readout convention) confirmed the band is driven by the spec (no zero-displacement edges), but different specs produced statistically unrelated crystals (pairwise Hamming ~30/64, chance is 32) β€” the register is being fought over, not un-addressed. That conflict is unresolved and is a second, independent reason compiler will need more than the SNR fix.

Sidecar quizzes at 930k (historical; see caveat above)

Quiz 930k What it actually measures
Compiler Hamming free 15 / 38 (chance β‰ˆ 19) Unconditioned 64-bit crystal vs 11 golds
Speech frame_match (max-over-shift) 1.49%, locked=false One global argmax tape vs many gold NYSV records β€” low ceiling by construction even under a working update
Vocoder codec (pre-rebuild) 21/256 atoms moved Superseded 2026-09-10; now 255/256
Chat (KOPG, SFT prefix) fragments, looped=false Static 2-hop neighbor walk, not HIP RK4 generation

930k chat probe (for the record, unaffected by any of the above fixes since language readout is a separate mechanism from the SNR issue affecting the trained signal β€” though the same broken contrastive update was training it the whole time):

  • hello β†’ based on you provide of your was a by the there are based, we importantafter
  • the result is β†’ l0_wra in the of the
  • nys is listening β†’ "_i3gw98h

Not a chatbot. Neighbor structure is not uniform noise; replies are still template / salad. Whether that ceiling is topology (briefing hypothesis: fixed ring+stubs never contain the right edge) or the same SNR problem as speech/compiler is now an open question again β€” it hasn't been re-tested under the fixed update yet.


What the organism is

NYS is one substrate with five slots. Slots 1, 2, and 4 train. Slot 3 is a crystal→beep readout. Slot 5 is a dump envelope (not trained).

Slot Trains? IDs / files What it is
1 Language Yes Frozen vocab_sparse.json (~3,664,150 IDs). Hard next-token + SDS1 distill + live mix. Hierarchical token IDs.
2 Compiler Yes 3,800,000 + bit, 64 active bits (8 bytes Γ— 8 spins). compiler_physics.nysa (412 recs). Spec-conditioned gold x86 gadgets. 11 gadgets share one 64-bit register β€” unresolved conflict, independent of the SNR fix.
3 Voice crystal No Digit tones from the crystal integer (350 Hz + 50 Hz/digit, 8 kHz). Beeps that report a number. Not speech.
4 Speech Yes 4,300,000 + (t\cdot 256) + code. speech_slot.nysv (6) + speech_align.nysv (305). Vocoder speech_vocoder.nyvc (rebuilt 2026-09-10, 255/256 atoms trained). Frame-code IDs on the same graph.
5 Dump No dump_slot.nysd (NYSD kinds 6/7). Inbox wrap only. Debug dumps. Do not train NYSD.

Compiler execution of open-ended x86 is a quench + crystallize + mprotect/CALL path (execute.asm). This sidecar does not execute the crystal. Demo gold (8-byte active, LSB-first spins):

  • f(x)=x+42: 48 89 F8 48 83 C0 2A C3
  • f(x)=x*x: 48 89 F8 48 0F AF C7 C3
  • plus ident, inc, dec, neg, not, shl1, add_self, zero, sub1 (11 gadgets on the same 64 bit nodes β€” this is the shared-register conflict noted above, separate from and in addition to the SNR issue)

Spins freeze LSB-first, 8 spins/byte, bit = (\mathrm{sign}(a_i \cos\theta_i)). 26 bits are tied across all padded golds; 38 are free. Grade free bits.


Physics and trainer

Equation and active set

The HIP fused kernel (eqprop_gpu.hip, hipcc, --offload-arch=gfx942) does not wrap all (N) oscillators and does not explode 512 neighbors of neighbors. As of the 2026-09-11 fix, each training step:

  1. Build (S = \mathrm{unique}(\mathrm{prompt} \cup \mathrm{nudge\ IDs})), hard cap 256. Prompt is packed first; if (|S|>256), the tail is dropped.
  2. Prompt nodes start at amplitude 1; others 0.
  3. Neighbor fill: up to 32 unseen CSR neighbors of the last prompt token are spliced into the prompt (amp 1) for exploration. The target-phase mean (next step) is computed from the real prompt only, excluding these β€” fixed 2026-09-11; previously the fill dominated the mean.
  4. Free RK4: 5 steps, (\mathrm{d}t = 0.02), (\omega) scaled by --omega-scale inside the integrator (default in the current run: 0.01; the checkpoint's stored (\omega) is never modified, so this is a launch-time choice, reversible by restarting with a different value).
  5. Branch point. From the free-phase end state:
    • Branch A (free-continued): 5 more RK4 steps, no nudge.
    • Branch B (nudged): 5 more RK4 steps from the same starting state, force (\beta \sin(\bar\theta - \theta_i)) on nudge targets. (\bar\theta) is the circular mean of the real prompt at the branch point (see step 3).
  6. Nudge targets get amplitude (\min(1, |\beta|)) for branch B. Amp gate for (K) updates is 0.1.
  7. Contrastive (K): on CSR edges whose both ends have amp (> 0.1), (\Delta K_{ij} \leftarrow \eta,(\cos\Delta\theta_{B} - \cos\Delta\theta_{A})) β€” both branches measured at the same elapsed time from the same starting state, so (\omega)-driven rotation is common to both and cancels. (K) is clamped to (\lvert K \rvert \le 3) after each update.
  8. Amplitudes written back to 0.

Before 2026-09-11, step 5 did not exist: branch B was compared directly against the step-4 (free) snapshot, 5 steps earlier in time. See the diagnosis at the top of this card.

Defaults: (\eta = 0.05), (\beta = 1.5). One HIP block, 256 threads. Model state ~20.64 GB HBM. Card share is capped (0.40 of device, max 77 GiB, leave 16 GiB free). No hipDeviceReset.

Contrast is the native signal

A printed contrast=0.00000 (or a slot column of 0.00000) means no amp-gated edges in (S), not a perfect model. It is a heartbeat (mean (\Delta\cos) on gated edges), not accuracy, and β€” as of 2026-09-11 β€” is understood to have been dominated by an artifact for every slot prior to the branch fix. Post-fix, it is the same statistic computed on a rule with verified nonzero SNR; still not a substitute for the clamped-pair / settle-based accuracy checks.

Live rotation

Each loop tries, in order: mix (tailed hard file) β†’ hard (train + SFT wrap) β†’ SDS1 distill β†’ NYSA compiler β†’ NYSV speech.

HIP holds NYSA/NYSV as FILE* for the life of the process. Slot updates must be atomic mv onto those paths; a truncate/scp onto an open slot file will crash the trainer.


Sparse checkpoint layout

Little-endian. Reject any file whose size is not 20,640,000,016. Saves are atomic (.writing then rename). Resume only from a complete file. Layout is unchanged by the 2026-09-11 fix β€” \omega in the file is the same value it always was; scaling happens only inside the integrator at load time via --omega-scale.

Field Type Count Notes
N uint32 1 5,000,000
k uint32 1 512
theta float64 (N) phase
amp float64 (N) 0 after a finished step
omega float64 (N) natural frequency, stored value unchanged by --omega-scale
K_row_offsets uint64 (N+1) CSR; row (i) is ([i k, (i+1)k))
K_col_indices uint32 (N \cdot k) ring + stubs; rewired at 925k (compiler/speech) and 2026-09-10 (speech reset)
K_values float32 (N \cdot k) learned couplings, clamped to (\lvert K \rvert \le 3) as of 2026-09-11

Offsets:

theta_off  = 8
amp_off    = 8 + 8N
omega_off  = 8 + 16N
row_off    = 8 + 24N
col_off    = row_off + 8(N+1)
val_off    = col_off + 4 N k
end        = val_off + 4 N k   = 20,640,000,016

Init topology (this lineage): for each row, (k-16) local ring (i - k/2 + j) mod N, last 16 random stubs. Training updates values on those edges. At 925k, a CPU pass replaced the weakest non-protected column in some rows so vocab ↔ compiler/speech IDs share an edge. At 2026-09-10, the speech band's edges were additionally force-reset to the wiring seed value once, to clear damage accumulated before the silence/ negative-collision fixes existed. (k) is still 512 throughout.

Filename: gpu_eqprop_<steps>_<YYYYMMDD>_<HHMMSS>.bin (UTC). The trainer parses steps from the name (or symlink target) for --resume-steps.


Tokenizer (do not retrain)

vocab_sparse.json (~107 MB). HierarchicalTokenizer: bytes 0–255, then greedy 4-gram phrases. The frozen space-join encode bug is part of the ID space β€” do not "fix" it or IDs will not match (K).

IDs above vocab and below (N) are reserved bands (compiler / voice entropy / speech / dump). They are first-class oscillators, not a second model.


Training data

Corpora live on YNSScarSaiyan/nys-corpus unless noted. Slot files also upload under slots/ on this repo.

Slot 1 β€” language

File Layout Records
train_corpus_sparse.bin hard: uint32 nrec + (uint16 len + uint32 toks[len]), last = target 255,435
sft_corpus_sparse.bin same hard layout 50,000
mix_sft_sparse.bin same hard layout (Rust sparse_convert from JSONL) 2,050,000
distill_corpus_sparse.bin SDS1: prompt + soft id/prob 40,000

Slot 2 β€” compiler physics (NYSA)

412 records: I/O triples (kind 1: 385), gold bytes (11), structural bit maps (5), traces (11). HIP nudges bit nodes at 3_800_000 + i with (+\beta) (spin 1) or (-\beta) (spin 0). All 11 gadgets share the same 64 oscillators β€” confirmed via settle-based conditioning probe to be a genuine unresolved conflict, not a readout artifact.

Slot 4 β€” speech (NYSV + NYVC)

File Records Role
speech_slot.nysv 6 Seed utterances
speech_align.nysv 305 Aligned text ↔ frame IDs (≀64 frames / rec)
speech_vocoder.nyvc 256 atoms Γ— 80 samples @ 8 kHz VQ table, rebuilt 2026-09-10 (255/256 atoms trained, was 21/256)

Frame IDs = 4_300_000 + t*256 + code. ALIGN_MAX_FRAMES=64 keeps (|S|) under the 256 cap. As of 2026-09-10, silence-coded frames are excluded from both nudge targets and negative examples; negatives are checked against a corpus-wide gold set to guarantee zero collision with another record's target.

Slot 5 β€” dump (NYSD)

Reserved envelope + inbox. The trainer does not EqProp NYSD.


Sidecar (CPU, after every complete snap)

sidecar.py mmaps the checkpoint. Does not run the HIP kernel. Its readouts (chat KOPG, crystal Hamming, speech frame_match) score stored theta directly β€” which, independent of the SNR issue, is the value a node was left at by whichever record touched it last, not a response to a live prompt. A separate read-only inference mode (eqprop_gpu --settle: clamp a prompt, run free RK4, read where target nodes land, no training-state mutation) exists and was used for the compiler conditioning check above; it is not yet wired into the routine sidecar quiz path.

Outputs (also uploaded here):

sidecar/gpu_eqprop_<step>_<utc>/
  train_acc.json  # clamped-pair acc (language / compiler / speech)
  eval.json       # train_acc + chat + free Hamming + speech lock + codec
  ...

HF upload: complete bins only (exact size, settle before upload). Retention deletes a local .bin only after this repo lists that filename at 20,640,000,016 B, and always keeps the newest local complete file for resume.


How to load (sparse)

import os, struct, mmap

FULL = 20_640_000_016
path = "gpu_eqprop_1350000_20260911_071500.bin"  # use the newest on this repo
assert os.path.getsize(path) == FULL

with open(path, "rb") as f:
    mm = mmap.mmap(f.fileno(), 0, access=mmap.ACCESS_READ)
    N, k = struct.unpack_from("<II", mm, 0)
    assert N == 5_000_000 and k == 512

Train / resume (HIP, AMD Instinct MI300X / gfx942). Never --init unless you intend to throw this lineage away.

hipcc -O3 -std=c++17 --offload-arch=gfx942 -o eqprop_gpu eqprop_gpu.hip
./eqprop_gpu \
  --data-dir ./data --live --live-mix ./data/mix_sft_sparse.bin \
  --ckpt ./gpu_eqprop_1350000_20260911_071500.bin \
  --save-every 5000 --save-dir ./ckpts \
  --vram-limit-gb 77 --device 0 --lr 0.05 --beta 1.5 \
  --omega-scale 0.01

Omit --omega-scale (or pass 1) to reproduce the pre-2026-09-11 dynamics exactly β€” the checkpoint format and stored values are identical either way.


Files in this repo

  • gpu_eqprop_<step>_<utc>.bin β€” sparse snapshots every 5,000 steps. Each file is the full substrate, not a delta. Reject any file that is not 20,640,000,016 B.
  • sidecar/gpu_eqprop_*/ β€” CPU readouts.
  • slots/ β€” NYSA / NYSV / NYVC / NYSD copies used by this run.
  • nys_sft_final.bin β€” legacy 65,536 dense SFT (see table above).
  • PROGRAM_CHARTER.md β€” cross-program charter.

Newest gpu_eqprop_* by step number is the current public checkpoint.


Hardware and stack

Item This public GPU run
GPU AMD Instinct MI300X (ROCm / HIP), one block of 256 threads
Compile hipcc, C++17, gfx942 β€” no PyTorch, no JAX in this process
HBM ~20.64 GB model; cap 77 GiB; leave 16 GiB for other jobs
Tokenizer Frozen hierarchical phrase vocab (~3.66M IDs)
Author Dakuwon Moody (YNSScarSaiyan)

Limitations (read this)

  • The single most important limitation is the one at the top of this card: every accuracy number from before 2026-09-11 was produced under a contrastive update with a measured signal-to-noise ratio of roughly 1:6,000. Treat those numbers as documentation of the failure mode, not as evidence about the organism's ceiling in either direction.
  • The fix is offline-verified, not yet live-verified as of this writing. Check the changelog or the latest checkpoint's sidecar output before citing any post-fix accuracy claim as settled.
  • Compiler has a second, independent problem: 11 gadgets share one 64-bit register, confirmed via a settle-based conditioning probe to be a genuine write conflict, not a readout artifact. Fixing the SNR does not by itself fix this.
  • Not a production chat model. No CE, no instruction-eval scores claimed here. Contrast β‰  quality grade. KOPG chat β‰  EqProp generation.
  • This GPU lineage started from random init, not the completed CPU sparse base.
  • Sparse β‰  dense. You cannot mmap these bins as a 65,536Γ—65,536 K.
  • Voice (slot 3) is digit-tone crystallize, not speech synthesis.
  • Do not retrain vocab_sparse.json.
  • Do not --init this lineage unless you mean to start over.
  • Research artifact. Architecture-locked to this (N,k) and tokenizer.

Related

Updated 2026-09-11 β€” diagnosed and fixed a contrastive-update SNR problem present since this lineage's --init (affects every slot); rebuilt the speech vocoder; fixed speech silence/negative-collision training bugs; added a settle-based (clamp-and-integrate) inference probe used to separate readout artifacts from training artifacts on the compiler band.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support