Resolving role consistency and memorization bottlenecks via 2-pass block recycling

#1
by AndrewThompson1233 - opened

Hi Siva,

Documenting the failure modes of deceptive validation metrics (the uninitialized RoPE inv_freq buffer and teacher-forcing loops) is one of the most transparent and insightful write-ups on Hugging Face.

Looking at your architecture trade-offs and the C8 (role consistency = 0.0) / C5 (field accuracy = 71.7%) gates:

  1. Parameter capacity vs training token exposure:
  • 3,800 documents across 256 tokens for 3 epochs equals ~2.9M tokens of total exposure.
  • At 53M parameters, the network has ~18x more weights than training tokens. As you noted regarding the 97.5% vs 61.2% split, 53M parameters easily memorize synthetic fixtures rather than learning generalized extraction transforms.
  1. Role inversion in unlabelled running prose:
    Disentangling debtor from creditor without canonical field tags requires multi-step syntactic binding: first identifying entity spans (IBANs, names, amounts), then traversing prepositional phrases ("on behalf of", "favoring") to assign semantic roles. With d_model=512 and intermediate FFN=768, 20 shallow layers struggle with this multi-hop resolution.

In an open architecture project called Maba (101M reference model: https://huggingface.co/AndrewThompson1233/maba-v1-architecture), we handle this regime using deterministic 2-pass block recycling:
Reduce physical depth from 20 blocks to 10, and cycle representations through them twice during the forward pass.
Condition each pass using Split RMSNorm (distinct scale vectors for pass 0 and pass 1). Pass 0 handles token-level span detection, while pass 1 refines relational role assignment.
This cuts physical weights from 53M down to ~33M parameters, curbing the capacity for verbatim memorization on small corpora while preserving 20 effective layers of relational depth.
Compute and memory fit effortlessly within the 4 GB VRAM ceiling of mobile GPUs like your RTX 2050.

Did you test whether scaling down physical layer depth or varying parameter-to-token ratios improved generalization on unseen hostile-difficulty draws?

Best,
Andrew

Sign up or log in to comment