Addressing entity drift and vocabulary allocation on RTX 4060 Laptop budgets

#1
by AndrewThompson1233 - opened

Hi Duoia,

Training a clean 39M causal Llama-style model from scratch on a single RTX 4060 Laptop in under 3.5 hours using LitGPT is a really impressive zero-budget project. Getting validation perplexity down to 2.99 on the SFT stage demonstrates clean data hygiene.

Looking at your parameter allocation (38.85M) and the limitation noted regarding late-sequence entity drift:

  1. Syntactic depth versus narrative drift:
    At 11 layers with 512 hidden width, each transformer block costs roughly 3.15M parameters.
    Across a 512-token narrative, an 11-layer stack has limited non-linear composition depth to maintain character state and relational bindings, which directly causes the entity confusion and repetition you noted in longer outputs.
    In an open architecture project called Maba (101M reference model: https://huggingface.co/AndrewThompson1233/maba-v1-architecture), we address this on TinyStories using deterministic 2-pass block recycling:
    Routing hidden states through your 11 physical blocks twice with Split RMSNorm (distinct norm vectors for pass 0 and pass 1) expands depth to 22 effective layers at zero additional parameter cost. This added depth provides the relational capacity needed to keep character identities stable throughout the story.

  2. Embedding parameter footprint:
    Even with tied weights, an 8,192 vocabulary at 512 hidden dimension consumes 4.19M parameters (10.8% of your entire 38.85M budget). That single lookup table costs more than an entire transformer block.
    Decoupling the input lookup via low-rank factorization (8,192 -> 64 -> 512 = ~0.56M params) reclaims roughly 3.63M parameters, freeing enough parameter budget to add a 12th full layer within the exact same 39M ceiling.

Did 8GB VRAM limits enforce the 11-layer depth during the RTX 4060 training run?

Best,
Andrew

Hi Andrew,

Thank you so much for the detailed feedback and thoughtful insights!

To address your question: since this project was primarily built for learning language model architecture from scratch, my main priority was fast iteration speed to validate my ideas quickly. At the same time, I was worried that making the model too shallow would hamper its ability to learn meaningful representations, so I settled on 11 layers. With a model of this small size, VRAM wasn't really the main bottleneck, so I wasn't too worried about OOM on the 8GB RTX 4060.

Your suggestions are spot on:

Embedding Table: You make a great point. For a dataset mostly focused on children's stories (TinyStories), an 8,192 vocabulary is indeed larger than necessary and takes up a disproportionate share of the parameter budget.

Block Recycling: The deterministic 2-pass block recycling with Split RMSNorm in Maba is super inspiring! Doubling the effective depth without adding parameter overhead is a brilliant solution for narrative drift.

I’ll definitely take a close look at the Maba architecture and test these ideas in my next iteration.

(P.S. English isn't my native language, so I had an LLM polish this response!)

Best regards,

Duoia

Hi Duoia,

Appreciate the candid response! Training from scratch on a laptop GPU is the best way to get a solid intuition for scaling laws and depth tradeoffs.

For TinyStories specifically, dropping the vocab or factorizing it really opens up room for extra active compute. And with 2-pass recycling, you get the representational depth of a 22-layer model while keeping the fast parameter iteration loop of an 11-layer run.

No worries at all about the English, everything came through perfectly clear.

Good luck with the v2 experiments, excited to see the next checkpoint!

Best,
Andrew

Sign up or log in to comment