The 39% vocabulary parameter tax and layer allocation on Colab T4 budgets

#1
by AndrewThompson1233 - opened

Hi westabdu,

Pretraining an 81.17M Turkish generative transformer completely from scratch on a free Colab Tesla T4, building a custom 50k Turkish BPE tokenizer, and driving validation loss down to 2.84 over 5,000 steps is a solid engineering milestone.

Looking at your exact parameter accounting (10 layers, 640 hidden size, 50,000 vocab) and the 512-token context ceiling:

  1. The 32M vocabulary footprint (39.4% of your entire network):
    With a 50,000-token vocabulary at hidden dimension 640, your tied embedding table consumes exactly 32,000,000 parameters.
    Out of your 81,165,440 total parameter budget, that single static lookup table accounts for 39.4% of the entire model.
    In a standard nanoGPT block, one transformer layer (attention plus 4x SwiGLU/GELU MLP) costs roughly 4.92M parameters.
    That means your static embedding table consumes the parameter budget of 6.5 full transformer layers. In a 10-layer model, four out of every ten parameters are sitting in static lookup rows rather than participating in active contextual reasoning.

  2. Reclaiming parameters via low-rank factorization:
    Decoupling token embeddings through a two-stage projection (50,000 -> 128 -> 640 = ~6.48M params) reclaims over 25.5M parameters.
    Reallocating those reclaimed weights directly into computational depth would fund 5 additional physical transformer layers. That would expand the backbone from 10 layers to 15 layers within the exact same 81M ceiling, significantly boosting semantic composition on agglutinative Turkish morphology.

  3. Context scaling and T4 throughput:
    On Colab T4 instances, standard multi-head attention quadratic scaling ($O(L^2)$) typically forces a 512-token limit to avoid out-of-memory errors during training and sluggish decode steps during generation.
    In an open architecture project called Maba v2 (101M reference release: https://huggingface.co/AndrewThompson1233/maba-v2-architecture), we target this exact sub-100M envelope on Tesla T4 hardware:
    We use factorized embeddings to keep the vocabulary tax under 5%, and route 75% of depth through linear recurrence (DGDA) paired with latent sparse attention.
    This eliminates the quadratic attention ceiling and achieves flat O(1) decode latency across long contexts while running fused Triton kernels directly optimized for T4.

If you are planning a v2 run, checking out the parameter allocation layout and factorized embedding setup in the Maba v2 repo might give you some practical reference points for squeezing more depth out of Colab runs.

Did you run into VRAM limits when experimenting with 12 or 14 layers during the Colab training runs?

Best,
Andrew

Hi Andrew,

Thank you so much for the detailed breakdown, compliment, and the architecture recommendations! The math behind factorized embeddings and reclaiming params for actual layer depth makes complete sense—allocating ~39% of the parameter budget purely to static embeddings is definitely something I want to optimize for v2.

To answer your question regarding VRAM limits:
I actually didn't hit any Out Of Memory (OOM) errors during the runs! My setup involves a local GTX 1650 (4GB VRAM) for testing/prototyping, and then offloading the full pre-training runs to Google Colab's Tesla T4.

To make it fit comfortably without choking the VRAM on the T4 (and keeping iteration speed fast), I used a tailored configuration:

  • Batch size: 2 to 8 (depending on the run/stage)
  • Gradient Accumulation: 20 steps
  • Effective batch size: 40 (when batch_size=2) or 160 (when batch_size=8)
  • Context/Block size: 256 to 512
  • Mixed Precision: float16 with use_gradient_checkpointing = True

With this config, the entire pre-training run on the T4 took only about 2.5 hours, running very smoothly without memory issues even when experimenting.

By the way, if you'd like to check out the exact training scripts, tokenization setups, and configuration files I used, they are all freely available on my GitHub repository:

👉 https://github.com/westabdu/Cores-AI-TR-v1.0.0-85M

I'm definitely going to check out the Maba v2 repository and study your factorized embedding setup before launching the v2 architecture. Reclaiming 6+ layers worth of parameters while keeping context processing fast on T4 is huge for Turkish morphology.

Thanks again for reaching out and sharing such valuable insights!

Best regards,
westabdu
GitHub: https://github.com/westabdu/Cores-AI-TR-v1.0.0-85M

Hi westabdu,

Prototyping locally on a 4GB GTX 1650 and offloading full pretraining runs to Colab T4 is a classic, disciplined workflow! Running gradient checkpointing with an effective batch size of 40-160 to finish 3.5B/5k-step runs in 2.5 hours without OOM is very clean execution.

One quick practical tip when you implement factorized embeddings in PyTorch:
To keep input and output embeddings tied with a two-stage projection:

  1. Embed tokens into rank-128: E = nn.Embedding(50000, 128)
  2. Project up to model dimension: W_up = nn.Linear(128, 640, bias=False)
  3. For the LM head at the top, project back down: W_down = nn.Linear(640, 128, bias=False), and then tie the logits projection to the embedding table: logits = F.linear(W_down(x), E.weight)

You can see the exact implementation in maba_sparse/model.py under FactorizedEmbeddings.

Because Turkish is agglutinative, keeping a large 50k vocabulary is critical so suffixes don't fragment into 4-5 tiny subwords. Factorization gives you the best of both worlds: you keep the rich 50k vocabulary coverage while freeing up enough weights to expand from 10 to 15 layers for deeper syntactic composition.

Best of luck with the v2 run, really looking forward to seeing how the 15-layer stack performs on Turkish benchmarks! :)

Best,
Andrew

Hi Andrew,

Thank you so much for the detailed explanation and the practical tip! Every piece of advice you share is genuinely golden for me, as I'm quite new to LLM development and this 85M architecture is actually my very first model.

The factorized embedding approach makes total sense for Turkish—saving parameters while preserving a 50k vocab to jump from 10 to 15 layers is a game changer. I'm definitely implementing this in the v2 architecture!

If you ever want to connect or chat more about model architectures, my Discord is aapo33. I'd be more than happy to add you there! Having your guidance throughout this journey means a lot to me.

I'll keep you updated on how the 15-layer v2 run performs on Turkish benchmarks!

Best,

westabdu

Sign up or log in to comment