--- license: odc-by datasets: - HuggingFaceFW/fineweb-edu language: - en library_name: transformers pipeline_tag: text-generation tags: - nanodex - tiny-lm - pretrained-from-scratch --- # useless-parameters A **114,838,272-parameter** decoder-only language model pre-trained **from scratch** on [fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu), using the [HyperDex Trainer](https://hugging-science-hyperdex-trainer.hf.space/) Space. ## Architecture A standard `LlamaForCausalLM` decoder-only transformer — SiLU MLP, RMSNorm, rotary position embeddings, grouped-query attention, tied embeddings, no biases — scaled down in width and depth to fit the parameter budget. | | | |---|---| | Parameters | 114,838,272 | | Hidden size | 768 | | Layers | 12 | | Attention heads | 12 (KV: 12) | | FFN size | 3072 | | Context length | 512 | | Vocab | 2,048 (custom BPE trained on fineweb-edu) | ## Training | | | |---|---| | Tokens seen | 524,288 | | Steps | 1 | | Tokens / step | 524,288 | | Optimizer | AdamW(0.9, 0.95) wd=0.1 clip=1.0 | | LR schedule | warmup 2% + cosine to 10% (peak 3e-04) | | Final loss | 8.2987 (ppl 4018.8) | | Wall time | 0.4 min | | Trained by | [@GGUFGuy](https://huggingface.co/GGUFGuy) | ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer tok = AutoTokenizer.from_pretrained("GGUFGuy/useless-parameters") model = AutoModelForCausalLM.from_pretrained("GGUFGuy/useless-parameters") ids = tok("The mitochondria is", return_tensors="pt").input_ids print(tok.decode(model.generate(ids, max_new_tokens=60, do_sample=True, temperature=0.8, top_k=50)[0])) ``` ## Caveats This is a **small-scale research artifact**. At this parameter count and token budget the model learns word shapes, common collocations and a little syntax — it is not a useful assistant and its output is not factual. It exists to make "pre-train a transformer from scratch" something you can actually watch happen.