DedeProGames commited on
Commit
df552ba
·
verified ·
1 Parent(s): 52811a8

NanoDex 8m · 199,753,728 fineweb-edu tokens · loss 3.8884

Browse files
README.md ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: odc-by
3
+ datasets:
4
+ - HuggingFaceFW/fineweb-edu
5
+ language:
6
+ - en
7
+ library_name: transformers
8
+ pipeline_tag: text-generation
9
+ tags:
10
+ - nanodex
11
+ - tiny-lm
12
+ - pretrained-from-scratch
13
+ ---
14
+
15
+ # LowOnMind-8M
16
+
17
+ A **8,060,256-parameter** decoder-only language model pre-trained
18
+ **from scratch** on [fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu),
19
+ using the [NanoDex Trainer](https://huggingface.co/spaces/hugging-science/nanodex-trainer) Space.
20
+
21
+ ## Architecture
22
+
23
+ A standard `LlamaForCausalLM` decoder-only transformer — SiLU MLP, RMSNorm,
24
+ rotary position embeddings, grouped-query attention, tied embeddings, no biases —
25
+ scaled down in width and depth to fit the parameter budget.
26
+
27
+ | | |
28
+ |---|---|
29
+ | Parameters | 8,060,256 |
30
+ | Hidden size | 288 |
31
+ | Layers | 9 |
32
+ | Attention heads | 9 (KV: 3) |
33
+ | FFN size | 704 |
34
+ | Context length | 512 |
35
+ | Vocab | 2,048 (custom BPE trained on fineweb-edu) |
36
+
37
+ ## Training
38
+
39
+ | | |
40
+ |---|---|
41
+ | Tokens seen | 199,753,728 |
42
+ | Steps | 381 |
43
+ | Tokens / step | 524,288 |
44
+ | Optimizer | AdamW(0.9, 0.95) wd=0.1 clip=1.0 |
45
+ | LR schedule | warmup 2% + cosine to 10% (peak 1e-03) |
46
+ | Final loss | 3.8884 (ppl 48.8) |
47
+ | Wall time | 29.3 min |
48
+ | Trained by | [@DedeProGames](https://huggingface.co/DedeProGames) |
49
+
50
+ ## Usage
51
+
52
+ ```python
53
+ from transformers import AutoModelForCausalLM, AutoTokenizer
54
+
55
+ tok = AutoTokenizer.from_pretrained("DedeProGames/LowOnMind-8M")
56
+ model = AutoModelForCausalLM.from_pretrained("DedeProGames/LowOnMind-8M")
57
+
58
+ ids = tok("The mitochondria is", return_tensors="pt").input_ids
59
+ print(tok.decode(model.generate(ids, max_new_tokens=60, do_sample=True,
60
+ temperature=0.8, top_k=50)[0]))
61
+ ```
62
+
63
+ ## Caveats
64
+
65
+ This is a **nano-scale research artifact**. At this parameter count and token
66
+ budget the model learns word shapes, common collocations and a little syntax —
67
+ it is not a useful assistant and its output is not factual. It exists to make
68
+ "pre-train a transformer from scratch" something you can actually watch happen.
config.json ADDED
@@ -0,0 +1,32 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "LlamaForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 0,
8
+ "dtype": "float32",
9
+ "eos_token_id": 0,
10
+ "head_dim": 32,
11
+ "hidden_act": "silu",
12
+ "hidden_size": 288,
13
+ "initializer_range": 0.041666666666666664,
14
+ "intermediate_size": 704,
15
+ "max_position_embeddings": 512,
16
+ "mlp_bias": false,
17
+ "model_type": "llama",
18
+ "num_attention_heads": 9,
19
+ "num_hidden_layers": 9,
20
+ "num_key_value_heads": 3,
21
+ "pad_token_id": 1,
22
+ "pretraining_tp": 1,
23
+ "rms_norm_eps": 1e-05,
24
+ "rope_parameters": {
25
+ "rope_theta": 10000.0,
26
+ "rope_type": "default"
27
+ },
28
+ "tie_word_embeddings": true,
29
+ "transformers_version": "5.17.0",
30
+ "use_cache": true,
31
+ "vocab_size": 2048
32
+ }
generation_config.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 0,
3
+ "eos_token_id": 0,
4
+ "pad_token_id": 1,
5
+ "_from_model_config": true
6
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ba1df7576bd7a53279bf6effcb4b621fa211f36ac3be342a669e9606989c1a85
3
+ size 32249984
special_tokens_map.json ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ {
2
+ "bos_token": "<|endoftext|>",
3
+ "eos_token": "<|endoftext|>",
4
+ "pad_token": "<|pad|>"
5
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "added_tokens_decoder": {
3
+ "0": {
4
+ "content": "<|endoftext|>",
5
+ "lstrip": false,
6
+ "normalized": false,
7
+ "rstrip": false,
8
+ "single_word": false,
9
+ "special": true
10
+ },
11
+ "1": {
12
+ "content": "<|pad|>",
13
+ "lstrip": false,
14
+ "normalized": false,
15
+ "rstrip": false,
16
+ "single_word": false,
17
+ "special": true
18
+ }
19
+ },
20
+ "bos_token": "<|endoftext|>",
21
+ "eos_token": "<|endoftext|>",
22
+ "pad_token": "<|pad|>",
23
+ "unk_token": null,
24
+ "clean_up_tokenization_spaces": false,
25
+ "model_max_length": 512,
26
+ "tokenizer_class": "PreTrainedTokenizerFast"
27
+ }
training_run.json ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "job_id": "e6414655b517",
3
+ "tier": "8m",
4
+ "label": "NanoDex-8M",
5
+ "n_params": 8060256,
6
+ "target_tokens": 200000000,
7
+ "tokens_seen": 199753728,
8
+ "steps": 381,
9
+ "seq_len": 512,
10
+ "final_loss": 3.8884217739105225,
11
+ "best_loss": 3.820189207792282,
12
+ "peak_lr": 0.00133,
13
+ "batch_tokens": 524288,
14
+ "optimizer": "AdamW(0.9, 0.95) wd=0.1 clip=1.0",
15
+ "schedule": "warmup 2% + cosine to 10%",
16
+ "dataset": "HuggingFaceFW/fineweb-edu (sample-10BT)",
17
+ "architecture": "LlamaForCausalLM (SiLU, RMSNorm, RoPE, GQA, tied embeddings)",
18
+ "wall_time_s": 1755.7315411567688,
19
+ "trained_by": "DedeProGames"
20
+ }