kipasyangin5 commited on
Commit
b3b4eef
·
verified ·
1 Parent(s): cba7fe0

Add clean model card README.md

Browse files
Files changed (1) hide show
  1. README.md +73 -0
README.md ADDED
@@ -0,0 +1,73 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ - id
5
+ - es
6
+ - fr
7
+ - de
8
+ license: mit
9
+ tags:
10
+ - text-generation
11
+ - llama
12
+ - pytorch
13
+ - terminal
14
+ - small-language-model
15
+ pipeline_tag: text-generation
16
+ widget:
17
+ - text: "User: How do I navigate up one directory?\nAssistant:"
18
+ - text: "$ cd ..\n$"
19
+ ---
20
+
21
+ # 5M Terminal & Multilingual Language Model (Chinchilla Optimal - 100M Tokens)
22
+
23
+ This repository contains a **5.0 Million Parameter Causal Language Model** trained from scratch on a compute-optimal token budget of **100 Million Tokens** ($20\times$ parameter count) following Chinchilla scaling laws.
24
+
25
+ The model is specialized in **Linux Terminal Commands, Shell Automation, and Multilingual Text Generation** (English, Indonesian, Spanish, French, German).
26
+
27
+ ## Model Architecture
28
+
29
+ * **Total Trainable Parameters**: 4,984,064 (~4.98M / 5.0M)
30
+ * **Architecture**: Decoder-only Transformer (LLaMA style)
31
+ * **Vocabulary Size**: 4,096 (Byte-Pair Encoding Tokenizer)
32
+ * **Hidden Dimension (`d_model`)**: 256
33
+ * **Number of Layers (`n_layer`)**: 6
34
+ * **Number of Attention Heads (`n_head`)**: 8 (Head dimension = 32)
35
+ * **MLP Hidden Dimension (`inter_dim`)**: 512 (SwiGLU activation)
36
+ * **Context Window (`max_seq_len`)**: 256 tokens
37
+ * **Weight Tying**: Tied Embedding and LM Head weights
38
+
39
+ ## Training Details
40
+
41
+ * **Token Budget**: 100,000,000 Tokens (Chinchilla Optimal: 20 tokens per parameter)
42
+ * **Dataset Composition**:
43
+ * **Terminal CLI & Commands**: 20% (Bash commands, `cd ..`, `ls -la`, `mkdir`, `grep`, `git`, `chmod`, `curl`, Q&A pairs)
44
+ * **English Text**: 24% (`wikitext-103-v1` + technical corpus)
45
+ * **Multilingual Text**: 56% (Indonesian, Spanish, French, German Wikipedia & general text)
46
+ * **Hardware**: Kaggle NVIDIA GPU (Dual Tesla T4 / P100) with FP16 AMP Mixed Precision.
47
+
48
+ ## Usage & Inference Example (PyTorch)
49
+
50
+ ```python
51
+ import torch
52
+ from tokenizers import Tokenizer
53
+ from huggingface_hub import hf_hub_download
54
+
55
+ # Download model assets
56
+ repo_id = "kipasyangin5/5m-terminal-lm-chinchilla"
57
+ tok_path = hf_hub_download(repo_id=repo_id, filename="tokenizer.json")
58
+ weights_path = hf_hub_download(repo_id=repo_id, filename="model.pt")
59
+
60
+ # Load Tokenizer
61
+ tokenizer = Tokenizer.from_file(tok_path)
62
+
63
+ # Prompt execution
64
+ prompt = "User: How do I list all files including hidden ones?\nAssistant:"
65
+ tokens = tokenizer.encode(prompt).ids
66
+ input_ids = torch.tensor([tokens], dtype=torch.long)
67
+
68
+ print("Input prompt:", prompt)
69
+ ```
70
+
71
+ ## Citation & License
72
+
73
+ MIT License. Developed by `kipasyangin5`.