File size: 2,443 Bytes
b3b4eef
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
---
language:
- en
- id
- es
- fr
- de
license: mit
tags:
- text-generation
- llama
- pytorch
- terminal
- small-language-model
pipeline_tag: text-generation
widget:
- text: "User: How do I navigate up one directory?\nAssistant:"
- text: "$ cd ..\n$"
---

# 5M Terminal & Multilingual Language Model (Chinchilla Optimal - 100M Tokens)

This repository contains a **5.0 Million Parameter Causal Language Model** trained from scratch on a compute-optimal token budget of **100 Million Tokens** ($20\times$ parameter count) following Chinchilla scaling laws.

The model is specialized in **Linux Terminal Commands, Shell Automation, and Multilingual Text Generation** (English, Indonesian, Spanish, French, German).

## Model Architecture

* **Total Trainable Parameters**: 4,984,064 (~4.98M / 5.0M)
* **Architecture**: Decoder-only Transformer (LLaMA style)
* **Vocabulary Size**: 4,096 (Byte-Pair Encoding Tokenizer)
* **Hidden Dimension (`d_model`)**: 256
* **Number of Layers (`n_layer`)**: 6
* **Number of Attention Heads (`n_head`)**: 8 (Head dimension = 32)
* **MLP Hidden Dimension (`inter_dim`)**: 512 (SwiGLU activation)
* **Context Window (`max_seq_len`)**: 256 tokens
* **Weight Tying**: Tied Embedding and LM Head weights

## Training Details

* **Token Budget**: 100,000,000 Tokens (Chinchilla Optimal: 20 tokens per parameter)
* **Dataset Composition**:
  * **Terminal CLI & Commands**: 20% (Bash commands, `cd ..`, `ls -la`, `mkdir`, `grep`, `git`, `chmod`, `curl`, Q&A pairs)
  * **English Text**: 24% (`wikitext-103-v1` + technical corpus)
  * **Multilingual Text**: 56% (Indonesian, Spanish, French, German Wikipedia & general text)
* **Hardware**: Kaggle NVIDIA GPU (Dual Tesla T4 / P100) with FP16 AMP Mixed Precision.

## Usage & Inference Example (PyTorch)

```python
import torch
from tokenizers import Tokenizer
from huggingface_hub import hf_hub_download

# Download model assets
repo_id = "kipasyangin5/5m-terminal-lm-chinchilla"
tok_path = hf_hub_download(repo_id=repo_id, filename="tokenizer.json")
weights_path = hf_hub_download(repo_id=repo_id, filename="model.pt")

# Load Tokenizer
tokenizer = Tokenizer.from_file(tok_path)

# Prompt execution
prompt = "User: How do I list all files including hidden ones?\nAssistant:"
tokens = tokenizer.encode(prompt).ids
input_ids = torch.tensor([tokens], dtype=torch.long)

print("Input prompt:", prompt)
```

## Citation & License

MIT License. Developed by `kipasyangin5`.