--- language: - en - id - es - fr - de license: mit tags: - text-generation - llama - pytorch - terminal - small-language-model pipeline_tag: text-generation widget: - text: "User: How do I navigate up one directory?\nAssistant:" - text: "$ cd ..\n$" --- # 5M Terminal & Multilingual Language Model (Chinchilla Optimal - 100M Tokens) This repository contains a **5.0 Million Parameter Causal Language Model** trained from scratch on a compute-optimal token budget of **100 Million Tokens** ($20\times$ parameter count) following Chinchilla scaling laws. The model is specialized in **Linux Terminal Commands, Shell Automation, and Multilingual Text Generation** (English, Indonesian, Spanish, French, German). ## Model Architecture * **Total Trainable Parameters**: 4,984,064 (~4.98M / 5.0M) * **Architecture**: Decoder-only Transformer (LLaMA style) * **Vocabulary Size**: 4,096 (Byte-Pair Encoding Tokenizer) * **Hidden Dimension (`d_model`)**: 256 * **Number of Layers (`n_layer`)**: 6 * **Number of Attention Heads (`n_head`)**: 8 (Head dimension = 32) * **MLP Hidden Dimension (`inter_dim`)**: 512 (SwiGLU activation) * **Context Window (`max_seq_len`)**: 256 tokens * **Weight Tying**: Tied Embedding and LM Head weights ## Training Details * **Token Budget**: 100,000,000 Tokens (Chinchilla Optimal: 20 tokens per parameter) * **Dataset Composition**: * **Terminal CLI & Commands**: 20% (Bash commands, `cd ..`, `ls -la`, `mkdir`, `grep`, `git`, `chmod`, `curl`, Q&A pairs) * **English Text**: 24% (`wikitext-103-v1` + technical corpus) * **Multilingual Text**: 56% (Indonesian, Spanish, French, German Wikipedia & general text) * **Hardware**: Kaggle NVIDIA GPU (Dual Tesla T4 / P100) with FP16 AMP Mixed Precision. ## Usage & Inference Example (PyTorch) ```python import torch from tokenizers import Tokenizer from huggingface_hub import hf_hub_download # Download model assets repo_id = "kipasyangin5/5m-terminal-lm-chinchilla" tok_path = hf_hub_download(repo_id=repo_id, filename="tokenizer.json") weights_path = hf_hub_download(repo_id=repo_id, filename="model.pt") # Load Tokenizer tokenizer = Tokenizer.from_file(tok_path) # Prompt execution prompt = "User: How do I list all files including hidden ones?\nAssistant:" tokens = tokenizer.encode(prompt).ids input_ids = torch.tensor([tokens], dtype=torch.long) print("Input prompt:", prompt) ``` ## Citation & License MIT License. Developed by `kipasyangin5`.