|
Download README.md from kipasyangin5/5m-terminal-lm-chinchilla: direct link, hf CLI and curl.
- Browser
- Download file 2.44 kB
-
https://huggingface.co/kipasyangin5/5m-terminal-lm-chinchilla/resolve/b3b4eef66f9c19c26b92edb3caaa94f2c91e55b1/README.md
- Command line
-
hf download hf://kipasyangin5/5m-terminal-lm-chinchilla@b3b4eef66f9c19c26b92edb3caaa94f2c91e55b1/README.md
-
curl -L -o README.md https://huggingface.co/kipasyangin5/5m-terminal-lm-chinchilla/resolve/b3b4eef66f9c19c26b92edb3caaa94f2c91e55b1/README.md
2.44 kB
| language: | |
| - en | |
| - id | |
| - es | |
| - fr | |
| - de | |
| license: mit | |
| tags: | |
| - text-generation | |
| - llama | |
| - pytorch | |
| - terminal | |
| - small-language-model | |
| pipeline_tag: text-generation | |
| widget: | |
| - text: "User: How do I navigate up one directory?\nAssistant:" | |
| - text: "$ cd ..\n$" | |
| # 5M Terminal & Multilingual Language Model (Chinchilla Optimal - 100M Tokens) | |
| This repository contains a **5.0 Million Parameter Causal Language Model** trained from scratch on a compute-optimal token budget of **100 Million Tokens** ($20\times$ parameter count) following Chinchilla scaling laws. | |
| The model is specialized in **Linux Terminal Commands, Shell Automation, and Multilingual Text Generation** (English, Indonesian, Spanish, French, German). | |
| ## Model Architecture | |
| * **Total Trainable Parameters**: 4,984,064 (~4.98M / 5.0M) | |
| * **Architecture**: Decoder-only Transformer (LLaMA style) | |
| * **Vocabulary Size**: 4,096 (Byte-Pair Encoding Tokenizer) | |
| * **Hidden Dimension (`d_model`)**: 256 | |
| * **Number of Layers (`n_layer`)**: 6 | |
| * **Number of Attention Heads (`n_head`)**: 8 (Head dimension = 32) | |
| * **MLP Hidden Dimension (`inter_dim`)**: 512 (SwiGLU activation) | |
| * **Context Window (`max_seq_len`)**: 256 tokens | |
| * **Weight Tying**: Tied Embedding and LM Head weights | |
| ## Training Details | |
| * **Token Budget**: 100,000,000 Tokens (Chinchilla Optimal: 20 tokens per parameter) | |
| * **Dataset Composition**: | |
| * **Terminal CLI & Commands**: 20% (Bash commands, `cd ..`, `ls -la`, `mkdir`, `grep`, `git`, `chmod`, `curl`, Q&A pairs) | |
| * **English Text**: 24% (`wikitext-103-v1` + technical corpus) | |
| * **Multilingual Text**: 56% (Indonesian, Spanish, French, German Wikipedia & general text) | |
| * **Hardware**: Kaggle NVIDIA GPU (Dual Tesla T4 / P100) with FP16 AMP Mixed Precision. | |
| ## Usage & Inference Example (PyTorch) | |
| ```python | |
| import torch | |
| from tokenizers import Tokenizer | |
| from huggingface_hub import hf_hub_download | |
| # Download model assets | |
| repo_id = "kipasyangin5/5m-terminal-lm-chinchilla" | |
| tok_path = hf_hub_download(repo_id=repo_id, filename="tokenizer.json") | |
| weights_path = hf_hub_download(repo_id=repo_id, filename="model.pt") | |
| # Load Tokenizer | |
| tokenizer = Tokenizer.from_file(tok_path) | |
| # Prompt execution | |
| prompt = "User: How do I list all files including hidden ones?\nAssistant:" | |
| tokens = tokenizer.encode(prompt).ids | |
| input_ids = torch.tensor([tokens], dtype=torch.long) | |
| print("Input prompt:", prompt) | |
| ``` | |
| ## Citation & License | |
| MIT License. Developed by `kipasyangin5`. | |