|
Download README.md from kipasyangin5/5m-terminal-lm-chinchilla: direct link, hf CLI and curl.
- Browser
- Download file 2.44 kB
-
https://huggingface.co/kipasyangin5/5m-terminal-lm-chinchilla/resolve/b3b4eef66f9c19c26b92edb3caaa94f2c91e55b1/README.md
- Command line
-
hf download hf://kipasyangin5/5m-terminal-lm-chinchilla@b3b4eef66f9c19c26b92edb3caaa94f2c91e55b1/README.md
-
curl -L -o README.md https://huggingface.co/kipasyangin5/5m-terminal-lm-chinchilla/resolve/b3b4eef66f9c19c26b92edb3caaa94f2c91e55b1/README.md
2.44 kB
metadata
language:
- en
- id
- es
- fr
- de
license: mit
tags:
- text-generation
- llama
- pytorch
- terminal
- small-language-model
pipeline_tag: text-generation
widget:
- text: |-
User: How do I navigate up one directory?
Assistant:
- text: |-
$ cd ..
$
5M Terminal & Multilingual Language Model (Chinchilla Optimal - 100M Tokens)
This repository contains a 5.0 Million Parameter Causal Language Model trained from scratch on a compute-optimal token budget of 100 Million Tokens ($20\times$ parameter count) following Chinchilla scaling laws.
The model is specialized in Linux Terminal Commands, Shell Automation, and Multilingual Text Generation (English, Indonesian, Spanish, French, German).
Model Architecture
- Total Trainable Parameters: 4,984,064 (~4.98M / 5.0M)
- Architecture: Decoder-only Transformer (LLaMA style)
- Vocabulary Size: 4,096 (Byte-Pair Encoding Tokenizer)
- Hidden Dimension (
d_model): 256 - Number of Layers (
n_layer): 6 - Number of Attention Heads (
n_head): 8 (Head dimension = 32) - MLP Hidden Dimension (
inter_dim): 512 (SwiGLU activation) - Context Window (
max_seq_len): 256 tokens - Weight Tying: Tied Embedding and LM Head weights
Training Details
- Token Budget: 100,000,000 Tokens (Chinchilla Optimal: 20 tokens per parameter)
- Dataset Composition:
- Terminal CLI & Commands: 20% (Bash commands,
cd ..,ls -la,mkdir,grep,git,chmod,curl, Q&A pairs) - English Text: 24% (
wikitext-103-v1+ technical corpus) - Multilingual Text: 56% (Indonesian, Spanish, French, German Wikipedia & general text)
- Terminal CLI & Commands: 20% (Bash commands,
- Hardware: Kaggle NVIDIA GPU (Dual Tesla T4 / P100) with FP16 AMP Mixed Precision.
Usage & Inference Example (PyTorch)
import torch
from tokenizers import Tokenizer
from huggingface_hub import hf_hub_download
# Download model assets
repo_id = "kipasyangin5/5m-terminal-lm-chinchilla"
tok_path = hf_hub_download(repo_id=repo_id, filename="tokenizer.json")
weights_path = hf_hub_download(repo_id=repo_id, filename="model.pt")
# Load Tokenizer
tokenizer = Tokenizer.from_file(tok_path)
# Prompt execution
prompt = "User: How do I list all files including hidden ones?\nAssistant:"
tokens = tokenizer.encode(prompt).ids
input_ids = torch.tensor([tokens], dtype=torch.long)
print("Input prompt:", prompt)
Citation & License
MIT License. Developed by kipasyangin5.