kipasyangin5's picture
Add clean model card README.md
b3b4eef verified
|
Raw History Blame
2.44 kB
metadata
language:
  - en
  - id
  - es
  - fr
  - de
license: mit
tags:
  - text-generation
  - llama
  - pytorch
  - terminal
  - small-language-model
pipeline_tag: text-generation
widget:
  - text: |-
      User: How do I navigate up one directory?
      Assistant:
  - text: |-
      $ cd ..
      $

5M Terminal & Multilingual Language Model (Chinchilla Optimal - 100M Tokens)

This repository contains a 5.0 Million Parameter Causal Language Model trained from scratch on a compute-optimal token budget of 100 Million Tokens ($20\times$ parameter count) following Chinchilla scaling laws.

The model is specialized in Linux Terminal Commands, Shell Automation, and Multilingual Text Generation (English, Indonesian, Spanish, French, German).

Model Architecture

  • Total Trainable Parameters: 4,984,064 (~4.98M / 5.0M)
  • Architecture: Decoder-only Transformer (LLaMA style)
  • Vocabulary Size: 4,096 (Byte-Pair Encoding Tokenizer)
  • Hidden Dimension (d_model): 256
  • Number of Layers (n_layer): 6
  • Number of Attention Heads (n_head): 8 (Head dimension = 32)
  • MLP Hidden Dimension (inter_dim): 512 (SwiGLU activation)
  • Context Window (max_seq_len): 256 tokens
  • Weight Tying: Tied Embedding and LM Head weights

Training Details

  • Token Budget: 100,000,000 Tokens (Chinchilla Optimal: 20 tokens per parameter)
  • Dataset Composition:
    • Terminal CLI & Commands: 20% (Bash commands, cd .., ls -la, mkdir, grep, git, chmod, curl, Q&A pairs)
    • English Text: 24% (wikitext-103-v1 + technical corpus)
    • Multilingual Text: 56% (Indonesian, Spanish, French, German Wikipedia & general text)
  • Hardware: Kaggle NVIDIA GPU (Dual Tesla T4 / P100) with FP16 AMP Mixed Precision.

Usage & Inference Example (PyTorch)

import torch
from tokenizers import Tokenizer
from huggingface_hub import hf_hub_download

# Download model assets
repo_id = "kipasyangin5/5m-terminal-lm-chinchilla"
tok_path = hf_hub_download(repo_id=repo_id, filename="tokenizer.json")
weights_path = hf_hub_download(repo_id=repo_id, filename="model.pt")

# Load Tokenizer
tokenizer = Tokenizer.from_file(tok_path)

# Prompt execution
prompt = "User: How do I list all files including hidden ones?\nAssistant:"
tokens = tokenizer.encode(prompt).ids
input_ids = torch.tensor([tokens], dtype=torch.long)

print("Input prompt:", prompt)

Citation & License

MIT License. Developed by kipasyangin5.