--- language: - en - th license: apache-2.0 tags: - frankenmoe - qwen2.5 - lora - peft - gguf - coding - math - chat - expert-models pipeline_tag: text-generation base_model: Qwen/Qwen2.5-1.5B-Instruct --- # FrankenMoE — Qwen2.5-1.5B Expert Models 🇹🇭 **3 specialized LoRA fine-tuned experts** — coding, math, and chat — built from Qwen2.5-1.5B-Instruct with 13,000 curated training samples. > 🛑 MoE merge skipped (mergekit does not support Qwen2 MoE architecture). > ✅ Each expert is independently usable as LoRA adapter or GGUF. --- ## 📦 What's Inside | Expert | Domain | LoRA | GGUF (Q4_K_M) | Train Loss | Eval Loss | |--------|--------|------|---------------|------------|-----------| | **coding** | Python/Algorithm/SWE | 71 MB | 941 MB | 1.03 | - | | **math** | Mathematics/Proofs | 70 MB | 941 MB | 1.18 | - | | **chat** | Instruction Following | 74 MB | 941 MB | 1.23 | 1.27 | --- ## 🚀 Quick Start ### Option 1: LoRA with PEFT (Python) ```python from peft import PeftModel from transformers import AutoModelForCausalLM, AutoTokenizer import torch base = "Qwen/Qwen2.5-1.5B-Instruct" model = AutoModelForCausalLM.from_pretrained(base, torch_dtype=torch.bfloat16) model = PeftModel.from_pretrained(model, "hotdogs/frankenmoe", subfolder="coding") tokenizer = AutoTokenizer.from_pretrained("hotdogs/frankenmoe", subfolder="coding") prompt = "Write a Python function to reverse a linked list" inputs = tokenizer(prompt, return_tensors="pt") outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0])) ``` ### Option 2: GGUF with llama.cpp ```bash # Download wget https://huggingface.co/hotdogs/frankenmoe/resolve/main/coding/frankenmoe_coding-Q4_K_M.gguf # Run llama.cpp/build/bin/llama-cli \ -m frankenmoe_coding-Q4_K_M.gguf \ -p "Write a Python function to reverse a linked list" \ -n 256 ``` ### Option 3: Ollama Modelfile ```dockerfile FROM ./frankenmoe_coding-Q4_K_M.gguf SYSTEM "You are a coding expert specialized in Python, algorithms, and software engineering." ``` --- ## 🔧 Training Details | Parameter | Value | |-----------|-------| | Base Model | Qwen2.5-1.5B-Instruct | | Method | LoRA (r=16, alpha=32) | | Precision | bfloat16 (no 4-bit quantization) | | Dataset | 13,000 curated samples (coding: 5K, math: 3K, chat: 5K) | | Epochs | 2 per expert | | GPU | RTX 4060 Ti 16GB | | Framework | transformers + peft + trl | | Optimizer | AdamW (torch) | | NEFTune α | 5-7 | --- ## 📁 Repository Structure ``` hotdogs/frankenmoe/ ├── README.md ├── coding/ │ ├── adapter_model.safetensors │ ├── adapter_config.json │ ├── tokenizer.json │ ├── tokenizer_config.json │ └── frankenmoe_coding-Q4_K_M.gguf ├── math/ │ ├── adapter_model.safetensors │ ├── adapter_config.json │ ├── tokenizer.json │ ├── tokenizer_config.json │ └── frankenmoe_math-Q4_K_M.gguf └── chat/ ├── adapter_model.safetensors ├── adapter_config.json ├── tokenizer.json ├── tokenizer_config.json └── frankenmoe_chat-Q4_K_M.gguf ``` --- ## 📊 Training Logs | Expert | Steps | Train Loss | Final LR | Time | |--------|-------|-----------|----------|------| | coding | 564 | 1.03 | - | ~15 min | | math | 338 | 1.18 | - | ~15 min | | chat | 564 | 1.23 | - | ~28 min | --- ## ⚠️ Known Limitations - **No MoE routing** — experts are independent models, not a single MoE - **Small base model** (1.5B) — good for experimentation, limited for production - **Qwen2 architecture** — not compatible with mergekit MoE (only Qwen3 MoE supported) --- ## 📜 License Same as base model: Apache 2.0 --- ## 🙏 Credits Trained by **UKA** (AI Agent) on FrankenMoE Pipeline v2.0 Thai AI infrastructure — local GPU only, zero cloud dependency 🇹🇭