frankenmoe / README.md
hotdogs's picture
Upload README.md with huggingface_hub
8c6e4c1 verified
|
Raw History Blame
3.92 kB
metadata
language:
  - en
  - th
license: apache-2.0
tags:
  - frankenmoe
  - qwen2.5
  - lora
  - peft
  - gguf
  - coding
  - math
  - chat
  - expert-models
pipeline_tag: text-generation
base_model: Qwen/Qwen2.5-1.5B-Instruct

FrankenMoE β€” Qwen2.5-1.5B Expert Models πŸ‡ΉπŸ‡­

3 specialized LoRA fine-tuned experts β€” coding, math, and chat β€” built from Qwen2.5-1.5B-Instruct with 13,000 curated training samples.

πŸ›‘ MoE merge skipped (mergekit does not support Qwen2 MoE architecture).
βœ… Each expert is independently usable as LoRA adapter or GGUF.


πŸ“¦ What's Inside

Expert Domain LoRA GGUF (Q4_K_M) Train Loss Eval Loss
coding Python/Algorithm/SWE 71 MB 941 MB 1.03 -
math Mathematics/Proofs 70 MB 941 MB 1.18 -
chat Instruction Following 74 MB 941 MB 1.23 1.27

πŸš€ Quick Start

Option 1: LoRA with PEFT (Python)

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

base = "Qwen/Qwen2.5-1.5B-Instruct"
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype=torch.bfloat16)
model = PeftModel.from_pretrained(model, "hotdogs/frankenmoe", subfolder="coding")
tokenizer = AutoTokenizer.from_pretrained("hotdogs/frankenmoe", subfolder="coding")

prompt = "Write a Python function to reverse a linked list"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0]))

Option 2: GGUF with llama.cpp

# Download
wget https://huggingface.co/hotdogs/frankenmoe/resolve/main/coding/frankenmoe_coding-Q4_K_M.gguf

# Run
llama.cpp/build/bin/llama-cli \
  -m frankenmoe_coding-Q4_K_M.gguf \
  -p "Write a Python function to reverse a linked list" \
  -n 256

Option 3: Ollama Modelfile

FROM ./frankenmoe_coding-Q4_K_M.gguf
SYSTEM "You are a coding expert specialized in Python, algorithms, and software engineering."

πŸ”§ Training Details

Parameter Value
Base Model Qwen2.5-1.5B-Instruct
Method LoRA (r=16, alpha=32)
Precision bfloat16 (no 4-bit quantization)
Dataset 13,000 curated samples (coding: 5K, math: 3K, chat: 5K)
Epochs 2 per expert
GPU RTX 4060 Ti 16GB
Framework transformers + peft + trl
Optimizer AdamW (torch)
NEFTune Ξ± 5-7

πŸ“ Repository Structure

hotdogs/frankenmoe/
β”œβ”€β”€ README.md
β”œβ”€β”€ coding/
β”‚   β”œβ”€β”€ adapter_model.safetensors
β”‚   β”œβ”€β”€ adapter_config.json
β”‚   β”œβ”€β”€ tokenizer.json
β”‚   β”œβ”€β”€ tokenizer_config.json
β”‚   └── frankenmoe_coding-Q4_K_M.gguf
β”œβ”€β”€ math/
β”‚   β”œβ”€β”€ adapter_model.safetensors
β”‚   β”œβ”€β”€ adapter_config.json
β”‚   β”œβ”€β”€ tokenizer.json
β”‚   β”œβ”€β”€ tokenizer_config.json
β”‚   └── frankenmoe_math-Q4_K_M.gguf
└── chat/
    β”œβ”€β”€ adapter_model.safetensors
    β”œβ”€β”€ adapter_config.json
    β”œβ”€β”€ tokenizer.json
    β”œβ”€β”€ tokenizer_config.json
    └── frankenmoe_chat-Q4_K_M.gguf

πŸ“Š Training Logs

Expert Steps Train Loss Final LR Time
coding 564 1.03 - ~15 min
math 338 1.18 - ~15 min
chat 564 1.23 - ~28 min

⚠️ Known Limitations

  • No MoE routing β€” experts are independent models, not a single MoE
  • Small base model (1.5B) β€” good for experimentation, limited for production
  • Qwen2 architecture β€” not compatible with mergekit MoE (only Qwen3 MoE supported)

πŸ“œ License

Same as base model: Apache 2.0


πŸ™ Credits

Trained by UKA (AI Agent) on FrankenMoE Pipeline v2.0
Thai AI infrastructure β€” local GPU only, zero cloud dependency πŸ‡ΉπŸ‡­