File size: 3,924 Bytes
8c6e4c1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a4d637f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
---
language:
  - en
  - th
license: apache-2.0
tags:
  - frankenmoe
  - qwen2.5
  - lora
  - peft
  - gguf
  - coding
  - math
  - chat
  - expert-models
pipeline_tag: text-generation
base_model: Qwen/Qwen2.5-1.5B-Instruct
---

# FrankenMoE β€” Qwen2.5-1.5B Expert Models πŸ‡ΉπŸ‡­

**3 specialized LoRA fine-tuned experts** β€” coding, math, and chat β€” built from Qwen2.5-1.5B-Instruct with 13,000 curated training samples.

> πŸ›‘ MoE merge skipped (mergekit does not support Qwen2 MoE architecture).  
> βœ… Each expert is independently usable as LoRA adapter or GGUF.

---

## πŸ“¦ What's Inside

| Expert | Domain | LoRA | GGUF (Q4_K_M) | Train Loss | Eval Loss |
|--------|--------|------|---------------|------------|-----------|
| **coding** | Python/Algorithm/SWE | 71 MB | 941 MB | 1.03 | - |
| **math** | Mathematics/Proofs | 70 MB | 941 MB | 1.18 | - |
| **chat** | Instruction Following | 74 MB | 941 MB | 1.23 | 1.27 |

---

## πŸš€ Quick Start

### Option 1: LoRA with PEFT (Python)

```python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

base = "Qwen/Qwen2.5-1.5B-Instruct"
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype=torch.bfloat16)
model = PeftModel.from_pretrained(model, "hotdogs/frankenmoe", subfolder="coding")
tokenizer = AutoTokenizer.from_pretrained("hotdogs/frankenmoe", subfolder="coding")

prompt = "Write a Python function to reverse a linked list"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0]))
```

### Option 2: GGUF with llama.cpp

```bash
# Download
wget https://huggingface.co/hotdogs/frankenmoe/resolve/main/coding/frankenmoe_coding-Q4_K_M.gguf

# Run
llama.cpp/build/bin/llama-cli \
  -m frankenmoe_coding-Q4_K_M.gguf \
  -p "Write a Python function to reverse a linked list" \
  -n 256
```

### Option 3: Ollama Modelfile

```dockerfile
FROM ./frankenmoe_coding-Q4_K_M.gguf
SYSTEM "You are a coding expert specialized in Python, algorithms, and software engineering."
```

---

## πŸ”§ Training Details

| Parameter | Value |
|-----------|-------|
| Base Model | Qwen2.5-1.5B-Instruct |
| Method | LoRA (r=16, alpha=32) |
| Precision | bfloat16 (no 4-bit quantization) |
| Dataset | 13,000 curated samples (coding: 5K, math: 3K, chat: 5K) |
| Epochs | 2 per expert |
| GPU | RTX 4060 Ti 16GB |
| Framework | transformers + peft + trl |
| Optimizer | AdamW (torch) |
| NEFTune Ξ± | 5-7 |

---

## πŸ“ Repository Structure

```
hotdogs/frankenmoe/
β”œβ”€β”€ README.md
β”œβ”€β”€ coding/
β”‚   β”œβ”€β”€ adapter_model.safetensors
β”‚   β”œβ”€β”€ adapter_config.json
β”‚   β”œβ”€β”€ tokenizer.json
β”‚   β”œβ”€β”€ tokenizer_config.json
β”‚   └── frankenmoe_coding-Q4_K_M.gguf
β”œβ”€β”€ math/
β”‚   β”œβ”€β”€ adapter_model.safetensors
β”‚   β”œβ”€β”€ adapter_config.json
β”‚   β”œβ”€β”€ tokenizer.json
β”‚   β”œβ”€β”€ tokenizer_config.json
β”‚   └── frankenmoe_math-Q4_K_M.gguf
└── chat/
    β”œβ”€β”€ adapter_model.safetensors
    β”œβ”€β”€ adapter_config.json
    β”œβ”€β”€ tokenizer.json
    β”œβ”€β”€ tokenizer_config.json
    └── frankenmoe_chat-Q4_K_M.gguf
```

---

## πŸ“Š Training Logs

| Expert | Steps | Train Loss | Final LR | Time |
|--------|-------|-----------|----------|------|
| coding | 564 | 1.03 | - | ~15 min |
| math | 338 | 1.18 | - | ~15 min |
| chat | 564 | 1.23 | - | ~28 min |

---

## ⚠️ Known Limitations

- **No MoE routing** β€” experts are independent models, not a single MoE
- **Small base model** (1.5B) β€” good for experimentation, limited for production
- **Qwen2 architecture** β€” not compatible with mergekit MoE (only Qwen3 MoE supported)

---

## πŸ“œ License

Same as base model: Apache 2.0

---

## πŸ™ Credits

Trained by **UKA** (AI Agent) on FrankenMoE Pipeline v2.0  
Thai AI infrastructure β€” local GPU only, zero cloud dependency πŸ‡ΉπŸ‡­