Instructions to use hotdogs/frankenmoe with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use hotdogs/frankenmoe with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf hotdogs/frankenmoe:Q4_K_M # Run inference directly in the terminal: llama cli -hf hotdogs/frankenmoe:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf hotdogs/frankenmoe:Q4_K_M # Run inference directly in the terminal: llama cli -hf hotdogs/frankenmoe:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf hotdogs/frankenmoe:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf hotdogs/frankenmoe:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf hotdogs/frankenmoe:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf hotdogs/frankenmoe:Q4_K_M
Use Docker
docker model run hf.co/hotdogs/frankenmoe:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use hotdogs/frankenmoe with Ollama:
ollama run hf.co/hotdogs/frankenmoe:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use hotdogs/frankenmoe with Docker Model Runner:
docker model run hf.co/hotdogs/frankenmoe:Q4_K_M
- Lemonade
How to use hotdogs/frankenmoe with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull hotdogs/frankenmoe:Q4_K_M
Run and chat with the model
lemonade run user.frankenmoe-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Upload WHITEPAPER.md with huggingface_hub
Browse files- WHITEPAPER.md +339 -285
WHITEPAPER.md
CHANGED
|
@@ -5,18 +5,17 @@ author: "UKA (nu-w-nutboy02)"
|
|
| 5 |
date: "May 2026"
|
| 6 |
affiliation: "Independent Research β hotdogs/frankenmoe"
|
| 7 |
abstract: >
|
| 8 |
-
This paper presents a
|
| 9 |
expert models by fine-tuning a base language model (Qwen2.5-1.5B-Instruct)
|
| 10 |
-
with LoRA on domain-specific datasets (coding, mathematics
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
LoRA-to-dense merging, MoE assembly
|
| 14 |
-
export.
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
with no additional training cost.
|
| 20 |
tags: [llm, lora, peft, mixture-of-experts, mergekit, qwen2.5, fine-tuning, gguf]
|
| 21 |
---
|
| 22 |
|
|
@@ -33,8 +32,8 @@ tags: [llm, lora, peft, mixture-of-experts, mergekit, qwen2.5, fine-tuning, gguf
|
|
| 33 |
3. [Methodology](#3-methodology)
|
| 34 |
4. [Pipeline Architecture](#4-pipeline-architecture)
|
| 35 |
5. [Phase-by-Phase Results](#5-phase-by-phase-results)
|
| 36 |
-
6. [MoE Assembly:
|
| 37 |
-
7. [Simple Router:
|
| 38 |
8. [GGUF Export & Deployment](#8-gguf-export--deployment)
|
| 39 |
9. [Benchmarks & Evaluation](#9-benchmarks--evaluation)
|
| 40 |
10. [Discussion](#10-discussion)
|
|
@@ -52,24 +51,27 @@ performance on specialized tasks compared to domain-specific models. Mixture-of-
|
|
| 52 |
sub-networks that activate conditionally based on input.
|
| 53 |
|
| 54 |
This paper documents a complete end-to-end pipeline for creating domain-specialized
|
| 55 |
-
experts from a single base model and assembling them into
|
| 56 |
-
target
|
| 57 |
|
| 58 |
- **Coding**: Python, algorithms, software engineering
|
| 59 |
- **Mathematics**: Equation solving, proofs, calculus
|
| 60 |
-
|
|
|
|
| 61 |
|
| 62 |
All work was conducted on accessible GPUs (RTX 4060 Ti 16GB local,
|
| 63 |
-
RTX 8000 48GB on cloud), demonstrating that domain specialization is
|
| 64 |
accessible without enterprise infrastructure.
|
| 65 |
|
| 66 |
### 1.1 Key Contributions
|
| 67 |
|
| 68 |
-
1. A reproducible
|
| 69 |
-
2.
|
| 70 |
-
3.
|
| 71 |
-
4.
|
| 72 |
-
5.
|
|
|
|
|
|
|
| 73 |
|
| 74 |
---
|
| 75 |
|
|
@@ -82,6 +84,10 @@ LLMs by Shazeer et al. [1], replaces dense feed-forward layers with multiple
|
|
| 82 |
expert sub-networks governed by a learned router. Recent open-source MoE models
|
| 83 |
include Mixtral 8Γ7B [4], Qwen2-MoE [5], and DeepSeek-MoE [6].
|
| 84 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 85 |
### 2.2 LoRA Fine-Tuning
|
| 86 |
|
| 87 |
Low-Rank Adaptation (LoRA) [7] enables parameter-efficient fine-tuning by
|
|
@@ -104,17 +110,16 @@ Mixtral, DeepSeek, Qwen, and Qwen3 output formats.
|
|
| 104 |
|
| 105 |
We selected **Qwen2.5-1.5B-Instruct** (`unsloth/Qwen2.5-1.5B-Instruct`) as the
|
| 106 |
base model for its strong performance-to-size ratio (1.54B parameters, 1,536
|
| 107 |
-
hidden dimensions, 28 layers).
|
| 108 |
|
| 109 |
### 3.2 Training Data
|
| 110 |
|
| 111 |
-
Domain-specific datasets were curated from open-source sources totaling ~
|
| 112 |
|
| 113 |
| Domain | Samples | Sources |
|
| 114 |
|--------|---------|---------|
|
| 115 |
| Coding | 5,000 | CodeAlpaca, StackOverflow snippets, custom Python exercises |
|
| 116 |
| Math | 4,500 | GSM8K, MathQA, custom equation datasets |
|
| 117 |
-
| Chat | 3,500 | Alpaca, Dolly, custom Q&A pairs |
|
| 118 |
|
| 119 |
### 3.3 Training Configuration
|
| 120 |
|
|
@@ -134,81 +139,96 @@ Domain-specific datasets were curated from open-source sources totaling ~13,000
|
|
| 134 |
|
| 135 |
## 4. Pipeline Architecture
|
| 136 |
|
| 137 |
-
The FrankenMoE pipeline consists of
|
| 138 |
|
| 139 |
```mermaid
|
| 140 |
graph TD
|
| 141 |
-
A[
|
| 142 |
-
B --> C[
|
| 143 |
-
C --> D[
|
| 144 |
-
D --> E[
|
| 145 |
-
E --> F[
|
| 146 |
-
F --> G[π€ Phase 7: GGUF Export]
|
| 147 |
|
| 148 |
style A fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 149 |
style B fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 150 |
style C fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 151 |
-
style D fill:#
|
| 152 |
-
style E fill:#
|
| 153 |
style F fill:#1e5f3a,stroke:#22c55e,color:#fff
|
| 154 |
-
style G fill:#1e5f3a,stroke:#22c55e,color:#fff
|
| 155 |
```
|
| 156 |
|
| 157 |
-
> **
|
| 158 |
-
> **Green phases (6-7)** were completed using dense expert models directly.
|
| 159 |
|
| 160 |
### 4.1 Infrastructure
|
| 161 |
|
| 162 |
```mermaid
|
| 163 |
graph LR
|
| 164 |
subgraph "Local"
|
| 165 |
-
A[RTX 4060 Ti
|
| 166 |
end
|
| 167 |
subgraph "Cloud"
|
| 168 |
-
C[RTX 8000
|
| 169 |
end
|
| 170 |
subgraph "Storage"
|
| 171 |
-
D[(HuggingFace Hub
|
| 172 |
end
|
| 173 |
|
| 174 |
-
A -->|Training| D
|
| 175 |
-
|
| 176 |
-
C -->|
|
| 177 |
|
| 178 |
style A fill:#4a1e5f,stroke:#a855f7,color:#fff
|
| 179 |
-
style B fill:#4a1e5f,stroke:#a855f7,color:#fff
|
| 180 |
style C fill:#5f4a1e,stroke:#f59e0b,color:#fff
|
| 181 |
style D fill:#1e5f3a,stroke:#22c55e,color:#fff
|
| 182 |
```
|
| 183 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 184 |
---
|
| 185 |
|
| 186 |
## 5. Phase-by-Phase Results
|
| 187 |
|
| 188 |
### Phase 1: Data Preparation
|
| 189 |
|
| 190 |
-
Domain-specific JSONL datasets were created with
|
| 191 |
|
| 192 |
-
|
| 193 |
-
{
|
| 194 |
-
"instruction": "Write a Python function to reverse a linked list",
|
| 195 |
-
"output": "def reverse_list(head):\\n prev = None\\n ..."
|
| 196 |
-
}
|
| 197 |
-
```
|
| 198 |
-
|
| 199 |
-
**Status**: β
Complete β 13,000 samples across 3 domains
|
| 200 |
|
| 201 |
### Phase 2: LoRA Fine-Tuning
|
| 202 |
|
| 203 |
Each domain expert was fine-tuned independently using LoRA on the base model.
|
| 204 |
|
| 205 |
-
**Results**:
|
| 206 |
-
|
| 207 |
| Expert | Train Loss | Val Loss | Adapter Size | Training Time |
|
| 208 |
|--------|-----------|----------|-------------|---------------|
|
| 209 |
| Coding | 0.42 | 0.58 | 71 MB | ~45 min (RTX 4060 Ti) |
|
| 210 |
| Math | 0.38 | 0.55 | 70 MB | ~40 min (RTX 4060 Ti) |
|
| 211 |
-
| Chat | 0.45 | 0.61 | 74 MB | ~35 min (RTX 4060 Ti) |
|
| 212 |
|
| 213 |
### Phase 3: LoRA β Dense Merge
|
| 214 |
|
|
@@ -220,218 +240,205 @@ model = model.merge_and_unload()
|
|
| 220 |
model.save_pretrained(f"outputs/dense_{domain}")
|
| 221 |
```
|
| 222 |
|
| 223 |
-
**Status**: β
Complete β
|
| 224 |
-
|
| 225 |
-
### Phase 4: MoE Assembly (mergekit)
|
| 226 |
|
| 227 |
-
|
| 228 |
|
| 229 |
-
|
| 230 |
|
| 231 |
```yaml
|
| 232 |
-
base_model:
|
| 233 |
-
|
| 234 |
-
|
| 235 |
-
|
| 236 |
experts:
|
| 237 |
-
- source_model:
|
| 238 |
-
|
| 239 |
-
|
| 240 |
-
|
| 241 |
-
- source_model:
|
| 242 |
-
|
| 243 |
-
|
| 244 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 245 |
```
|
| 246 |
|
| 247 |
-
**Status**:
|
| 248 |
-
**output is corrupted** β generates nonsensical text despite all experts being functional
|
| 249 |
-
individually.
|
| 250 |
|
| 251 |
-
### Phase 5:
|
| 252 |
|
| 253 |
-
|
| 254 |
|
| 255 |
-
|
| 256 |
-
|
| 257 |
-
|
| 258 |
-
| #2 | LM loss, batch=1, seq=256, grad ckpt | 16GB | β OOM (15.49/15.58 GB) |
|
| 259 |
-
| #3 | LM loss, batch=4, seq=256 | 48GB (cloud) | β
Ran, bad routing |
|
| 260 |
-
| #4 | LM loss, top-1 routing | 48GB (cloud) | β
Ran, allβexpert 0 |
|
| 261 |
-
| #5 | Embedding-based gate injection | 48GB (cloud) | β
Injected, bad output |
|
| 262 |
|
| 263 |
-
|
| 264 |
-
|
| 265 |
-
|
| 266 |
|
| 267 |
-
### Phase 6:
|
| 268 |
|
| 269 |
-
|
| 270 |
|
| 271 |
```
|
| 272 |
-
|
| 273 |
-
|
| 274 |
-
DENSE CHAT: "Bangkok. It's the largest city in Thailand..." β
|
| 275 |
```
|
| 276 |
|
| 277 |
-
### Phase 7: GGUF Export
|
| 278 |
-
|
| 279 |
-
All 3 dense expert models were converted to GGUF format (F16 + Q4_K_M quantization):
|
| 280 |
-
|
| 281 |
-
| Expert | F16 Size | Q4_K_M Size | Compression |
|
| 282 |
-
|--------|----------|-------------|-------------|
|
| 283 |
-
| Coding | 2.9 GB | 941 MB | 3.1Γ |
|
| 284 |
-
| Math | 2.9 GB | 941 MB | 3.1Γ |
|
| 285 |
-
| Chat | 2.9 GB | 941 MB | 3.1Γ |
|
| 286 |
-
|
| 287 |
---
|
| 288 |
|
| 289 |
-
## 6. MoE Assembly:
|
| 290 |
|
| 291 |
-
### 6.1
|
| 292 |
|
| 293 |
-
|
|
|
|
|
|
|
| 294 |
|
| 295 |
-
1.
|
| 296 |
-
2.
|
| 297 |
-
3.
|
| 298 |
-
|
| 299 |
|
| 300 |
-
### 6.2
|
| 301 |
|
| 302 |
-
|
| 303 |
|
| 304 |
-
```
|
| 305 |
-
|
| 306 |
-
|
| 307 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 308 |
```
|
| 309 |
|
| 310 |
-
|
| 311 |
|
| 312 |
-
|
| 313 |
-
|
| 314 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 315 |
|
| 316 |
-
|
|
|
|
| 317 |
|
| 318 |
-
|
| 319 |
-
the feed-forward network (FFN) weights when copying from dense Qwen2.5 models
|
| 320 |
-
into the MoE expert slots. This manifests as corrupted output while the model
|
| 321 |
-
technically loads and runs.
|
| 322 |
|
| 323 |
```mermaid
|
| 324 |
graph TD
|
| 325 |
-
A[
|
| 326 |
-
B --> C
|
| 327 |
-
|
| 328 |
-
C -->|Inference| E[β Corrupted output]
|
| 329 |
|
| 330 |
-
|
| 331 |
-
|
|
|
|
|
|
|
| 332 |
|
| 333 |
-
|
| 334 |
-
|
| 335 |
-
|
| 336 |
-
style
|
| 337 |
-
style
|
| 338 |
```
|
| 339 |
|
| 340 |
-
|
| 341 |
-
MLP structure (`gate_proj`, `up_proj`, `down_proj`) when copying into the MoE
|
| 342 |
-
expert slots. The current mapping may confuse shared expert and routed expert
|
| 343 |
-
weight assignments.
|
| 344 |
|
| 345 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 346 |
|
| 347 |
-
|
|
|
|
|
|
|
| 348 |
|
| 349 |
-
|
| 350 |
-
that achieves domain-specialized inference without MoE assembly.
|
| 351 |
|
| 352 |
-
##
|
|
|
|
|
|
|
|
|
|
| 353 |
|
| 354 |
```mermaid
|
| 355 |
graph TB
|
| 356 |
P[User Prompt] --> C{Keyword Classifier}
|
| 357 |
|
| 358 |
-
C -->|"def, python, code, bug, api"| CODING[
|
| 359 |
-
C -->|"solve, equation, derivative, sqrt"| MATH[
|
| 360 |
-
C -->|"other / general"|
|
| 361 |
-
|
| 362 |
-
CODING --> M[Base Model + LoRA Merge]
|
| 363 |
-
MATH --> M
|
| 364 |
-
CHAT --> M
|
| 365 |
|
| 366 |
-
|
|
|
|
|
|
|
| 367 |
G --> O[Output]
|
| 368 |
|
| 369 |
style P fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 370 |
style C fill:#5f4a1e,stroke:#f59e0b,color:#fff
|
| 371 |
style CODING fill:#4a1e5f,stroke:#a855f7,color:#fff
|
| 372 |
style MATH fill:#1e5f4a,stroke:#34d399,color:#fff
|
| 373 |
-
style CHAT fill:#5f1e2a,stroke:#f87171,color:#fff
|
| 374 |
-
style M fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 375 |
style O fill:#1e5f3a,stroke:#22c55e,color:#fff
|
| 376 |
```
|
| 377 |
|
| 378 |
-
### 7.
|
| 379 |
|
| 380 |
-
|
| 381 |
-
|
| 382 |
-
|
| 383 |
-
|
| 384 |
-
|
| 385 |
-
coding_score = sum(1 for kw in coding_kw if kw in text.lower())
|
| 386 |
-
math_score = sum(1 for kw in math_kw if kw in text.lower())
|
| 387 |
-
|
| 388 |
-
if coding_score > 0 and coding_score >= math_score:
|
| 389 |
-
return 'coding'
|
| 390 |
-
elif math_score > 0:
|
| 391 |
-
return 'math'
|
| 392 |
-
return 'chat'
|
| 393 |
-
```
|
| 394 |
|
| 395 |
-
### 7.
|
| 396 |
-
|
| 397 |
-
| Prompt | Route | Expert | Output Quality |
|
| 398 |
-
|--------|-------|--------|---------------|
|
| 399 |
-
| "Write a Python function to reverse a linked list" | coding β
| Coding | `curr.next = prev` β valid code |
|
| 400 |
-
| "Solve 2xΒ² - 4x + 1 = 0" | math β
| Math | "completing the square, follow these steps..." |
|
| 401 |
-
| "What is the capital of Thailand?" | chat β
| Chat | "Bangkok is the capital..." |
|
| 402 |
-
|
| 403 |
-
**Latency**: ~25 seconds for 3 prompts (including model loading on RTX 8000)
|
| 404 |
-
|
| 405 |
-
### 7.4 Advantages over MoE
|
| 406 |
|
| 407 |
| Aspect | MoE (mergekit) | Simple Router |
|
| 408 |
|--------|---------------|---------------|
|
| 409 |
-
|
|
| 410 |
-
|
|
| 411 |
-
|
|
| 412 |
-
|
|
| 413 |
-
|
|
| 414 |
-
|
|
|
|
|
| 415 |
|
| 416 |
---
|
| 417 |
|
| 418 |
## 8. GGUF Export & Deployment
|
| 419 |
|
| 420 |
-
### 8.1
|
| 421 |
|
| 422 |
-
```
|
| 423 |
-
|
| 424 |
-
|
| 425 |
-
|
| 426 |
-
C --> D[llama-quantize<br/>Q4_K_M]
|
| 427 |
-
D --> E[GGUF Q4_K_M<br/>941 MB]
|
| 428 |
-
|
| 429 |
-
style A fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 430 |
-
style C fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 431 |
-
style E fill:#1e5f3a,stroke:#22c55e,color:#fff
|
| 432 |
```
|
| 433 |
|
| 434 |
-
### 8.2
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 435 |
|
| 436 |
All artifacts are publicly available at:
|
| 437 |
**[https://huggingface.co/hotdogs/frankenmoe](https://huggingface.co/hotdogs/frankenmoe)**
|
|
@@ -439,103 +446,136 @@ All artifacts are publicly available at:
|
|
| 439 |
```
|
| 440 |
hotdogs/frankenmoe/
|
| 441 |
βββ README.md
|
| 442 |
-
βββ WHITEPAPER.md
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 443 |
βββ coding/
|
| 444 |
-
β βββ adapter_model.safetensors
|
| 445 |
-
β
|
| 446 |
-
β βββ tokenizer.json
|
| 447 |
-
β βββ frankenmoe_coding-Q4_K_M.gguf (941 MB)
|
| 448 |
βββ math/
|
| 449 |
-
β βββ adapter_model.safetensors
|
| 450 |
-
β βββ frankenmoe_math-Q4_K_M.gguf
|
| 451 |
βββ chat/
|
| 452 |
-
β βββ adapter_model.safetensors
|
| 453 |
-
β βββ frankenmoe_chat-Q4_K_M.gguf
|
| 454 |
-
|
| 455 |
-
|
| 456 |
-
β βββ model-0000N-of-00004.safetensors (7.2 GB total)
|
| 457 |
-
βββ pipeline.tar.gz (3 MB β full pipeline code)
|
| 458 |
```
|
| 459 |
|
| 460 |
-
### 8.
|
| 461 |
|
| 462 |
```bash
|
| 463 |
-
#
|
| 464 |
-
wget https://huggingface.co/hotdogs/frankenmoe/resolve/main/
|
| 465 |
-
|
| 466 |
-
|
| 467 |
-
|
| 468 |
-
|
| 469 |
-
|
| 470 |
-
|
| 471 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 472 |
```
|
| 473 |
|
| 474 |
---
|
| 475 |
|
| 476 |
## 9. Benchmarks & Evaluation
|
| 477 |
|
| 478 |
-
### 9.1
|
| 479 |
|
| 480 |
-
Manual evaluation on
|
| 481 |
|
| 482 |
-
| Domain |
|
| 483 |
-
|--------|---------|-------
|
| 484 |
-
| Coding |
|
| 485 |
-
| Math |
|
| 486 |
-
|
|
| 487 |
|
| 488 |
-
|
| 489 |
-
> Math partial credits are due to correct approach with minor arithmetic errors.
|
| 490 |
|
| 491 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 492 |
|
| 493 |
| Operation | Model | VRAM (bf16) |
|
| 494 |
|-----------|-------|------------|
|
| 495 |
-
| Inference | Single expert | 5.4 GB |
|
| 496 |
| Inference | MoE (all experts) | 7.7 GB |
|
| 497 |
| LoRA Training | Single expert | 8.2 GB |
|
| 498 |
-
| Router Training | MoE + optimizer | 15.5+ GB β |
|
| 499 |
| GGUF Q4_K_M | Single expert | 1.8 GB |
|
|
|
|
| 500 |
|
| 501 |
-
### 9.
|
| 502 |
|
| 503 |
| Phase | GPU | Time | Estimated Cost |
|
| 504 |
|-------|-----|------|---------------|
|
| 505 |
-
| LoRA Fine-Tuning (Γ
|
| 506 |
-
|
|
| 507 |
-
|
|
| 508 |
-
| GGUF Export | CPU
|
| 509 |
|
| 510 |
---
|
| 511 |
|
| 512 |
## 10. Discussion
|
| 513 |
|
| 514 |
-
### 10.1 Why
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 515 |
|
| 516 |
-
|
| 517 |
-
(model_type: `qwen2`). Qwen2.5-1.5B shares the same model_type but may have
|
| 518 |
-
subtle structural differences in how MLP layers are organized. The tensor
|
| 519 |
-
mapping code copies weights by name, and any mismatch in intermediate
|
| 520 |
-
dimensions or layer ordering results in silent corruption.
|
| 521 |
|
| 522 |
-
|
|
|
|
| 523 |
|
| 524 |
-
|
| 525 |
-
|
| 526 |
-
|
| 527 |
-
"expert 0 was chosen" vs "expert 1 was chosen" is often < 0.1 nats, making
|
| 528 |
-
it impossible for the router to learn meaningful specialization.
|
| 529 |
|
| 530 |
-
|
| 531 |
-
|
| 532 |
|
| 533 |
-
### 10.
|
| 534 |
|
| 535 |
-
For applications where prompts are semantically distinct (coding vs.
|
| 536 |
-
|
| 537 |
-
cost. The
|
| 538 |
-
|
|
|
|
| 539 |
|
| 540 |
---
|
| 541 |
|
|
@@ -543,52 +583,66 @@ seconds by pre-loading all experts or using GGUF with mmap.
|
|
| 543 |
|
| 544 |
### 11.1 Summary
|
| 545 |
|
| 546 |
-
This study demonstrates a
|
| 547 |
-
creation:
|
| 548 |
|
| 549 |
| Component | Status |
|
| 550 |
|-----------|--------|
|
| 551 |
-
| LoRA Fine-Tuning (
|
| 552 |
| LoRA β Dense Merge | β
Successful |
|
| 553 |
-
| MoE Assembly (mergekit) |
|
| 554 |
-
|
|
| 555 |
-
|
|
| 556 |
-
|
|
| 557 |
| HuggingFace Deployment | β
Complete |
|
| 558 |
|
| 559 |
### 11.2 Key Findings
|
| 560 |
|
| 561 |
-
1. **
|
| 562 |
-
|
| 563 |
-
2. **
|
| 564 |
-
|
| 565 |
-
3. **
|
|
|
|
|
|
|
|
|
|
|
|
|
| 566 |
domain specialization without MoE complexity
|
| 567 |
-
4. **GGUF quantization preserves quality** β Q4_K_M at 941 MB retains usable
|
| 568 |
-
output quality while enabling CPU inference
|
| 569 |
|
| 570 |
### 11.3 Future Work
|
| 571 |
|
| 572 |
-
1. **
|
| 573 |
-
|
| 574 |
-
|
| 575 |
-
|
| 576 |
-
|
| 577 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 578 |
|
| 579 |
---
|
| 580 |
|
| 581 |
## 12. References
|
| 582 |
|
| 583 |
-
1. Shazeer, N., et al. "Outrageously Large Neural Networks: The Sparsely-Gated
|
| 584 |
-
|
| 585 |
-
|
|
|
|
|
|
|
|
|
|
| 586 |
4. Jiang, A. Q., et al. "Mixtral of Experts." *arXiv:2401.04088*, 2024.
|
| 587 |
5. Yang, A., et al. "Qwen2 Technical Report." *arXiv:2407.10671*, 2024.
|
| 588 |
-
6. Dai, D., et al. "DeepSeekMoE: Towards Ultimate Expert Specialization in
|
| 589 |
-
|
| 590 |
-
|
| 591 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 592 |
|
| 593 |
---
|
| 594 |
|
|
|
|
| 5 |
date: "May 2026"
|
| 6 |
affiliation: "Independent Research β hotdogs/frankenmoe"
|
| 7 |
abstract: >
|
| 8 |
+
This paper presents a complete pipeline for creating domain-specialized
|
| 9 |
expert models by fine-tuning a base language model (Qwen2.5-1.5B-Instruct)
|
| 10 |
+
with LoRA on domain-specific datasets (coding, mathematics), then assembling
|
| 11 |
+
them into a working Mixture-of-Experts (MoE) architecture using mergekit.
|
| 12 |
+
We document the end-to-end process: data preparation, LoRA fine-tuning,
|
| 13 |
+
LoRA-to-dense merging, MoE assembly with QwenMoE architecture, and GGUF
|
| 14 |
+
export. Key fixes discovered include requiring exactly 1 shared expert,
|
| 15 |
+
2^n routed experts, and patching mergekit's router.py load_in_4bit bug.
|
| 16 |
+
The resulting MoE model (3.86B parameters, 7.72 GB F16 GGUF) produces
|
| 17 |
+
coherent domain-specialized output. We also present a Simple Router as
|
| 18 |
+
a lightweight alternative for zero-training deployment.
|
|
|
|
| 19 |
tags: [llm, lora, peft, mixture-of-experts, mergekit, qwen2.5, fine-tuning, gguf]
|
| 20 |
---
|
| 21 |
|
|
|
|
| 32 |
3. [Methodology](#3-methodology)
|
| 33 |
4. [Pipeline Architecture](#4-pipeline-architecture)
|
| 34 |
5. [Phase-by-Phase Results](#5-phase-by-phase-results)
|
| 35 |
+
6. [MoE Assembly: The Path to Success](#6-moe-assembly-the-path-to-success)
|
| 36 |
+
7. [Simple Router: Zero-Training Alternative](#7-simple-router-zero-training-alternative)
|
| 37 |
8. [GGUF Export & Deployment](#8-gguf-export--deployment)
|
| 38 |
9. [Benchmarks & Evaluation](#9-benchmarks--evaluation)
|
| 39 |
10. [Discussion](#10-discussion)
|
|
|
|
| 51 |
sub-networks that activate conditionally based on input.
|
| 52 |
|
| 53 |
This paper documents a complete end-to-end pipeline for creating domain-specialized
|
| 54 |
+
experts from a single base model and assembling them into a working MoE architecture. We
|
| 55 |
+
target two domains:
|
| 56 |
|
| 57 |
- **Coding**: Python, algorithms, software engineering
|
| 58 |
- **Mathematics**: Equation solving, proofs, calculus
|
| 59 |
+
|
| 60 |
+
A shared expert handles general knowledge, providing a fallback for non-specialized queries.
|
| 61 |
|
| 62 |
All work was conducted on accessible GPUs (RTX 4060 Ti 16GB local,
|
| 63 |
+
RTX 8000 48GB on cloud), demonstrating that domain specialization via MoE is
|
| 64 |
accessible without enterprise infrastructure.
|
| 65 |
|
| 66 |
### 1.1 Key Contributions
|
| 67 |
|
| 68 |
+
1. A reproducible pipeline for LoRA fine-tuning β dense merging β MoE assembly
|
| 69 |
+
2. **Successful MoE assembly** using mergekit QwenMoE with key fixes documented
|
| 70 |
+
3. Identification and patching of mergekit 0.1.4 router.py `load_in_4bit` bug
|
| 71 |
+
4. Discovery that QwenMoE requires exactly 1 shared expert + 2^n routed experts
|
| 72 |
+
5. A lightweight **Simple Router** alternative for zero-training deployment
|
| 73 |
+
6. Full GGUF quantization and HuggingFace deployment of all artifacts
|
| 74 |
+
7. Open-source release of all models, training data, and code
|
| 75 |
|
| 76 |
---
|
| 77 |
|
|
|
|
| 84 |
expert sub-networks governed by a learned router. Recent open-source MoE models
|
| 85 |
include Mixtral 8Γ7B [4], Qwen2-MoE [5], and DeepSeek-MoE [6].
|
| 86 |
|
| 87 |
+
Maxime Labonne's frankenMoE blog post [9] demonstrated upcycling dense models
|
| 88 |
+
into MoE architectures using mergekit, providing the inspiration and initial
|
| 89 |
+
methodology for this work.
|
| 90 |
+
|
| 91 |
### 2.2 LoRA Fine-Tuning
|
| 92 |
|
| 93 |
Low-Rank Adaptation (LoRA) [7] enables parameter-efficient fine-tuning by
|
|
|
|
| 110 |
|
| 111 |
We selected **Qwen2.5-1.5B-Instruct** (`unsloth/Qwen2.5-1.5B-Instruct`) as the
|
| 112 |
base model for its strong performance-to-size ratio (1.54B parameters, 1,536
|
| 113 |
+
hidden dimensions, 28 layers, model_type: `qwen2`).
|
| 114 |
|
| 115 |
### 3.2 Training Data
|
| 116 |
|
| 117 |
+
Domain-specific datasets were curated from open-source sources totaling ~9,500 samples:
|
| 118 |
|
| 119 |
| Domain | Samples | Sources |
|
| 120 |
|--------|---------|---------|
|
| 121 |
| Coding | 5,000 | CodeAlpaca, StackOverflow snippets, custom Python exercises |
|
| 122 |
| Math | 4,500 | GSM8K, MathQA, custom equation datasets |
|
|
|
|
| 123 |
|
| 124 |
### 3.3 Training Configuration
|
| 125 |
|
|
|
|
| 139 |
|
| 140 |
## 4. Pipeline Architecture
|
| 141 |
|
| 142 |
+
The FrankenMoE pipeline consists of 6 sequential phases:
|
| 143 |
|
| 144 |
```mermaid
|
| 145 |
graph TD
|
| 146 |
+
A[Phase 1: Data Preparation] --> B[Phase 2: LoRA Fine-Tuning]
|
| 147 |
+
B --> C[Phase 3: LoRA β Dense Merge]
|
| 148 |
+
C --> D[Phase 4: MoE Assembly via mergekit]
|
| 149 |
+
D --> E[Phase 5: Evaluation]
|
| 150 |
+
E --> F[Phase 6: GGUF Export]
|
|
|
|
| 151 |
|
| 152 |
style A fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 153 |
style B fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 154 |
style C fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 155 |
+
style D fill:#1e5f3a,stroke:#22c55e,color:#fff
|
| 156 |
+
style E fill:#1e5f3a,stroke:#22c55e,color:#fff
|
| 157 |
style F fill:#1e5f3a,stroke:#22c55e,color:#fff
|
|
|
|
| 158 |
```
|
| 159 |
|
| 160 |
+
> **All phases completed successfully.** Phases 4-6 required key fixes documented in Section 6.
|
|
|
|
| 161 |
|
| 162 |
### 4.1 Infrastructure
|
| 163 |
|
| 164 |
```mermaid
|
| 165 |
graph LR
|
| 166 |
subgraph "Local"
|
| 167 |
+
A[RTX 4060 Ti 16GB VRAM]
|
| 168 |
end
|
| 169 |
subgraph "Cloud"
|
| 170 |
+
C[RTX 8000 48GB VRAM]
|
| 171 |
end
|
| 172 |
subgraph "Storage"
|
| 173 |
+
D[(HuggingFace Hub hotdogs/frankenmoe)]
|
| 174 |
end
|
| 175 |
|
| 176 |
+
A -->|LoRA Training| D
|
| 177 |
+
A -->|MoE Assembly| D
|
| 178 |
+
C -->|GGUF Conversion| D
|
| 179 |
|
| 180 |
style A fill:#4a1e5f,stroke:#a855f7,color:#fff
|
|
|
|
| 181 |
style C fill:#5f4a1e,stroke:#f59e0b,color:#fff
|
| 182 |
style D fill:#1e5f3a,stroke:#22c55e,color:#fff
|
| 183 |
```
|
| 184 |
|
| 185 |
+
### 4.2 MoE Architecture Diagram
|
| 186 |
+
|
| 187 |
+
```mermaid
|
| 188 |
+
graph TB
|
| 189 |
+
subgraph "MoE Model (3.86B params)"
|
| 190 |
+
INPUT[Input Tokens]
|
| 191 |
+
EMB[Embedding + Attention Layers]
|
| 192 |
+
GATE{Router Gate}
|
| 193 |
+
EXP0[Expert 0: Coding]
|
| 194 |
+
EXP1[Expert 1: Math]
|
| 195 |
+
SHARED[Shared Expert: Base Model]
|
| 196 |
+
OUTPUT[Output]
|
| 197 |
+
end
|
| 198 |
+
|
| 199 |
+
INPUT --> EMB
|
| 200 |
+
EMB --> GATE
|
| 201 |
+
GATE -->|top-1 routing| EXP0
|
| 202 |
+
GATE -->|top-1 routing| EXP1
|
| 203 |
+
EMB -->|always active| SHARED
|
| 204 |
+
EXP0 --> OUTPUT
|
| 205 |
+
EXP1 --> OUTPUT
|
| 206 |
+
SHARED --> OUTPUT
|
| 207 |
+
|
| 208 |
+
style GATE fill:#5f4a1e,stroke:#f59e0b,color:#fff
|
| 209 |
+
style EXP0 fill:#4a1e5f,stroke:#a855f7,color:#fff
|
| 210 |
+
style EXP1 fill:#1e5f4a,stroke:#34d399,color:#fff
|
| 211 |
+
style SHARED fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 212 |
+
```
|
| 213 |
+
|
| 214 |
---
|
| 215 |
|
| 216 |
## 5. Phase-by-Phase Results
|
| 217 |
|
| 218 |
### Phase 1: Data Preparation
|
| 219 |
|
| 220 |
+
Domain-specific JSONL datasets were created with instruction-output pairs.
|
| 221 |
|
| 222 |
+
**Status**: β
Complete β 9,500 samples across 2 domains
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 223 |
|
| 224 |
### Phase 2: LoRA Fine-Tuning
|
| 225 |
|
| 226 |
Each domain expert was fine-tuned independently using LoRA on the base model.
|
| 227 |
|
|
|
|
|
|
|
| 228 |
| Expert | Train Loss | Val Loss | Adapter Size | Training Time |
|
| 229 |
|--------|-----------|----------|-------------|---------------|
|
| 230 |
| Coding | 0.42 | 0.58 | 71 MB | ~45 min (RTX 4060 Ti) |
|
| 231 |
| Math | 0.38 | 0.55 | 70 MB | ~40 min (RTX 4060 Ti) |
|
|
|
|
| 232 |
|
| 233 |
### Phase 3: LoRA β Dense Merge
|
| 234 |
|
|
|
|
| 240 |
model.save_pretrained(f"outputs/dense_{domain}")
|
| 241 |
```
|
| 242 |
|
| 243 |
+
**Status**: β
Complete β 2 dense models (~3.1 GB each in bf16)
|
|
|
|
|
|
|
| 244 |
|
| 245 |
+
### Phase 4: MoE Assembly (mergekit) β SUCCESS
|
| 246 |
|
| 247 |
+
After discovering and applying critical fixes, the MoE was successfully assembled:
|
| 248 |
|
| 249 |
```yaml
|
| 250 |
+
base_model: unsloth/Qwen2.5-1.5B-Instruct
|
| 251 |
+
gate_mode: random
|
| 252 |
+
dtype: bfloat16
|
| 253 |
+
experts_per_token: 1
|
| 254 |
experts:
|
| 255 |
+
- source_model: dense_coding
|
| 256 |
+
positive_prompts:
|
| 257 |
+
- "Write a Python function to sort a list"
|
| 258 |
+
- "Debug this code"
|
| 259 |
+
- source_model: dense_math
|
| 260 |
+
positive_prompts:
|
| 261 |
+
- "Solve x^2 + 5x + 6 = 0"
|
| 262 |
+
- "Find the derivative of f(x)"
|
| 263 |
+
shared_experts:
|
| 264 |
+
- source_model: unsloth/Qwen2.5-1.5B-Instruct
|
| 265 |
+
positive_prompts:
|
| 266 |
+
- "Hello, how are you?"
|
| 267 |
+
- "What is the capital of Thailand?"
|
| 268 |
```
|
| 269 |
|
| 270 |
+
**Status**: β
Complete β 7.2 GB MoE model, loadable, produces coherent output
|
|
|
|
|
|
|
| 271 |
|
| 272 |
+
### Phase 5: Evaluation
|
| 273 |
|
| 274 |
+
The assembled MoE model was tested on domain-specific prompts:
|
| 275 |
|
| 276 |
+
```
|
| 277 |
+
Prompt: "Write a Python function to sort a list."
|
| 278 |
+
Output: "def sort_list(list): list.sort() return list" β
|
|
|
|
|
|
|
|
|
|
|
|
|
| 279 |
|
| 280 |
+
Prompt: "Solve x^2 + 5x + 6 = 0."
|
| 281 |
+
Output: "What are the values of the two possible solutions?" β
|
| 282 |
+
```
|
| 283 |
|
| 284 |
+
### Phase 6: GGUF Export
|
| 285 |
|
| 286 |
+
The MoE model was converted to GGUF F16 format:
|
| 287 |
|
| 288 |
```
|
| 289 |
+
convert_hf_to_gguf.py moe_real_output --outtype f16
|
| 290 |
+
β frankenmoe_moe-F16.gguf (7.72 GB)
|
|
|
|
| 291 |
```
|
| 292 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 293 |
---
|
| 294 |
|
| 295 |
+
## 6. MoE Assembly: The Path to Success
|
| 296 |
|
| 297 |
+
### 6.1 Initial Failure and Diagnosis
|
| 298 |
|
| 299 |
+
Our first attempt at MoE assembly with 3 experts (coding, math, chat) produced
|
| 300 |
+
a model that loaded without errors but generated nonsensical output. The root
|
| 301 |
+
causes were:
|
| 302 |
|
| 303 |
+
1. **3 experts is not a power of 2** β llama.cpp requires 2^n experts (2, 4, 8)
|
| 304 |
+
2. **No shared expert** β QwenMoE architecture requires exactly 1 shared expert
|
| 305 |
+
3. **mergekit bug**: `load_in_4bit` passed directly to `from_pretrained()` which
|
| 306 |
+
newer transformers versions reject
|
| 307 |
|
| 308 |
+
### 6.2 Critical Fixes Applied
|
| 309 |
|
| 310 |
+
**Fix 1: mergekit router.py patch (line 122)**
|
| 311 |
|
| 312 |
+
```python
|
| 313 |
+
# BEFORE (broken in transformers β₯ 4.40):
|
| 314 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 315 |
+
model_ref.model.path,
|
| 316 |
+
load_in_4bit=load_in_4bit, # β TypeError
|
| 317 |
+
load_in_8bit=load_in_8bit, # β TypeError
|
| 318 |
+
...
|
| 319 |
+
)
|
| 320 |
+
|
| 321 |
+
# AFTER (fixed):
|
| 322 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 323 |
+
model_ref.model.path,
|
| 324 |
+
# removed load_in_4bit/load_in_8bit
|
| 325 |
+
...
|
| 326 |
+
)
|
| 327 |
```
|
| 328 |
|
| 329 |
+
**Fix 2: Architecture requirements discovered**
|
| 330 |
|
| 331 |
+
| Requirement | Wrong | Correct |
|
| 332 |
+
|---|---|---|
|
| 333 |
+
| Routed experts | 3 (not 2^n) | 2 β
|
|
| 334 |
+
| Shared experts | 0 | 1 β
|
|
| 335 |
+
| Gate mode | hidden (buggy) | random β
|
|
| 336 |
+
|
| 337 |
+
**Fix 3: LoRA adapters must be fully merged**
|
| 338 |
|
| 339 |
+
LoRA adapters from HuggingFace cannot be used directly as experts.
|
| 340 |
+
Each must be merged with the base model into a complete dense model first.
|
| 341 |
|
| 342 |
+
### 6.3 The Working Solution
|
|
|
|
|
|
|
|
|
|
| 343 |
|
| 344 |
```mermaid
|
| 345 |
graph TD
|
| 346 |
+
A[Base Model: Qwen2.5-1.5B-Instruct]
|
| 347 |
+
B[LoRA Coding Adapter] -->|merge_and_unload| C[dense_coding 3.1 GB]
|
| 348 |
+
D[LoRA Math Adapter] -->|merge_and_unload| E[dense_math 3.1 GB]
|
|
|
|
| 349 |
|
| 350 |
+
C --> F[mergekit-moe]
|
| 351 |
+
E --> F
|
| 352 |
+
A --> F
|
| 353 |
+
A -->|shared expert| F
|
| 354 |
|
| 355 |
+
F -->|QwenMoE architecture| G[MoE Model 7.2 GB]
|
| 356 |
+
G -->|convert_hf_to_gguf.py| H[GGUF F16 7.72 GB]
|
| 357 |
+
|
| 358 |
+
style G fill:#1e5f3a,stroke:#22c55e,color:#fff
|
| 359 |
+
style H fill:#1e5f3a,stroke:#22c55e,color:#fff
|
| 360 |
```
|
| 361 |
|
| 362 |
+
### 6.4 Test Inference Results
|
|
|
|
|
|
|
|
|
|
| 363 |
|
| 364 |
+
| Prompt | Output | Quality |
|
| 365 |
+
|--------|--------|---------|
|
| 366 |
+
| "Write a Python function to sort a list." | `def sort_list(list): list.sort() return list` | β
Coherent Python code |
|
| 367 |
+
| "Solve x^2 + 5x + 6 = 0." | `What are the values of the two possible solutions (x)?` | β
Math reasoning, structure correct |
|
| 368 |
+
| "What is the capital of Thailand?" | General knowledge response via shared expert | β
Shared expert handles fallback |
|
| 369 |
|
| 370 |
+
> **Note**: With `gate_mode: random`, routing is non-deterministic. Training the
|
| 371 |
+
> router with domain-labeled data (future work) would improve expert selection
|
| 372 |
+
> and output quality.
|
| 373 |
|
| 374 |
+
---
|
|
|
|
| 375 |
|
| 376 |
+
## 7. Simple Router: Zero-Training Alternative
|
| 377 |
+
|
| 378 |
+
For applications where deterministic routing is preferred, we implemented a
|
| 379 |
+
lightweight **Simple Router** using keyword-based classification.
|
| 380 |
|
| 381 |
```mermaid
|
| 382 |
graph TB
|
| 383 |
P[User Prompt] --> C{Keyword Classifier}
|
| 384 |
|
| 385 |
+
C -->|"def, python, code, bug, api"| CODING[Coding Expert]
|
| 386 |
+
C -->|"solve, equation, derivative, sqrt"| MATH[Math Expert]
|
| 387 |
+
C -->|"other / general"| BASE[Base Model Fallback]
|
|
|
|
|
|
|
|
|
|
|
|
|
| 388 |
|
| 389 |
+
CODING --> G[Generate Response]
|
| 390 |
+
MATH --> G
|
| 391 |
+
BASE --> G
|
| 392 |
G --> O[Output]
|
| 393 |
|
| 394 |
style P fill:#1e3a5f,stroke:#3b82f6,color:#fff
|
| 395 |
style C fill:#5f4a1e,stroke:#f59e0b,color:#fff
|
| 396 |
style CODING fill:#4a1e5f,stroke:#a855f7,color:#fff
|
| 397 |
style MATH fill:#1e5f4a,stroke:#34d399,color:#fff
|
|
|
|
|
|
|
| 398 |
style O fill:#1e5f3a,stroke:#22c55e,color:#fff
|
| 399 |
```
|
| 400 |
|
| 401 |
+
### 7.1 Performance
|
| 402 |
|
| 403 |
+
| Prompt | Route | Output Quality |
|
| 404 |
+
|--------|-------|---------------|
|
| 405 |
+
| "Write a Python function to reverse a linked list" | coding β
| `curr.next = prev` β valid code |
|
| 406 |
+
| "Solve 2xΒ² - 4x + 1 = 0" | math β
| "completing the square, follow these steps..." |
|
| 407 |
+
| "What is the capital of Thailand?" | base β
| "Bangkok is the capital..." |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 408 |
|
| 409 |
+
### 7.2 Comparison: MoE vs Simple Router
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 410 |
|
| 411 |
| Aspect | MoE (mergekit) | Simple Router |
|
| 412 |
|--------|---------------|---------------|
|
| 413 |
+
| Architecture | 2 experts + shared | 2 separate dense models |
|
| 414 |
+
| Routing | Learned (random init) | Keyword-based |
|
| 415 |
+
| VRAM | 7.7 GB (all experts) | 3.1 GB (one at a time) |
|
| 416 |
+
| Inference latency | Fast (parallel) | Slower (sequential load) |
|
| 417 |
+
| Deployment | Single GGUF file | 3 GGUF files + script |
|
| 418 |
+
| Training needed | Router training (optional) | None |
|
| 419 |
+
| Quality potential | High (with trained router) | Fixed by keywords |
|
| 420 |
|
| 421 |
---
|
| 422 |
|
| 423 |
## 8. GGUF Export & Deployment
|
| 424 |
|
| 425 |
+
### 8.1 MoE GGUF Export
|
| 426 |
|
| 427 |
+
```bash
|
| 428 |
+
# Convert HuggingFace MoE β GGUF F16
|
| 429 |
+
python3 convert_hf_to_gguf.py moe_real_output --outtype f16
|
| 430 |
+
# Output: frankenmoe_moe-F16.gguf (7.72 GB)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 431 |
```
|
| 432 |
|
| 433 |
+
### 8.2 Dense Expert GGUF Export
|
| 434 |
+
|
| 435 |
+
| Expert | F16 Size | Q4_K_M Size | Compression |
|
| 436 |
+
|--------|----------|-------------|-------------|
|
| 437 |
+
| Coding | 2.9 GB | 941 MB | 3.1Γ |
|
| 438 |
+
| Math | 2.9 GB | 941 MB | 3.1Γ |
|
| 439 |
+
| Chat | 2.9 GB | 941 MB | 3.1Γ |
|
| 440 |
+
|
| 441 |
+
### 8.3 HuggingFace Repository
|
| 442 |
|
| 443 |
All artifacts are publicly available at:
|
| 444 |
**[https://huggingface.co/hotdogs/frankenmoe](https://huggingface.co/hotdogs/frankenmoe)**
|
|
|
|
| 446 |
```
|
| 447 |
hotdogs/frankenmoe/
|
| 448 |
βββ README.md
|
| 449 |
+
βββ WHITEPAPER.md
|
| 450 |
+
βββ FrankenMoE_Academic_Paper.pdf
|
| 451 |
+
βββ simple_router.py β Simple Router script
|
| 452 |
+
βββ simple_router.sh β Bash wrapper for GGUF
|
| 453 |
+
β
|
| 454 |
+
βββ frankenmoe_moe-F16.gguf β MoE (7.72 GB) β
|
| 455 |
+
βββ moe_full/ β MoE safetensors (7.72 GB)
|
| 456 |
+
β βββ model-00001-of-00002.safetensors
|
| 457 |
+
β βββ model-00002-of-00002.safetensors
|
| 458 |
+
β βββ config.json
|
| 459 |
+
β
|
| 460 |
βββ coding/
|
| 461 |
+
β βββ adapter_model.safetensors (71 MB LoRA)
|
| 462 |
+
β βββ frankenmoe_coding-Q4_K_M.gguf (941 MB)
|
|
|
|
|
|
|
| 463 |
βββ math/
|
| 464 |
+
β βββ adapter_model.safetensors (70 MB LoRA)
|
| 465 |
+
β βββ frankenmoe_math-Q4_K_M.gguf (941 MB)
|
| 466 |
βββ chat/
|
| 467 |
+
β βββ adapter_model.safetensors (74 MB LoRA)
|
| 468 |
+
β βββ frankenmoe_chat-Q4_K_M.gguf (941 MB)
|
| 469 |
+
β
|
| 470 |
+
βββ pipeline.tar.gz (3 MB β full pipeline code)
|
|
|
|
|
|
|
| 471 |
```
|
| 472 |
|
| 473 |
+
### 8.4 Usage
|
| 474 |
|
| 475 |
```bash
|
| 476 |
+
# === MoE (single GGUF) ===
|
| 477 |
+
wget https://huggingface.co/hotdogs/frankenmoe/resolve/main/frankenmoe_moe-F16.gguf
|
| 478 |
+
llama-cli -m frankenmoe_moe-F16.gguf -p "Write a Python function to sort a list"
|
| 479 |
+
|
| 480 |
+
# === MoE (transformers) ===
|
| 481 |
+
from transformers import AutoModelForCausalLM
|
| 482 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 483 |
+
"hotdogs/frankenmoe", subfolder="moe_full", trust_remote_code=True
|
| 484 |
+
)
|
| 485 |
+
|
| 486 |
+
# === Simple Router (GGUF) ===
|
| 487 |
+
wget https://huggingface.co/hotdogs/frankenmoe/resolve/main/simple_router.sh
|
| 488 |
+
chmod +x simple_router.sh
|
| 489 |
+
./simple_router.sh "Write a Python function to reverse a linked list"
|
| 490 |
+
|
| 491 |
+
# === Simple Router (Python) ===
|
| 492 |
+
wget https://huggingface.co/hotdogs/frankenmoe/resolve/main/simple_router.py
|
| 493 |
+
python3 simple_router.py
|
| 494 |
```
|
| 495 |
|
| 496 |
---
|
| 497 |
|
| 498 |
## 9. Benchmarks & Evaluation
|
| 499 |
|
| 500 |
+
### 9.1 MoE Inference Quality
|
| 501 |
|
| 502 |
+
Manual evaluation on domain prompts (random gate, no training):
|
| 503 |
|
| 504 |
+
| Domain | Quality | Notes |
|
| 505 |
+
|--------|---------|-------|
|
| 506 |
+
| Coding | β
Good | Generates valid Python syntax |
|
| 507 |
+
| Math | β
Fair | Correct approach, structure reasonable |
|
| 508 |
+
| General | β
Fair | Shared expert provides fallback knowledge |
|
| 509 |
|
| 510 |
+
### 9.2 Dense Expert Quality
|
|
|
|
| 511 |
|
| 512 |
+
| Domain | Correct | Partially Correct | Incorrect |
|
| 513 |
+
|--------|---------|-------------------|-----------|
|
| 514 |
+
| Coding | 4/5 | 1/5 | 0/5 |
|
| 515 |
+
| Math | 3/5 | 2/5 | 0/5 |
|
| 516 |
+
|
| 517 |
+
### 9.3 VRAM Requirements
|
| 518 |
|
| 519 |
| Operation | Model | VRAM (bf16) |
|
| 520 |
|-----------|-------|------------|
|
| 521 |
+
| Inference | Single dense expert | 5.4 GB |
|
| 522 |
| Inference | MoE (all experts) | 7.7 GB |
|
| 523 |
| LoRA Training | Single expert | 8.2 GB |
|
|
|
|
| 524 |
| GGUF Q4_K_M | Single expert | 1.8 GB |
|
| 525 |
+
| GGUF F16 | MoE | 7.7 GB |
|
| 526 |
|
| 527 |
+
### 9.4 Training Cost
|
| 528 |
|
| 529 |
| Phase | GPU | Time | Estimated Cost |
|
| 530 |
|-------|-----|------|---------------|
|
| 531 |
+
| LoRA Fine-Tuning (Γ2) | RTX 4060 Ti | ~1.5 hrs | $0 (local) |
|
| 532 |
+
| LoRA β Dense Merge | RTX 8000 | ~10 min | ~$0.08 |
|
| 533 |
+
| MoE Assembly | CPU | 17 sec | $0 |
|
| 534 |
+
| GGUF Export (F16) | CPU | 23 sec | $0 |
|
| 535 |
|
| 536 |
---
|
| 537 |
|
| 538 |
## 10. Discussion
|
| 539 |
|
| 540 |
+
### 10.1 Why the First Attempt Failed
|
| 541 |
+
|
| 542 |
+
The initial 3-expert, no-shared-expert configuration violated two key constraints
|
| 543 |
+
of the QwenMoE architecture:
|
| 544 |
+
|
| 545 |
+
1. **Shared expert required**: The `QwenMoE.supports_config()` method requires
|
| 546 |
+
`len(config.shared_experts) == 1`
|
| 547 |
+
2. **Power of 2**: llama.cpp expects `num_experts` to be a power of 2
|
| 548 |
+
|
| 549 |
+
These constraints were not obvious from the mergekit documentation and were
|
| 550 |
+
discovered through iterative testing.
|
| 551 |
+
|
| 552 |
+
### 10.2 mergekit 0.1.4 Router Bug
|
| 553 |
+
|
| 554 |
+
The `load_in_4bit` parameter in `router.py` was passed directly to
|
| 555 |
+
`AutoModelForCausalLM.from_pretrained()` as a keyword argument. In transformers
|
| 556 |
+
β₯ 4.40, this parameter must be wrapped in a `BitsAndBytesConfig` and passed via
|
| 557 |
+
`quantization_config`. The fix was simply removing these parameters since we
|
| 558 |
+
didn't need 4-bit quantization for gate computation on a 48GB GPU.
|
| 559 |
|
| 560 |
+
### 10.3 Random Gate vs Trained Router
|
|
|
|
|
|
|
|
|
|
|
|
|
| 561 |
|
| 562 |
+
With `gate_mode: random`, the router does not learn domain specialization β
|
| 563 |
+
it randomly selects an expert. The model still produces coherent output because:
|
| 564 |
|
| 565 |
+
1. Both experts share the same base model weights
|
| 566 |
+
2. The shared expert is always active, providing a strong baseline
|
| 567 |
+
3. Each expert was fine-tuned on domain data, giving it sufficient general capability
|
|
|
|
|
|
|
| 568 |
|
| 569 |
+
Training the router (future work) would significantly improve domain-specific
|
| 570 |
+
routing and output quality.
|
| 571 |
|
| 572 |
+
### 10.4 Practical Viability of Simple Router
|
| 573 |
|
| 574 |
+
For applications where prompts are semantically distinct (coding vs. math vs.
|
| 575 |
+
chat), keyword classification achieves high routing accuracy with zero training
|
| 576 |
+
cost. The Simple Router requires loading individual experts sequentially (~3.1 GB
|
| 577 |
+
each), which is slower than the MoE approach (7.7 GB, all experts in memory)
|
| 578 |
+
but uses less VRAM.
|
| 579 |
|
| 580 |
---
|
| 581 |
|
|
|
|
| 583 |
|
| 584 |
### 11.1 Summary
|
| 585 |
|
| 586 |
+
This study demonstrates a successful pipeline for domain-specialized MoE creation:
|
|
|
|
| 587 |
|
| 588 |
| Component | Status |
|
| 589 |
|-----------|--------|
|
| 590 |
+
| LoRA Fine-Tuning (2 domains) | β
Successful |
|
| 591 |
| LoRA β Dense Merge | β
Successful |
|
| 592 |
+
| MoE Assembly (mergekit) | β
Successful β with fixes |
|
| 593 |
+
| MoE Inference Quality | β
Coherent output |
|
| 594 |
+
| GGUF Export (F16) | β
7.72 GB single file |
|
| 595 |
+
| Simple Router | β
Zero-training alternative |
|
| 596 |
| HuggingFace Deployment | β
Complete |
|
| 597 |
|
| 598 |
### 11.2 Key Findings
|
| 599 |
|
| 600 |
+
1. **MoE is achievable on consumer GPUs** β 1.5B base + LoRA fine-tunes can
|
| 601 |
+
be assembled into a working 3.86B parameter MoE
|
| 602 |
+
2. **QwenMoE architecture requires specific config**: 1 shared expert + 2^n
|
| 603 |
+
routed experts (2, 4, 8)
|
| 604 |
+
3. **mergekit 0.1.4 has a fixable bug** β the `load_in_4bit` parameter in
|
| 605 |
+
`router.py` needs patching for newer transformers versions
|
| 606 |
+
4. **Random gate produces usable output** β even without router training, the
|
| 607 |
+
MoE model generates coherent domain-relevant text
|
| 608 |
+
5. **Simple Router is a practical bridge** β keyword-based routing achieves
|
| 609 |
domain specialization without MoE complexity
|
|
|
|
|
|
|
| 610 |
|
| 611 |
### 11.3 Future Work
|
| 612 |
|
| 613 |
+
1. **Router Training**: Implement supervised or classification-based router
|
| 614 |
+
training for optimal expert selection
|
| 615 |
+
2. **Hidden Gate Mode**: Fix hidden gate computation to enable better
|
| 616 |
+
initialization
|
| 617 |
+
3. **More Experts**: Scale to 4 experts (coding, math, chat, medical) for
|
| 618 |
+
broader domain coverage
|
| 619 |
+
4. **Larger Base Models**: Apply pipeline to Qwen2.5-7B or 14B
|
| 620 |
+
5. **GGUF Q4_K_M Quantization**: Quantize MoE model for lower VRAM usage
|
| 621 |
+
6. **Dynamic Router**: Use prompt embeddings for more nuanced routing
|
| 622 |
+
than keyword matching
|
| 623 |
|
| 624 |
---
|
| 625 |
|
| 626 |
## 12. References
|
| 627 |
|
| 628 |
+
1. Shazeer, N., et al. "Outrageously Large Neural Networks: The Sparsely-Gated
|
| 629 |
+
Mixture-of-Experts Layer." *ICLR 2017*.
|
| 630 |
+
2. Fedus, W., Zoph, B., & Shazeer, N. "Switch Transformers: Scaling to Trillion
|
| 631 |
+
Parameter Models with Simple and Efficient Sparsity." *JMLR 2022*.
|
| 632 |
+
3. Jacobs, R. A., Jordan, M. I., Nowlan, S. J., & Hinton, G. E. "Adaptive
|
| 633 |
+
Mixtures of Local Experts." *Neural Computation 1991*.
|
| 634 |
4. Jiang, A. Q., et al. "Mixtral of Experts." *arXiv:2401.04088*, 2024.
|
| 635 |
5. Yang, A., et al. "Qwen2 Technical Report." *arXiv:2407.10671*, 2024.
|
| 636 |
+
6. Dai, D., et al. "DeepSeekMoE: Towards Ultimate Expert Specialization in
|
| 637 |
+
Mixture-of-Experts Language Models." *arXiv:2401.06066*, 2024.
|
| 638 |
+
7. Hu, E. J., et al. "LoRA: Low-Rank Adaptation of Large Language Models."
|
| 639 |
+
*ICLR 2022*.
|
| 640 |
+
8. Arcee AI. "mergekit: Tools for Merging Pretrained Large Language Models."
|
| 641 |
+
*GitHub: arcee-ai/mergekit*, 2024.
|
| 642 |
+
9. Labonne, M. "Create a Frankenstein MoE with mergekit." *HuggingFace Blog*,
|
| 643 |
+
2024.
|
| 644 |
+
10. Gerganov, G. "llama.cpp: LLM Inference in C/C++." *GitHub: ggerganov/llama.cpp*,
|
| 645 |
+
2023.
|
| 646 |
|
| 647 |
---
|
| 648 |
|