hotdogs commited on
Commit
3164ccc
·
verified ·
1 Parent(s): 0e31eb6

Update README — PoC disclaimer

Browse files
Files changed (1) hide show
  1. README.md +46 -120
README.md CHANGED
@@ -1,151 +1,77 @@
1
  ---
 
2
  language:
3
  - en
4
  - th
5
- license: apache-2.0
6
  tags:
7
- - frankenmoe
 
 
8
  - qwen2.5
9
- - lora
10
- - peft
11
- - gguf
12
- - coding
13
- - math
14
- - chat
15
- - expert-models
16
- pipeline_tag: text-generation
17
- base_model: Qwen/Qwen2.5-1.5B-Instruct
18
- ---
19
-
20
- # FrankenMoE — Qwen2.5-1.5B Expert Models 🇹🇭
21
-
22
- **3 specialized LoRA fine-tuned experts** — coding, math, and chat — built from Qwen2.5-1.5B-Instruct with 13,000 curated training samples.
23
-
24
- > 🛑 MoE merge skipped (mergekit does not support Qwen2 MoE architecture).
25
- > ✅ Each expert is independently usable as LoRA adapter or GGUF.
26
-
27
  ---
28
 
29
- ## 📦 What's Inside
30
 
31
- | Expert | Domain | LoRA | GGUF (Q4_K_M) | Train Loss | Eval Loss |
32
- |--------|--------|------|---------------|------------|-----------|
33
- | **coding** | Python/Algorithm/SWE | 71 MB | 941 MB | 1.03 | - |
34
- | **math** | Mathematics/Proofs | 70 MB | 941 MB | 1.18 | - |
35
- | **chat** | Instruction Following | 74 MB | 941 MB | 1.23 | 1.27 |
36
 
37
- ---
38
-
39
- ## 🚀 Quick Start
40
 
41
- ### Option 1: LoRA with PEFT (Python)
42
 
43
- ```python
44
- from peft import PeftModel
45
- from transformers import AutoModelForCausalLM, AutoTokenizer
46
- import torch
47
 
48
- base = "Qwen/Qwen2.5-1.5B-Instruct"
49
- model = AutoModelForCausalLM.from_pretrained(base, torch_dtype=torch.bfloat16)
50
- model = PeftModel.from_pretrained(model, "hotdogs/frankenmoe", subfolder="coding")
51
- tokenizer = AutoTokenizer.from_pretrained("hotdogs/frankenmoe", subfolder="coding")
52
 
53
- prompt = "Write a Python function to reverse a linked list"
54
- inputs = tokenizer(prompt, return_tensors="pt")
55
- outputs = model.generate(**inputs, max_new_tokens=256)
56
- print(tokenizer.decode(outputs[0]))
57
- ```
58
 
59
- ### Option 2: GGUF with llama.cpp
60
 
61
- ```bash
62
- # Download
63
- wget https://huggingface.co/hotdogs/frankenmoe/resolve/main/coding/frankenmoe_coding-Q4_K_M.gguf
64
-
65
- # Run
66
- llama.cpp/build/bin/llama-cli \
67
- -m frankenmoe_coding-Q4_K_M.gguf \
68
- -p "Write a Python function to reverse a linked list" \
69
- -n 256
70
  ```
71
-
72
- ### Option 3: Ollama Modelfile
73
-
74
- ```dockerfile
75
- FROM ./frankenmoe_coding-Q4_K_M.gguf
76
- SYSTEM "You are a coding expert specialized in Python, algorithms, and software engineering."
77
  ```
78
 
79
- ---
80
-
81
- ## 🔧 Training Details
82
-
83
- | Parameter | Value |
84
- |-----------|-------|
85
- | Base Model | Qwen2.5-1.5B-Instruct |
86
- | Method | LoRA (r=16, alpha=32) |
87
- | Precision | bfloat16 (no 4-bit quantization) |
88
- | Dataset | 13,000 curated samples (coding: 5K, math: 3K, chat: 5K) |
89
- | Epochs | 2 per expert |
90
- | GPU | RTX 4060 Ti 16GB |
91
- | Framework | transformers + peft + trl |
92
- | Optimizer | AdamW (torch) |
93
- | NEFTune α | 5-7 |
94
 
95
- ---
 
 
 
 
 
96
 
97
- ## 📁 Repository Structure
98
 
99
  ```
100
- hotdogs/frankenmoe/
101
- ├── README.md
102
- ├── coding/
103
- │ ├── adapter_model.safetensors
104
- │ ├── adapter_config.json
105
- │ ├── tokenizer.json
106
- │ ├── tokenizer_config.json
107
- │ └── frankenmoe_coding-Q4_K_M.gguf
108
- ├── math/
109
- │ ├── adapter_model.safetensors
110
- │ ├── adapter_config.json
111
- │ ├── tokenizer.json
112
- │ ├── tokenizer_config.json
113
- │ └── frankenmoe_math-Q4_K_M.gguf
114
- └── chat/
115
- ├── adapter_model.safetensors
116
- ├── adapter_config.json
117
- ├── tokenizer.json
118
- ├── tokenizer_config.json
119
- └── frankenmoe_chat-Q4_K_M.gguf
120
  ```
121
 
122
- ---
123
-
124
- ## 📊 Training Logs
125
-
126
- | Expert | Steps | Train Loss | Final LR | Time |
127
- |--------|-------|-----------|----------|------|
128
- | coding | 564 | 1.03 | - | ~15 min |
129
- | math | 338 | 1.18 | - | ~15 min |
130
- | chat | 564 | 1.23 | - | ~28 min |
131
-
132
- ---
133
 
134
- ## ⚠️ Known Limitations
 
 
 
 
135
 
136
- - **No MoE routing** — experts are independent models, not a single MoE
137
- - **Small base model** (1.5B) — good for experimentation, limited for production
138
- - **Qwen2 architecture** — not compatible with mergekit MoE (only Qwen3 MoE supported)
139
 
140
- ---
 
 
 
 
141
 
142
- ## 📜 License
143
-
144
- Same as base model: Apache 2.0
145
 
146
  ---
147
-
148
- ## 🙏 Credits
149
-
150
- Trained by **UKA** (AI Agent) on FrankenMoE Pipeline v2.0
151
- Thai AI infrastructure — local GPU only, zero cloud dependency 🇹🇭
 
1
  ---
2
+ license: apache-2.0
3
  language:
4
  - en
5
  - th
 
6
  tags:
7
+ - moe
8
+ - test
9
+ - proof-of-concept
10
  - qwen2.5
11
+ - mergekit
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
12
  ---
13
 
14
+ # 🧪 FrankenMoE — Proof of Concept (NOT production)
15
 
16
+ **This is a technical experiment, not a useful model.**
 
 
 
 
17
 
18
+ ## ⚠️ Important Warning
 
 
19
 
20
+ This repository documents a **proof-of-concept** MoE pipeline. The model quality is **NOT good** — it produces incoherent / random outputs because:
21
 
22
+ 1. The router uses (no training)
23
+ 2. Experts were fine-tuned with only ~5K samples each
24
+ 3. Base model is only Qwen2.5-1.5B-Instruct
 
25
 
26
+ **Do NOT use this model for anything serious.** It exists purely to demonstrate that the FrankenMoE pipeline can be built end-to-end.
 
 
 
27
 
28
+ ## What We Actually Built
 
 
 
 
29
 
30
+ A working MoE pipeline from dense LoRA experts → GGUF:
31
 
 
 
 
 
 
 
 
 
 
32
  ```
33
+ Qwen2.5-1.5B-Instruct (base)
34
+ ├── Expert 0: Coding (LoRA fine-tuned)
35
+ ├── Expert 1: Math (LoRA fine-tuned)
36
+ └── Shared Expert: Base model
 
 
37
  ```
38
 
39
+ ## Key Technical Discoveries
 
 
 
 
 
 
 
 
 
 
 
 
 
 
40
 
41
+ | Discovery | Detail |
42
+ |---|---|
43
+ | mergekit 0.1.4 bug | param incompatible with transformers >= 4.40 — must patch |
44
+ | QwenMoE requirements | Exactly 1 shared expert + 2^n routed experts (2, 4, 8) |
45
+ | Tied embeddings fix | Qwen2.5 uses tied embeddings → must clone → before GGUF conversion, set |
46
+ | LoRA must be merged | Adapters must be before MoE assembly |
47
 
48
+ ## Repository Structure
49
 
50
  ```
51
+ 📦 frankenmoe_moe_v2-F16.gguf — MoE GGUF (fixed, has output.weight)
52
+ 📁 moe_full/ — Full safetensors model
53
+ 📁 coding/ math/ chat/ — Individual dense experts (LoRA + GGUF)
54
+ 📄 FrankenMoE_Academic_Paper.pdf — Research paper
55
+ 🐍 simple_router.py — Keyword-based router (functional alternative)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
56
  ```
57
 
58
+ ## Quick Test
 
 
 
 
 
 
 
 
 
 
59
 
60
+ ```bash
61
+ wget https://huggingface.co/hotdogs/frankenmoe/resolve/main/frankenmoe_moe_v2-F16.gguf
62
+ llama-cli -m frankenmoe_moe_v2-F16.gguf -p "Write a Python function"
63
+ # Output: Random/incoherent — this is expected! See warning above.
64
+ ```
65
 
66
+ ## Future: Real Model
 
 
67
 
68
+ The pipeline will be re-run with:
69
+ - Larger base model (Qwen2.5-7B/14B)
70
+ - Trained router (classification loss)
71
+ - More training data per domain
72
+ - 4 experts for proper 2^n routing
73
 
74
+ **Stay tuned — the real model is coming.**
 
 
75
 
76
  ---
77
+ Built by UKA 🇹🇭 | May 2026