MikeKuykendall commited on
Commit
2211fcb
·
verified ·
1 Parent(s): 4e46dc6

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +89 -152
README.md CHANGED
@@ -1,201 +1,138 @@
1
  ---
 
 
 
2
  tags:
3
- - pytorch
4
- - deepseek
5
  - mixture-of-experts
6
- - text-generation
7
- - cpu-offloading
8
  - gguf
9
- - llama-cpp
10
- - memory-efficient
11
- - local-inference
12
- - moe
 
13
  language:
14
  - en
15
- license: other
16
- model_type: deepseek
17
- inference: true
18
  pipeline_tag: text-generation
19
- library_name: transformers
20
  ---
21
 
22
- # DeepSeek MoE 16B with CPU Expert Offloading
23
 
24
- ## Model Description
25
 
26
- **DeepSeek MoE 16B CPU Offload** is a memory-optimized GGUF conversion of DeepSeek's MoE 16B model, enhanced with revolutionary CPU expert offloading technology. This enables running a 16.38 billion parameter Mixture of Experts model with minimal GPU memory requirements through innovative expert tensor offloading.
27
 
28
- ### Key Features
 
 
 
 
 
 
29
 
30
- - **🧠 Advanced Architecture**: 64 regular experts + 2 shared experts, 6 active per token
31
- - **💾 Minimal VRAM Usage**: CPU expert offloading dramatically reduces GPU memory requirements
32
- - **⚡ Efficient Inference**: Optimized for local deployment with acceptable load times (~40s)
33
- - **🔧 Production Ready**: Validated working implementation with coherent text generation
34
- - **📏 Reasonable Context**: 4K token context length for focused tasks
35
 
36
- ## Model Specifications
37
 
38
- | Specification | Value |
39
- |---------------|-------|
40
- | **Parameters** | 16.38B (total) |
41
- | **Architecture** | DeepSeek MoE with dual expert system |
42
- | **Expert Configuration** | 64 regular experts + 2 shared experts |
43
- | **Active Experts** | 6 per token |
44
- | **Context Length** | 4,096 tokens |
45
- | **Precision** | F16 |
46
- | **File Size** | 32.8GB (GGUF) |
47
- | **Base Model** | [deepseek-ai/deepseek-moe-16b-base](https://huggingface.co/deepseek-ai/deepseek-moe-16b-base) |
48
 
49
- ## Memory Requirements
50
 
51
- ### Traditional Inference (Estimated)
52
- - **Full GPU Loading**: ~33-35GB VRAM (based on model size)
53
- - **CPU RAM**: ~2GB
 
54
 
55
- ### With CPU Expert Offloading ⚡
56
- - **GPU VRAM**: Minimal (expert tensors offloaded to CPU)
57
- - **CPU RAM**: ~35GB (includes expert tensors)
58
- - **Memory Savings**: Significant VRAM reduction while maintaining performance
59
 
60
- ## Installation & Usage
61
 
62
- ### Prerequisites
 
 
63
 
64
- ```bash
65
- # Install required dependencies
66
- pip install llama-cpp-python
67
- # OR build llama.cpp with MoE CPU offloading support
68
- git clone https://github.com/ggerganov/llama.cpp
69
- cd llama.cpp
70
- make LLAMA_CUDA=1
71
- ```
72
-
73
- ### Download Model
74
 
75
  ```bash
76
- # Using HuggingFace CLI
77
  huggingface-cli download MikeKuykendall/deepseek-moe-16b-cpu-offload-gguf \
78
- deepseek-moe-16b-f16.gguf --local-dir ./models
 
79
  ```
80
 
81
- ### Basic Usage
82
 
83
- ```bash
84
- # Using llama.cpp with CPU expert offloading
85
- ./main -m ./models/deepseek-moe-16b-f16.gguf \
86
- --cpu-moe \
87
- --prompt "What is mixture of experts in AI?" \
88
- --n-predict 100
89
- ```
90
 
91
- ### Python Integration
92
-
93
- ```python
94
- from llama_cpp import Llama
95
-
96
- # Initialize model with CPU expert offloading
97
- llm = Llama(
98
- model_path="./models/deepseek-moe-16b-f16.gguf",
99
- n_ctx=4096,
100
- cpu_moe=True, # Enable CPU expert offloading
101
- verbose=True
102
- )
103
 
104
- # Generate text
105
- response = llm("What is mixture of experts in AI?", max_tokens=100)
106
- print(response['choices'][0]['text'])
107
  ```
108
 
109
- ## Performance Benchmarks
110
-
111
- ### Model Loading
112
- - **Load Time**: ~40 seconds (including expert tensor initialization)
113
- - **Memory Initialization**: Expert tensors successfully moved to CPU
114
- - **Architecture Detection**: 64+2 expert configuration properly recognized
115
-
116
- ### Generation Quality
117
- - **Coherence**: Maintains logical flow and context understanding
118
- - **Technical Accuracy**: Produces contextually appropriate responses
119
- - **Response Length**: Generates coherent text within token limits
120
- - **Expert Activation**: All 6 active experts properly utilized
121
-
122
- ### Memory Efficiency
123
- - **Expert Tensor Offloading**: ✅ All expert tensors successfully moved to CPU
124
- - **GPU Memory**: Minimal usage with CPU offloading enabled
125
- - **Total Model Size**: 32.8GB efficiently distributed between GPU and CPU
126
 
127
- ## Technical Architecture
128
-
129
- ### Unique Dual Expert System
130
- DeepSeek MoE implements an innovative architecture combining:
131
-
132
- 1. **64 Regular Experts**: Standard MoE experts for specialized processing
133
- 2. **2 Shared Experts**: Always-active experts for common patterns
134
- 3. **6 Active Per Token**: 6 experts activated for each token (highest among tested models)
135
-
136
- ### Expert Tensor Distribution
137
- ```
138
- Expert Tensors: ffn_gate_exps.weight, ffn_down_exps.weight, ffn_up_exps.weight
139
- Shared Experts: shared_expert.gate_proj.weight, shared_expert.up_proj.weight, shared_expert.down_proj.weight
140
- Buffer Override: All expert tensors moved to CPU for memory efficiency
 
 
 
 
141
  ```
142
 
143
- ## Comparison with Other MoE Models
144
-
145
- | Model | Parameters | Experts | Active/Token | VRAM Reduction | Context |
146
- |-------|------------|---------|--------------|----------------|---------|
147
- | **DeepSeek MoE 16B** | 16.38B | 64+2 shared | 6 | High | 4K |
148
- | GPT-OSS 20B | 20B | 32 | 4 | 99.9% | 131K |
149
- | Phi-3.5-MoE 41.9B | 41.9B | 16 | 2 | 97.1% | 131K |
150
 
151
- ## Limitations
 
 
 
 
152
 
153
- 1. **Context Length**: 4K tokens (shorter than other tested models)
154
- 2. **Generation Patterns**: May exhibit some repetitive patterns requiring parameter tuning
155
- 3. **Expert Complexity**: Dual expert system may require specialized handling for optimal performance
156
- 4. **Load Time**: ~40 second initialization due to large model size and expert configuration
 
157
 
158
- ## Use Cases
159
 
160
- ### Ideal For:
161
- - **Local AI Development**: Efficient local inference for development and testing
162
- - **Memory-Constrained Environments**: Systems with limited GPU VRAM but adequate CPU RAM
163
- - **Research Applications**: Studying MoE architectures and expert activation patterns
164
- - **Educational Purposes**: Understanding dual expert system architectures
165
 
166
- ### Best Practices:
167
- - Use with sufficient CPU RAM (>35GB) for optimal performance
168
- - Consider parameter tuning to reduce repetitive generation patterns
169
- - Monitor expert activation patterns for insights into model behavior
170
- - Combine with other models for diverse inference capabilities
171
 
172
- ## Model Card Authors
173
-
174
- **MikeKuykendall** - Conversion, optimization, and CPU offloading implementation
175
 
176
  ## Citation
177
 
178
- If you use this model in your research, please cite:
179
-
180
  ```bibtex
181
- @misc{deepseek-moe-16b-cpu-offload,
182
- title={DeepSeek MoE 16B with CPU Expert Offloading},
183
- author={MikeKuykendall},
184
- year={2025},
185
- url={https://huggingface.co/MikeKuykendall/deepseek-moe-16b-cpu-offload-gguf}
186
  }
187
  ```
188
 
189
- ## License
190
-
191
- This model follows the original DeepSeek license terms. Please refer to the [base model](https://huggingface.co/deepseek-ai/deepseek-moe-16b-base) for complete licensing information.
192
-
193
- ## Acknowledgments
194
-
195
- - **DeepSeek Team**: Original model architecture and training
196
- - **GGML/llama.cpp Community**: GGUF format and inference optimization
197
- - **MoE CPU Offloading Research**: Breakthrough memory optimization techniques
198
-
199
  ---
200
 
201
- *Model converted and optimized as part of comprehensive MoE CPU offloading research - October 2025*
 
1
  ---
2
+ license: apache-2.0
3
+ license_link: https://huggingface.co/deepseek-ai/deepseek-moe-16b-base/blob/main/LICENSE
4
+ base_model: deepseek-ai/deepseek-moe-16b-base
5
  tags:
6
+ - moe
 
7
  - mixture-of-experts
 
 
8
  - gguf
9
+ - llama.cpp
10
+ - shimmy
11
+ - rust
12
+ - cpu-offload
13
+ quantized_by: MikeKuykendall
14
  language:
15
  - en
16
+ - zh
 
 
17
  pipeline_tag: text-generation
18
+ library_name: llama.cpp
19
  ---
20
 
21
+ # DeepSeek MoE 16B Base - F16 GGUF with MoE CPU Offloading Support
22
 
23
+ F16 GGUF conversion of [deepseek-ai/deepseek-moe-16b-base](https://huggingface.co/deepseek-ai/deepseek-moe-16b-base) with Rust bindings for llama.cpp's MoE CPU offloading functionality.
24
 
25
+ ## Model Details
26
 
27
+ - **Base Model**: [deepseek-ai/deepseek-moe-16b-base](https://huggingface.co/deepseek-ai/deepseek-moe-16b-base)
28
+ - **Format**: GGUF F16 precision
29
+ - **File Size**: 31GB
30
+ - **Parameters**: 16.4B total (2.8B active per token)
31
+ - **Architecture**: 28 layers, 64 regular experts + 2 shared experts, 6 active per token
32
+ - **Context Length**: 4K tokens
33
+ - **Converted by**: [MikeKuykendall](https://huggingface.co/MikeKuykendall)
34
 
35
+ ## MoE CPU Offloading
 
 
 
 
36
 
37
+ This model supports **MoE CPU offloading** via llama.cpp (implemented in [PR #15077](https://github.com/ggml-org/llama.cpp/pull/15077)). Shimmy provides Rust bindings for this functionality, enabling:
38
 
39
+ - **VRAM Reduction**: 92.5% (30.1GB → 2.3GB measured on GH200)
40
+ - **Performance Trade-off**: 4.1x slower generation (26.8 → 6.5 TPS)
41
+ - **Use Case**: Running 16B parameter MoE on consumer GPUs (<4GB VRAM)
 
 
 
 
 
 
 
42
 
43
+ ### Controlled Baseline (NVIDIA GH200, N=3)
44
 
45
+ | Configuration | VRAM | TPS | TTFT |
46
+ |---------------|------|-----|------|
47
+ | **GPU-only** | 30.1GB | 26.8 | 426ms |
48
+ | **CPU Offload** | 2.3GB | 6.5 | 1,643ms |
49
 
50
+ **Trade-off**: Memory for speed. Best for VRAM-constrained scenarios where generation speed is less critical than model size.
 
 
 
51
 
52
+ ### Unique Architecture
53
 
54
+ DeepSeek MoE uses a **dual-expert architecture** (64 regular + 2 shared experts), validated to work correctly with CPU offloading:
55
+ - Regular experts: `ffn_gate_exps.weight`, `ffn_down_exps.weight`, `ffn_up_exps.weight`
56
+ - Shared experts: `ffn_gate_shexp.weight`, `ffn_down_shexp.weight`, `ffn_up_shexp.weight`
57
 
58
+ ## Download
 
 
 
 
 
 
 
 
 
59
 
60
  ```bash
 
61
  huggingface-cli download MikeKuykendall/deepseek-moe-16b-cpu-offload-gguf \
62
+ --include "deepseek-moe-16b-f16.gguf" \
63
+ --local-dir ./models
64
  ```
65
 
66
+ ## Usage
67
 
68
+ ### llama.cpp (CPU Offloading)
 
 
 
 
 
 
69
 
70
+ ```bash
71
+ # Standard loading (requires ~32GB VRAM)
72
+ ./llama-server -m deepseek-moe-16b-f16.gguf -c 4096
 
 
 
 
 
 
 
 
 
73
 
74
+ # With MoE CPU offloading (requires ~3GB VRAM + 32GB RAM)
75
+ ./llama-server -m deepseek-moe-16b-f16.gguf -c 4096 --cpu-moe
 
76
  ```
77
 
78
+ ### Shimmy (Rust Bindings)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
79
 
80
+ ```bash
81
+ # Install Shimmy
82
+ cargo install --git https://github.com/Michael-A-Kuykendall/shimmy --features llama-cuda
83
+
84
+ # Standard loading
85
+ shimmy serve --model deepseek-moe-16b-f16.gguf
86
+
87
+ # With MoE CPU offloading
88
+ shimmy serve --model deepseek-moe-16b-f16.gguf --cpu-moe
89
+
90
+ # Query the API
91
+ curl http://localhost:11435/api/generate \
92
+ -d '{
93
+ "model": "deepseek-moe-16b",
94
+ "prompt": "Explain the architecture of DeepSeek MoE",
95
+ "max_tokens": 256,
96
+ "stream": false
97
+ }'
98
  ```
99
 
100
+ ## Performance Notes
 
 
 
 
 
 
101
 
102
+ **Standard GPU Loading**:
103
+ - VRAM: 30.1GB
104
+ - Speed: 26.8 TPS
105
+ - Latency: 426ms TTFT
106
+ - Use when: VRAM is plentiful, speed is critical
107
 
108
+ **CPU Offloading**:
109
+ - VRAM: 2.3GB (92.5% reduction)
110
+ - Speed: 6.5 TPS (4.1x slower)
111
+ - Latency: 1,643ms TTFT
112
+ - Use when: Limited VRAM, speed less critical
113
 
114
+ ## Original Model
115
 
116
+ - **Developers**: DeepSeek AI
117
+ - **License**: Apache 2.0
118
+ - **Paper**: [DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models](https://arxiv.org/abs/2401.06066)
119
+ - **Languages**: English, Chinese
 
120
 
121
+ ## Technical Validation
 
 
 
 
122
 
123
+ Full validation report with controlled baselines: [Shimmy MoE CPU Offloading Technical Report](https://github.com/Michael-A-Kuykendall/shimmy/blob/feat/moe-cpu-offload/docs/MOE-TECHNICAL-REPORT.md)
 
 
124
 
125
  ## Citation
126
 
 
 
127
  ```bibtex
128
+ @article{dai2024deepseekmoe,
129
+ title={DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models},
130
+ author={Dai, Damai and others},
131
+ journal={arXiv preprint arXiv:2401.06066},
132
+ year={2024}
133
  }
134
  ```
135
 
 
 
 
 
 
 
 
 
 
 
136
  ---
137
 
138
+ *GGUF conversion and MoE offloading validation by [MikeKuykendall](https://huggingface.co/MikeKuykendall)*