MikeKuykendall commited on
Commit
6f8a736
·
verified ·
1 Parent(s): 850be8b

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +41 -0
README.md CHANGED
@@ -1,3 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  # GPT-OSS 20B MoE CPU Offload - GGUF
2
 
3
  **🚀 First Implementation of MoE CPU Offloading Technology**
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: WeOpenML/GPT-OSS-20B
4
+ tags:
5
+ - mixture-of-experts
6
+ - moe
7
+ - cpu-offload
8
+ - gguf
9
+ - llama.cpp
10
+ - shimmy
11
+ - memory-efficient
12
+ - first-implementation
13
+ library_name: llama.cpp
14
+ model_type: gpt-oss
15
+ quantized_by: MikeKuykendall
16
+ language:
17
+ - en
18
+ - multilingual
19
+ pipeline_tag: text-generation
20
+ widget:
21
+ - text: "Write a Python function for fibonacci sequence"
22
+ example_title: "Code Generation"
23
+ - text: "Explain quantum computing in simple terms"
24
+ example_title: "Explanation Task"
25
+ model-index:
26
+ - name: gpt-oss-20b-moe-cpu-offload-gguf
27
+ results:
28
+ - task:
29
+ type: text-generation
30
+ dataset:
31
+ type: cpu-offload-benchmark
32
+ name: MoE CPU Offloading (First Implementation)
33
+ metrics:
34
+ - type: vram_reduction
35
+ value: 99.9
36
+ name: VRAM Reduction %
37
+ - type: memory_usage_mb
38
+ value: 2
39
+ name: GPU Memory Usage (MB)
40
+ ---
41
+
42
  # GPT-OSS 20B MoE CPU Offload - GGUF
43
 
44
  **🚀 First Implementation of MoE CPU Offloading Technology**