bupalinyu commited on
Commit
8da036e
·
verified ·
1 Parent(s): 16e0668

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +19 -9
README.md CHANGED
@@ -30,9 +30,20 @@ pipeline_tag: text-generation
30
  **Edge0-8b-a1b** — an 8B MoE LLM that runs at viable speed on portable devices in under **1.0 GiB of active memory** (1/4 of its 4.2 GB weight footprint),
31
  via the [edge0](https://github.com/Edge0-AI/edge0) streaming inference framework.
32
 
33
- The key is streaming: experts are memory-mapped and fetched from SSD only as routed,
34
- so RAM holds just the active weights. What makes that viable — instead of stalling like plain parameter offloading —
35
- is a trained prerouter head that **predicts the next token's expert routing one step ahead**, hiding storage latency behind compute.
 
 
 
 
 
 
 
 
 
 
 
36
 
37
 
38
  > **Preview status:** this is an early preview release of the edge0
@@ -54,13 +65,12 @@ The LoRA and prerouter adapters are co-located with the base checkpoint
54
  and load automatically — this repository is a complete, ready-to-run
55
  model directory for `edge0`.
56
 
57
- ## Quality
58
 
59
- All benchmarks were run by us with
60
- [OpenCompass](https://github.com/open-compass/opencompass) under identical
61
- settings and parameters for both models. The loss of the edge0 pipeline
62
- (int4 + adapters) relative to the fp16 base model is small: **2.8 points on average**, with MMLU-Pro
63
- above the base (max 100):
64
 
65
  | Benchmark | edge0-8b (int4) | Ling 3.0 tiny (fp16) |
66
  |---|---:|---:|
 
30
  **Edge0-8b-a1b** — an 8B MoE LLM that runs at viable speed on portable devices in under **1.0 GiB of active memory** (1/4 of its 4.2 GB weight footprint),
31
  via the [edge0](https://github.com/Edge0-AI/edge0) streaming inference framework.
32
 
33
+ Three mechanisms make this work:
34
+
35
+ - **SSD expert offload**: expert weights are streamed from storage on
36
+ demand — fetched only as routed, so RAM holds just the active
37
+ weights. Peak memory is bounded by the active set, not the
38
+ parameter count.
39
+ - **Prerouter**: a trained head predicts expert routing one step
40
+ ahead, so expert loads overlap the forward pass instead of stalling
41
+ it — **up to +59%** decode throughput; the gain grows with storage
42
+ latency, model size, and routed width *K*.
43
+ - **Recover-LoRA**: the int4 base is frozen and LoRA adapters are
44
+ trained by distillation from the FP teacher, recovering most of the
45
+ quantization loss at 4-bit (see Quality below). Adapters stay
46
+ unmerged: one read-only base serves multiple adapter sets.
47
 
48
 
49
  > **Preview status:** this is an early preview release of the edge0
 
65
  and load automatically — this repository is a complete, ready-to-run
66
  model directory for `edge0`.
67
 
68
+ ## Quality (self-evaluation)
69
 
70
+ Internal self-evaluation of this checkpoint (int4 + adapters) relative to
71
+ the fp16 base model — the loss of the edge0 pipeline is small: **2.8
72
+ points on average**, with MMLU-Pro above the base (max 100, all
73
+ self-run):
 
74
 
75
  | Benchmark | edge0-8b (int4) | Ling 3.0 tiny (fp16) |
76
  |---|---:|---:|