bupalinyu commited on
Commit
130ffec
·
verified ·
1 Parent(s): 54d130b

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +5 -5
README.md CHANGED
@@ -30,7 +30,7 @@ pipeline_tag: text-generation
30
  </div>
31
 
32
 
33
- **Edge0-8b-a1b** — an 8B MoE LLM that runs at viable speed on portable devices in under **1.0 GiB of active memory** (1/4 of its 4.2 GB weight footprint),
34
  via the [edge0](https://github.com/Edge0-AI/edge0) streaming inference framework.
35
 
36
 
@@ -42,10 +42,10 @@ via the [edge0](https://github.com/Edge0-AI/edge0) streaming inference framework
42
 
43
  - **Runs in phone-class memory**: the full 4-bit checkpoint stays on
44
  storage and experts are streamed on demand, so only the active
45
- weights are in RAM — under **1.0 GiB**, with no sharding and no
46
  upfront download of the weights into memory.
47
- - **Fast enough for interactive use**: 23.9–25.3 tok/s decode on a
48
- Mac mini M4 Pro; long prompts fill in at 500–1428 tok/s.
49
  - **Quality kept after quantization**: Recover-LoRA distillation keeps
50
  the int4 model within **2.8 points** of its fp16 base (and above it
51
  on MMLU-Pro).
@@ -130,7 +130,7 @@ Expert weights stream from SSD on demand and are not resident.
130
  - The MLX backend currently targets Apple Silicon; other backends are
131
  on the edge0 roadmap.
132
  - Long contexts grow the KV cache (≈3.3 GiB at 3.3k tokens); use
133
- shorter contexts to keep peak memory at 1.0 GiB.
134
 
135
  ## Quick start
136
 
 
30
  </div>
31
 
32
 
33
+ **Edge0-8b-a1b** — an 8B MoE LLM that runs at viable speed on a phone in under **1 GiB of active memory** (1/4 of its 4.2 GB weight footprint),
34
  via the [edge0](https://github.com/Edge0-AI/edge0) streaming inference framework.
35
 
36
 
 
42
 
43
  - **Runs in phone-class memory**: the full 4-bit checkpoint stays on
44
  storage and experts are streamed on demand, so only the active
45
+ weights are in RAM — under **1 GiB**, with no sharding and no
46
  upfront download of the weights into memory.
47
+ - **Fast enough for interactive use**: 25 tok/s decode on a phone;
48
+ long prompts fill in at 1400 tok/s.
49
  - **Quality kept after quantization**: Recover-LoRA distillation keeps
50
  the int4 model within **2.8 points** of its fp16 base (and above it
51
  on MMLU-Pro).
 
130
  - The MLX backend currently targets Apple Silicon; other backends are
131
  on the edge0 roadmap.
132
  - Long contexts grow the KV cache (≈3.3 GiB at 3.3k tokens); use
133
+ shorter contexts to keep peak memory at 1 GiB.
134
 
135
  ## Quick start
136