bupalinyu's picture
Upload README.md with huggingface_hub
fbee063 verified
|
Raw History Blame
3.89 kB
metadata
license: apache-2.0
library_name: mlx
tags:
  - moe
  - edge-inference
  - prerouter
  - lora
  - ssd-offload
base_model:
  - inclusionAI/Ling-3.0-tiny-base
pipeline_tag: text-generation
edge0

Edge0-8b-a1b Preview

An 8B-class sparse MoE that runs on a phone β€” 1.0 GiB of active memory, experts streamed from SSD.

GitHub Hugging Face Hugging Face License

This repository hosts the edge0-8b checkpoint of the edge0 streaming MoE inference framework. The full 4-bit checkpoint (β‰ˆ4.2 GB) stays on storage; edge0 mmaps it and streams MoE experts from SSD on demand, with a trained prerouter head that predicts the next token's expert routing one step ahead so expert loads hide completely behind the forward pass. The result: an 8B-class MoE with β‰ˆ1.0 GiB of active memory β€” a phone-class memory budget, with no upfront weight download into RAM and no model sharding.

Preview status: this is an early preview release of the edge0 pipeline. The checkpoint ships as int4 quantization plus LoRA and prerouter adapters trained for this framework.

Model summary

Base model inclusionAI Ling 3.0 tiny (bailing hybrid, MLA + MoE, β‰ˆ7.9B total / β‰ˆ1.2B active)
Quantization 4-bit
Layers 24
Experts / active per token 128 / 8 (K=8)
Framework edge0 (MLX backend)
Contents base checkpoint + lora_edge0_8b.safetensors + prerouter_edge0_8b.safetensors

The LoRA and prerouter adapters are co-located with the base checkpoint and load automatically β€” this repository is a complete, ready-to-run model directory for edge0.

Quality (self-evaluation)

Internal self-evaluation of this checkpoint (int4 + adapters) relative to the fp16 base model β€” the loss of the edge0 pipeline is small: 2.8 points on average, with MMLU-Pro above the base (max 100, all self-run):

Benchmark edge0-8b (int4) Base fp16
AIME 2026 63.3 73.3
HumanEval 91.5 92.7
GPQA-Diamond 70.7 71.2
MMLU-Pro 70.1 65.8
IFBench 53.9 60.6
Average 69.9 72.7

Performance

Measured with examples/bench.py on a Mac mini M4 Pro, 24 GB:

Decode speed Prefill throughput (cold / warm) Peak active memory*
23.9–25.3 tok/s 500 / 1428 tok/s 1.0 GiB

*Short contexts; long contexts add KV cache (β‰ˆ3.3 GiB at 3.3k tokens). Expert weights stream from SSD via mmap and are not resident.

Quick start

pip install -e 'git+https://github.com/Edge0-AI/edge0.git#egg=edge0[fetch]'

# Download this repository into a local directory
huggingface-cli download Edge0/Edge0-8b-a1b-preview --local-dir ./Edge0-8b-a1b-preview

# Run it
export EDGE0_8B_MODEL=$PWD/Edge0-8b-a1b-preview
edge0 chat --name edge0-8b --prompt "Introduce yourself"

# Or serve an OpenAI-compatible HTTP API
edge0 serve --name edge0-8b --port 8083

For full usage (Python API, streaming options, prerouter details), see the edge0 documentation.

License

Apache 2.0. See LICENSE.