Text Generation
MLX
Safetensors
Mixture of Experts
edge-inference
prerouter
lora
ssd-offload
conversational
custom_code
4-bit precision
Instructions to use Edge0/Edge0-8B-A1B-preview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Edge0/Edge0-8B-A1B-preview with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Edge0/Edge0-8B-A1B-preview") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Edge0/Edge0-8B-A1B-preview with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Edge0/Edge0-8B-A1B-preview"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Edge0/Edge0-8B-A1B-preview" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use Edge0/Edge0-8B-A1B-preview with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Edge0/Edge0-8B-A1B-preview"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Edge0/Edge0-8B-A1B-preview" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Edge0/Edge0-8B-A1B-preview", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use Edge0/Edge0-8B-A1B-preview with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Edge0/Edge0-8B-A1B-preview"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Edge0/Edge0-8B-A1B-preview
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Edge0/Edge0-8B-A1B-preview with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Edge0/Edge0-8B-A1B-preview"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Edge0/Edge0-8B-A1B-preview" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Upload README.md
Browse files
README.md
CHANGED
|
@@ -25,6 +25,9 @@ pipeline_tag: text-generation
|
|
| 25 |
[](https://github.com/Edge0-AI/edge0)
|
| 26 |
[](https://huggingface.co/Edge0/Edge0-35b-a3b-preview)
|
| 27 |
[](https://huggingface.co/Edge0/Edge0-8b-a1b-preview)
|
|
|
|
|
|
|
|
|
|
| 28 |
[](https://github.com/Edge0-AI/edge0/blob/main/LICENSE)
|
| 29 |
|
| 30 |
</div>
|
|
@@ -165,3 +168,19 @@ For full usage (Python API, streaming options, prerouter details), see the
|
|
| 165 |
## License
|
| 166 |
|
| 167 |
Apache 2.0. See [LICENSE](https://github.com/Edge0-AI/edge0/blob/main/LICENSE).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
[](https://github.com/Edge0-AI/edge0)
|
| 26 |
[](https://huggingface.co/Edge0/Edge0-35b-a3b-preview)
|
| 27 |
[](https://huggingface.co/Edge0/Edge0-8b-a1b-preview)
|
| 28 |
+
[](https://www.modelscope.cn/models/Edge0/Edge0-35B-A3B-preview)
|
| 29 |
+
[](https://www.modelscope.cn/models/Edge0/Edge0-8B-A1B-preview)
|
| 30 |
+
[](https://arxiv.org/abs/2609.18063)
|
| 31 |
[](https://github.com/Edge0-AI/edge0/blob/main/LICENSE)
|
| 32 |
|
| 33 |
</div>
|
|
|
|
| 168 |
## License
|
| 169 |
|
| 170 |
Apache 2.0. See [LICENSE](https://github.com/Edge0-AI/edge0/blob/main/LICENSE).
|
| 171 |
+
|
| 172 |
+
## Citation
|
| 173 |
+
|
| 174 |
+
If you find Edge0 useful in your research, please cite our paper:
|
| 175 |
+
|
| 176 |
+
```bibtex
|
| 177 |
+
@misc{lin2026halfmemorywallserving,
|
| 178 |
+
title={The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction},
|
| 179 |
+
author={Yu Lin and Yiming Wang and Runyuan Cai and Hanze Liu and Xiaodong Zeng},
|
| 180 |
+
year={2026},
|
| 181 |
+
eprint={2609.18063},
|
| 182 |
+
archivePrefix={arXiv},
|
| 183 |
+
primaryClass={cs.AI},
|
| 184 |
+
url={https://arxiv.org/abs/2609.18063},
|
| 185 |
+
}
|
| 186 |
+
```
|