Hayula: A Local-First AI Infrastructure on Consumer Hardware — Architecture, Services, and Lessons from Production Deployment — Hayula Research
Hayula AI Lab
Abstract
Consumer hardware is sufficient but not trivial. The M2 Ultra's unified memory architecture is particularly well-suited to AI inference. The primary limitation is not compute but memory bandwidth when serving multiple concurrent large models.
Files
| File | Description |
|---|---|
paper.md |
Full paper (Markdown) |
README.md |
This model card |
Citation
@techreport{hayulalab2026hayulaecosysteminfrastructure,
title={Hayula: A Local-First AI Infrastructure on Consumer Hardware — Architecture, Services, and Lessons from Production Deployment — Hayula Research},
author={Hayula AI Lab},
year={2026},
url={https://huggingface.co/hayulalab/hayula-ecosystem-infrastructure-paper}
}
hayulalab — Open Source AI Research
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support