--- license: apache-2.0 library_name: coreml pipeline_tag: text-classification base_model: convaiinnovations/laya tags: - coreml - laya - apple-silicon - decision-model - local-ai - modernbert --- # laya-coreml **Laya typed decisions on Apple Silicon, using CPU + GPU.** This is a portable Core ML bundle for [laya-coreml](https://github.com/mizorewww/laya-coreml), converted from [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya). It outputs `choice`, `score`, and `noul` probabilities with **zero generated tokens**. Inference needs no PyTorch, Transformers, MLX, remote code, or cloud API. ## Run Apple Silicon, macOS 15+, Python 3.11–3.13. Tested on M3 Max / macOS 27.2. ```bash pip install laya-coreml ``` ```python import laya_coreml as laya agent = laya.load("aac6fef/laya-coreml") # Download once; Core ML runs locally. result = agent.predict( "The customer asks for a refund of a duplicate payment.", {"refund": {"type": "noul", "instructions": "Does the customer request a refund?"}}, ) print(result["answers"]) ``` To download explicitly and then run entirely offline: ```bash hf download aac6fef/laya-coreml --local-dir models/laya pip install 'laya-coreml[demo]' laya-coreml-snake --model models/laya --fps 12 ``` Use `laya.load("aac6fef/laya-coreml", local_files_only=True)` for a cached snapshot or pass a local directory. Use `revision=""` to pin a remote revision. ## Format and fidelity This FP16 export retains the original model architecture and decision schema. The enumerated-length GPU export is the validated general-purpose configuration. The complete 63-question fixture agrees with upstream selected answers, with 100 stable repeated calls. The release bundle is checked again after packaging. The exported capacity is **512 total tokens**, batch **1**, and **32** option slots. Questions/options and state share this budget. The ANE short exports reject over-capacity prompts. Snake uses planner features and a visible optional cycle safety shield; survival is not a claim of unaided game intelligence. `coreml_config.json` records shapes, source revisions and per-file SHA256 checksums. `validation.json` contains the packaging-time validation. Port fidelity on this regression suite does not establish general task accuracy or preserved calibration on arbitrary inputs. ## Performance and limits The multilingual **ANE L96 FP16** runtime measured **4.98 / 5.31 ms P50 / P95** for one short question on M3 Max; W8 measured **4.88 / 5.23 ms**. Whole-system energy per decision improved **2.78× / 3.19×**, respectively, against compiled MLX FP16 in that experiment. Those numbers apply to the named short ANE variants, not every bundle, long contexts, or complete Snake frames. The requested 10× improvement was not achieved. [Measurements and scope](https://github.com/mizorewww/laya-coreml/blob/main/docs/ANE_BENCHMARKS.md) · [General Core ML benchmarks](https://github.com/mizorewww/laya-coreml/blob/main/BENCHMARKS.md) · [Snake demo](https://github.com/mizorewww/laya-coreml/blob/main/docs/SNAKE_DEMO.md). ## Provenance - Original checkpoint: `convaiinnovations/laya` at `c5d78730f3493e4fe16d61507ef4b78eef7318cf`. - Original weights SHA256: `891102d372688fc2a094dac56a384bc537b87c63f21f9f3dac0be2b7cbc8d86c`. - Upstream implementation: [NandhaKishorM/laya](https://github.com/NandhaKishorM/laya), commit `6a5819129eb220570792e417e49723d697efd76f`. - Original models and code are by Convai Innovations and contributors, Apache-2.0. - Independent conversion; not an official Convai Innovations or Apple release. See `LICENSE` and `NOTICE`. Model quality and task/language limitations originate with Laya; this runtime is an inference port, not a newly trained decision model.