File size: 3,783 Bytes
fff78b2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
---
license: apache-2.0
library_name: coreml
pipeline_tag: text-classification
base_model: convaiinnovations/laya
tags:
- coreml
- laya
- apple-silicon
- decision-model
- local-ai
- modernbert
---

# laya-coreml

**Laya typed decisions on Apple Silicon, using CPU + GPU.**
This is a portable Core ML bundle for [laya-coreml](https://github.com/mizorewww/laya-coreml),
converted from [convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya).
It outputs `choice`, `score`, and `noul` probabilities with **zero generated tokens**.
Inference needs no PyTorch, Transformers, MLX, remote code, or cloud API.

## Run

Apple Silicon, macOS 15+, Python 3.11–3.13. Tested on M3 Max / macOS 27.2.

```bash
pip install laya-coreml
```

```python
import laya_coreml as laya

agent = laya.load("aac6fef/laya-coreml")  # Download once; Core ML runs locally.
result = agent.predict(
    "The customer asks for a refund of a duplicate payment.",
    {"refund": {"type": "noul", "instructions": "Does the customer request a refund?"}},
)
print(result["answers"])
```

To download explicitly and then run entirely offline:

```bash
hf download aac6fef/laya-coreml --local-dir models/laya
pip install 'laya-coreml[demo]'
laya-coreml-snake --model models/laya --fps 12
```

Use `laya.load("aac6fef/laya-coreml", local_files_only=True)` for a cached snapshot or pass a
local directory. Use `revision="<Hub commit SHA>"` to pin a remote revision.

## Format and fidelity

This FP16 export retains the original model architecture and decision schema. The enumerated-length GPU export is the validated general-purpose configuration.

The complete 63-question fixture agrees with upstream selected answers, with 100 stable repeated calls. The release bundle is checked again after packaging.

The exported capacity is **512 total tokens**, batch **1**,
and **32** option slots. Questions/options and state share this budget.
The ANE short exports reject over-capacity prompts. Snake uses planner features and a
visible optional cycle safety shield; survival is not a claim of unaided game intelligence.

`coreml_config.json` records shapes, source revisions and per-file SHA256 checksums.
`validation.json` contains the packaging-time validation. Port fidelity on this regression
suite does not establish general task accuracy or preserved calibration on arbitrary inputs.

## Performance and limits

The multilingual **ANE L96 FP16** runtime measured **4.98 / 5.31 ms P50 / P95** for
one short question on M3 Max; W8 measured **4.88 / 5.23 ms**. Whole-system energy per
decision improved **2.78× / 3.19×**, respectively, against compiled MLX FP16 in that
experiment. Those numbers apply to the named short ANE variants, not every bundle,
long contexts, or complete Snake frames. The requested 10× improvement was not achieved.

[Measurements and scope](https://github.com/mizorewww/laya-coreml/blob/main/docs/ANE_BENCHMARKS.md)
· [General Core ML benchmarks](https://github.com/mizorewww/laya-coreml/blob/main/BENCHMARKS.md)
· [Snake demo](https://github.com/mizorewww/laya-coreml/blob/main/docs/SNAKE_DEMO.md).

## Provenance

- Original checkpoint: `convaiinnovations/laya` at `c5d78730f3493e4fe16d61507ef4b78eef7318cf`.
- Original weights SHA256: `891102d372688fc2a094dac56a384bc537b87c63f21f9f3dac0be2b7cbc8d86c`.
- Upstream implementation: [NandhaKishorM/laya](https://github.com/NandhaKishorM/laya),
  commit `6a5819129eb220570792e417e49723d697efd76f`.
- Original models and code are by Convai Innovations and contributors, Apache-2.0.
- Independent conversion; not an official Convai Innovations or Apple release.

See `LICENSE` and `NOTICE`. Model quality and task/language limitations originate
with Laya; this runtime is an inference port, not a newly trained decision model.