--- title: Laya API emoji: 🧭 colorFrom: indigo colorTo: purple sdk: gradio app_file: app.py pinned: false short_description: Typed decisions over any state, as HTTP --- # Laya API [Laya](https://huggingface.co/convaiinnovations/laya) (Apache-2.0, Convai Innovations) answers **typed questions** about a state in one forward pass and returns calibrated probabilities — `choice`, `score` (ordinal rubric) and `noul` (P(true)). It writes no prose, so it does not replace an LLM; it replaces the hand-written rules an app uses to route, classify and verify. This Space wraps it in three endpoints. It is a **Gradio** Space (Docker Spaces are a paid feature): HF runs `python app.py`, which serves FastAPI on port 7860 with a Gradio test page mounted at `/ui`. ## Endpoints | | | |---|---| | `GET /health` | what is loaded, whether the weights are in yet, whether auth is on | | `POST /v1/decide` | generic: `{state, questions}` → typed answers | | `POST /v1/route` | pick one of N candidates — returns the candidate's own `id` | | `GET /docs` | interactive OpenAPI page | | `GET /ui` | a page to try it by hand | Auth: send `x-laya-token: ` on every call when that secret is set. A Space URL is public, so set it (Space → Settings → Variables and secrets → **LAYA_TOKEN**). ### Route — the Tally OS case Only the user's question and the candidate report titles are sent. **No ledger figures leave the app**, which is what makes it safe to run off our own infrastructure. ```bash curl -s https://-laya-api.hf.space/v1/route \ -H 'content-type: application/json' -H "x-laya-token: $LAYA_TOKEN" \ -d '{ "question": "is saal ka total sales kitna hua", "candidates": [ {"id": 71, "label": "Revenue booked this financial year"}, {"id": 13, "label": "Today total sales and invoice count"}, {"id": 33, "label": "Last 12 months daybook grouped by month — sales and purchase"}, {"id": 73, "label": "Total purchases between two dates"} ] }' # → {"id":71,"label":"Revenue booked this financial year","p":0.87,"ranked":[…],"ms":128} ``` ### Decide — anything else ```bash curl -s https://-laya-api.hf.space/v1/decide \ -H 'content-type: application/json' -H "x-laya-token: $LAYA_TOKEN" \ -d '{ "state": "Party asks for 45 days credit; they have 3 bills over 90 days overdue.", "questions": { "risk": {"type": "score", "instructions": "How risky is extending credit?", "criteria": ["low", "watch", "high"]}, "needs_period": {"type": "noul", "instructions": "Does answering this need a date range?"} } }' ``` ## Configuration | Variable | Default | What it does | |---|---|---| | `LAYA_TOKEN` | *(unset)* | shared secret; when set, every call must send it | | `LAYA_MODEL` | `convaiinnovations/laya` | checkpoint (`…/laya-multilingual` for Hindi/Gujarati, `…/laya-typed-decisions` for workflows) | | `LAYA_BACKEND` | auto | `laya` (torch, any OS) or `laya_mlx` (Apple Silicon) | | `LAYA_DTYPE` | `float16` | `float32` matches upstream maths more closely | | `LAYA_CORS_ORIGINS` | *(none)* | comma-separated origins, only if a browser calls this directly | | `LAYA_EAGER` | `1` | load the weights at boot so the first real call is fast | ## Footprint Weights are ~843 MB (421M params, FP16) and the process holds ~1 GB; 2 GB RAM is comfortable. The free CPU Space (2 vCPU / 16 GB) fits it easily. A free Space **sleeps when idle** and has no persistent disk, so a cold start re-downloads the weights — expect 1–3 minutes on the first call after a nap, then normal speed. Hit `/health` to wake it before a demo. Latency on CPU is not the 7–13 ms quoted for an M3 Max GPU; measure `/health` → `ms` on a real call before designing around it. ## Running it elsewhere ```bash # local (any OS with Python 3.10+) pip install -r requirements.txt && python app.py # or: uvicorn app:app --port 8000 # on a Mac, with the MLX build instead pip install laya-mlx fastapi 'uvicorn[standard]' && LAYA_BACKEND=laya_mlx uvicorn app:app --port 8000 # any container host EXCEPT Hugging Face (Docker Spaces are paid there) docker build -t laya-api . && docker run -p 7860:7860 -e LAYA_TOKEN=... laya-api ```