--- license: apache-2.0 language: - en - hi pipeline_tag: text-generation tags: - text-to-json - home-automation - robotics - hinglish - tiny-model - from-scratch - edge-ai - numpy - raspberry-pi - conditional-rules - experimental datasets: - sraivante/home-commands-json-v3 metrics: - exact_match model-index: - name: superfast-tiny-home-robotics-json-1m-v3 results: - task: type: text-generation name: instruction to JSON device command dataset: type: sraivante/home-commands-json-v3 name: home-commands-json-v3 (1,500 held-out rows, greedy decode) metrics: - type: exact_match value: 96.73 name: exact match (all) verified: false - type: exact_match value: 96.89 name: exact match (plain commands, n=900) verified: false - type: exact_match value: 96.50 name: exact match (conditional rules, n=600) verified: false --- # Superfast Tiny Home Robotics JSON 1M v3 A **1,143,328-parameter character-level transformer, trained from scratch**, that turns an English or Hinglish sentence into one JSON device command, including **conditional rules** ("if temperature goes above 42...", "roz shaam 7 baje..."). Weights are **2.1 MB (fp16)**. Inference is **numpy only**: no PyTorch, no tokenizer library, no internet, no GPU. ```text "bhai bedroom ka pankha high pe chala do" -> {"activity":"air","subject":"bedroom_ceiling_fan","action":"HIGH"} "agar room ka temperature 42 se upar jaye to bedroom ka heater band kar do" -> {"activity":"heating","subject":"bedroom_room_heater","action":"OFF", "when":{"metric":"temperature","op":">","value":42}} "roz shaam 7 baje garden ki batti chala do" -> {"activity":"light","subject":"garden_light","action":"ON", "when":{"metric":"time","op":"==","value":"19:00"}} "when soil moisture drops below 20 start the garden sprinkler" -> {"activity":"garden","subject":"garden_sprinkler","action":"ON", "when":{"metric":"soil_moisture","op":"<","value":20}} ``` This is the v3 of the series. Compared with [v1](https://huggingface.co/sraivante/superfast-tiny-home-robotics-json-1m-v1) (30 commands, 10 device types, transformers runtime) it covers **1,004 devices, 12 actions, 7 sensor metrics for rules**, and ships with a numpy engine plus a web UI and Docker image. | Release | Value | | --- | --- | | Publisher | sraivante | | Release date | 2026-09-26 | | Version | v3.0.0 | | Training data | [home-commands-json-v3](https://huggingface.co/datasets/sraivante/home-commands-json-v3) (synthetic, templated) | | Architecture | GPT-style causal decoder, char-level vocab (78 symbols), d=160, 4 layers, 4 heads, FFN 512, context 360 | | Parameters | 1,143,328 | | Weights | `model/model_v3_fp16.npz`, 2,146,783 bytes | | Runtime deps | `numpy` (plus `flask` for the optional server) | | License | Apache-2.0 | ## Raspberry Pi 5 edge-device performance Measured **2026-09-26** on a **Raspberry Pi 5 Model B, 16 GB RAM**, Debian 13.2/aarch64, Python 3.13.5. Measurement update: `edge-rpi5-20260926.1`; weights and training data are unchanged. | Threads | Median ms | p95 ms | p99 ms | Serial requests/s | Process RSS MiB | CPU °C min–mean–max | Timed correct | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | 1 | 410.38 | 553.20 | 558.65 | 2.53 | 43.77 | 59.0–60.9–63.4 | 500/500 | | 4 | 393.10 | 550.26 | 553.48 | 2.55 | 44.67 | 60.0–62.0–64.5 | 500/500 | Batch 1; 20 warmups and 500 timed calls per row, cycling over 10 frozen model-specific inputs. **Each timed reading has a matching CPU-temperature sample**, captured immediately afterward outside the latency timer; the table shows min–mean–max temperature. RSS includes Python/runtime overhead. Timings include local text-to-output inference and exclude loading, SSH/HTTP, speech recognition and actuation. Throughput is serial reciprocal mean. Runs used one and four CPU/BLAS threads sequentially. These synthetic workloads do not establish cross-model superiority or robot-control reliability. On-device quality check: **192/200** on training-set diagnostic, NOT held-out accuracy. For v3 this is explicitly a training-set diagnostic, not held-out accuracy; the earlier 96.73% result was not rerun. [Full Pi methodology, temperature plot and reproduction](edge_benchmarks/rpi5-2026-09-26/REPORT.md) · [Every latency/temperature reading](edge_benchmarks/rpi5-2026-09-26/readings.csv) · [Raw measurements](edge_benchmarks/rpi5-2026-09-26/results) · [Versioned evaluation data](https://huggingface.co/datasets/sraivante/home-commands-json-v3/tree/edge-rpi5-20260926.1/edge_benchmarks/rpi5-2026-09-26). ## Why a tiny model and not an LLM A home robot or a Raspberry Pi behind a microphone does not need to know the capital of France. It needs to turn "pankha tez kar do" into a command in a few milliseconds, offline, on a CPU that is also doing wake-word detection. This model is that one component: the **parser** between speech-to-text and the actuator. Everything else (safety, permissions, device state) stays in deterministic code you control. Measured on this release, single request, Python 3.13, numpy 2.3, Intel Core i7-1360P, default threads: | Input | Latency | | --- | ---: | | plain command (about 30 chars) | 70 to 90 ms | | conditional rule (about 70 chars) | 130 to 150 ms | | process RSS, engine loaded + warm (Python + numpy included) | 73 MB | Latency is dominated by Python overhead per generated character, not by FLOPs. Raspberry Pi 5 measurements are now recorded above; Jetson Orin Nano remains unmeasured. The table here retains the earlier desktop measurements. ## Quick start ```bash pip install numpy huggingface_hub python -c "from huggingface_hub import snapshot_download; snapshot_download('sraivante/superfast-tiny-home-robotics-json-1m-v3', local_dir='superfast-tiny-home-robotics-json-1m-v3')" cd superfast-tiny-home-robotics-json-1m-v3 python engine.py "kitchen ki light band kar do" # {"activity":"light","subject":"kitchen_light","action":"OFF"} ``` `git clone https://huggingface.co/sraivante/superfast-tiny-home-robotics-json-1m-v3` also works. In Python: ```python import engine engine.init("model") # loads the npz once raw, confidence = engine.parse_scored("turn off the porch light") # raw -> '{"activity":"light","subject":"porch_light","action":"OFF"}' # confidence -> avg log-prob per generated char; near 0.0 = in-domain ``` Web UI + REST API (Docker, no GPU): ```bash docker compose up --build # open http://localhost:8080 curl -s localhost:8080/api/parse -H 'Content-Type: application/json' \ -d '{"text":"if aqi goes above 250 set the bedroom air purifier on high"}' ``` The API returns the command, `valid_json`, catalog `issues`, `confidence`, `latency_ms` and a single `ok` flag. **Act only on `ok: true`.** ## How it is meant to be used (examples/) The model is a parser. `model/devices-catalog.json` is the law. Every example below runs the output through catalog validation before anything moves. | Example | What it shows | | --- | --- | | `examples/quickstart.py` | text in, validated command out, off-domain input rejected | | `examples/rule_engine.py` | what to do with the `when` clause: store it, watch sensors, fire once per crossing | | `examples/mqtt_bridge.py` | publish to `home//set` for ESP32 / Tasmota / Zigbee2MQTT devices | | `examples/home_assistant.py` | map subjects to Home Assistant entities and call services locally | Run any of them from the repo root, all have a `--dry-run` or simulated mode: ```bash python examples/quickstart.py python examples/rule_engine.py python examples/mqtt_bridge.py --dry-run "bedroom ka pankha high pe chala do" python examples/home_assistant.py --dry-run "garden ki light band kar do" ``` ### The pipeline this slots into ```text mic -> wake word -> Whisper/Vosk (Hindi+English) -> THIS MODEL -> catalog check -> actuator ~100 ms ~0 ms GPIO / MQTT / ROS ``` Plain commands (no `when`) go straight to the actuator. Conditional commands are stored as rules and evaluated on every sensor tick. `rule_engine.py` output on a simulated day: ```text "agar temperature 35 se upar jaye to bedroom ka ac high kar do" RULE stored: bedroom_ac -> HIGH when temperature > 35 "roz subah 6 baje garden sprinkler chalu kar do" RULE stored: garden_sprinkler -> ON when time == 06:00 [sensor] temperature = 36 ACTUATE bedroom_ac HIGH (temperature=36 > 35) [sensor] time = 06:00 ACTUATE garden_sprinkler ON (time=06:00 == 06:00) ``` ### Robotics use For a mobile robot or a robotic arm the same pattern applies: swap the device catalog for your robot's capabilities (`arm_gripper: OPEN/CLOSE`, `base_motor: START/STOP/LOW/MEDIUM/HIGH`, `head_camera: ON/OFF/STATUS`), regenerate the data with the scripts in the dataset repo, retrain for about an hour on two CPU cores, and you have a command parser that fits in the L2 cache. The `when` clause maps directly to sensor-gated behaviours ("agar battery 15 se neeche jaye to charging dock pe jao"). ## Output schema ```json {"activity": "...", "subject": "...", "action": "...", "when": {"metric": "...", "op": ">|<|==", "value": 42}} ``` - **1,004 subjects** in `model/devices-catalog.json`, each with its `activity` (18 of them: light, power, air, access, security, air_quality, cooling, heating, entertainment, music, garden, kitchen, office, network, water, laundry, cleaning, aquarium) and its allowed actions. - **12 actions**: ON, OFF, START, STOP, LOW, MEDIUM, HIGH, OPEN, CLOSE, LOCK, UNLOCK, STATUS. - **7 metrics** for `when`: temperature, humidity, aqi, battery, water_level, soil_moisture, time (value `"HH:MM"`, op `==`). - `when` is optional. Rules in the training data are sensible by construction: high temperature turns cooling ON or heating OFF, never the reverse. ## Confidence guard `parse_scored` returns the mean log-probability of the generated characters. In-domain commands decode at about 0.0; gibberish or off-domain sentences score clearly lower and the server flags them (default threshold -0.006, env `HCAI_CONF_THRESHOLD`). Example: "what is the capital of france" scores -0.022 and is rejected. This is a useful guard, not a guarantee: command-shaped off-domain input can still slip through. ## Evaluation 1,500 held-out rows from the same generator (900 plain, 600 conditional), greedy decode, exact string match on the JSON: | Metric | Value | | --- | ---: | | exact match, all | 96.73 % | | exact match, plain | 96.89 % | | exact match, conditional | 96.50 % | | valid JSON | 100 % | Full numbers and 12 sample failures are in `model/metrics_v3.json`. The held-out rows themselves were not preserved in this release; the dataset repo ships the generators and the deterministic `prep_mix.py` (seed 108) that describes the train/eval construction. The preserved 340,000-row training mix differs from the shipped script's configured counts; exact reconstruction of the missing historical evaluation split is unverified. The new Pi diagnostic uses a frozen sample from the preserved training mix and does not re-estimate the historical held-out score. ## Honest limitations - **Synthetic, templated training data.** The accuracy above measures interpolation inside the template space, not the wild. Novel or indirect phrasing ("andhera kar do", "thoda tez") fails. - **Observed misses that pass validation.** "kitchen ki chimney high pe karo" returned `STATUS` in our tests, and that is an allowed action for the chimney, so the catalog check passed it. Treat STATUS on an imperative sentence with suspicion in your own guard. - **Char-level spelling errors** in subjects happen (see sample failures); the catalog check exists exactly for that. Reject anything with issues. - **No negation, no reject class.** "fan mat chalao" can become `ON`. - **One instruction, one command.** No multi-intent, no context or pronouns. - Time rules are one-shot `==` triggers; recurrence ("roz") is not encoded separately, your scheduler decides that. - Romanised Hindi only. Devanagari input is not in the vocabulary and is silently dropped by the tokenizer. ## Adapting to your devices 1. Edit the device tables in `generators/generate_dataset_v2.py` and the rule tables in `generators/generate_dataset_v3_conditional.py` (dataset repo). 2. Regenerate and mix: `python generate_dataset_v2.py && python generate_dataset_v3_conditional.py && python prep_mix.py` 3. Retrain (needs torch; about 80 minutes on two CPU cores, minutes on a GPU): `python training/train_v3.py train --seconds 4200 && python training/train_v3.py export` 4. Drop `model_v3_fp16.npz`, `model_v3_config.json` and the new `devices-catalog.json` into `model/`. ## Files | Path | What | | --- | --- | | `config.json` | release metadata; also the file Hugging Face uses to count downloads | | `model/model_v3_fp16.npz` | weights, fp16, 2.1 MB | | `model/model_v3_config.json` | vocab (`itos`) + architecture config | | `model/devices-catalog.json` | the 1,004 subjects with allowed actions | | `model/metrics_v3.json` | eval numbers and sample failures | | `engine.py` | numpy inference, kv-cached greedy decode | | `server.py`, `static/index.html` | Flask API + offline web UI | | `examples/` | quickstart, rule engine, MQTT bridge, Home Assistant bridge | | `training/train_v3.py` | full PyTorch training and export source | | `Dockerfile`, `docker-compose.yml` | 150 MB image, healthcheck included | ## License Apache-2.0 for the weights, code and catalog. The training data is fully synthetic and released under the same license in the dataset repo. No warranty: this model suggests commands, your code decides whether to obey. Copyright (c) 2026 sraivante applies only to original material and selection/arrangement contributed by the user; third-party ownership and licenses remain unchanged.