Superfast Tiny Home Robotics JSON 1M v3
A 1,143,328-parameter character-level transformer, trained from scratch, that turns an English or Hinglish sentence into one JSON device command, including conditional rules ("if temperature goes above 42...", "roz shaam 7 baje..."). Weights are 2.1 MB (fp16). Inference is numpy only: no PyTorch, no tokenizer library, no internet, no GPU.
"bhai bedroom ka pankha high pe chala do"
-> {"activity":"air","subject":"bedroom_ceiling_fan","action":"HIGH"}
"agar room ka temperature 42 se upar jaye to bedroom ka heater band kar do"
-> {"activity":"heating","subject":"bedroom_room_heater","action":"OFF",
"when":{"metric":"temperature","op":">","value":42}}
"roz shaam 7 baje garden ki batti chala do"
-> {"activity":"light","subject":"garden_light","action":"ON",
"when":{"metric":"time","op":"==","value":"19:00"}}
"when soil moisture drops below 20 start the garden sprinkler"
-> {"activity":"garden","subject":"garden_sprinkler","action":"ON",
"when":{"metric":"soil_moisture","op":"<","value":20}}
This is the v3 of the series. Compared with v1 (30 commands, 10 device types, transformers runtime) it covers 1,004 devices, 12 actions, 7 sensor metrics for rules, and ships with a numpy engine plus a web UI and Docker image.
| Release | Value |
|---|---|
| Publisher | sraivante |
| Release date | 2026-09-26 |
| Version | v3.0.0 |
| Training data | home-commands-json-v3 (synthetic, templated) |
| Architecture | GPT-style causal decoder, char-level vocab (78 symbols), d=160, 4 layers, 4 heads, FFN 512, context 360 |
| Parameters | 1,143,328 |
| Weights | model/model_v3_fp16.npz, 2,146,783 bytes |
| Runtime deps | numpy (plus flask for the optional server) |
| License | Apache-2.0 |
Raspberry Pi 5 edge-device performance
Measured 2026-09-26 on a Raspberry Pi 5 Model B, 16 GB RAM, Debian 13.2/aarch64, Python 3.13.5. Measurement update: edge-rpi5-20260926.1; weights and training data are unchanged.
| Threads | Median ms | p95 ms | p99 ms | Serial requests/s | Process RSS MiB | CPU °C min–mean–max | Timed correct |
|---|---|---|---|---|---|---|---|
| 1 | 410.38 | 553.20 | 558.65 | 2.53 | 43.77 | 59.0–60.9–63.4 | 500/500 |
| 4 | 393.10 | 550.26 | 553.48 | 2.55 | 44.67 | 60.0–62.0–64.5 | 500/500 |
Batch 1; 20 warmups and 500 timed calls per row, cycling over 10 frozen model-specific inputs. Each timed reading has a matching CPU-temperature sample, captured immediately afterward outside the latency timer; the table shows min–mean–max temperature. RSS includes Python/runtime overhead. Timings include local text-to-output inference and exclude loading, SSH/HTTP, speech recognition and actuation. Throughput is serial reciprocal mean. Runs used one and four CPU/BLAS threads sequentially. These synthetic workloads do not establish cross-model superiority or robot-control reliability.
On-device quality check: 192/200 on training-set diagnostic, NOT held-out accuracy. For v3 this is explicitly a training-set diagnostic, not held-out accuracy; the earlier 96.73% result was not rerun.
Full Pi methodology, temperature plot and reproduction · Every latency/temperature reading · Raw measurements · Versioned evaluation data.
Why a tiny model and not an LLM
A home robot or a Raspberry Pi behind a microphone does not need to know the capital of France. It needs to turn "pankha tez kar do" into a command in a few milliseconds, offline, on a CPU that is also doing wake-word detection. This model is that one component: the parser between speech-to-text and the actuator. Everything else (safety, permissions, device state) stays in deterministic code you control.
Measured on this release, single request, Python 3.13, numpy 2.3, Intel Core i7-1360P, default threads:
| Input | Latency |
|---|---|
| plain command (about 30 chars) | 70 to 90 ms |
| conditional rule (about 70 chars) | 130 to 150 ms |
| process RSS, engine loaded + warm (Python + numpy included) | 73 MB |
Latency is dominated by Python overhead per generated character, not by FLOPs. Raspberry Pi 5 measurements are now recorded above; Jetson Orin Nano remains unmeasured. The table here retains the earlier desktop measurements.
Quick start
pip install numpy huggingface_hub
python -c "from huggingface_hub import snapshot_download; snapshot_download('sraivante/superfast-tiny-home-robotics-json-1m-v3', local_dir='superfast-tiny-home-robotics-json-1m-v3')"
cd superfast-tiny-home-robotics-json-1m-v3
python engine.py "kitchen ki light band kar do"
# {"activity":"light","subject":"kitchen_light","action":"OFF"}
git clone https://huggingface.co/sraivante/superfast-tiny-home-robotics-json-1m-v3 also works.
In Python:
import engine
engine.init("model") # loads the npz once
raw, confidence = engine.parse_scored("turn off the porch light")
# raw -> '{"activity":"light","subject":"porch_light","action":"OFF"}'
# confidence -> avg log-prob per generated char; near 0.0 = in-domain
Web UI + REST API (Docker, no GPU):
docker compose up --build # open http://localhost:8080
curl -s localhost:8080/api/parse -H 'Content-Type: application/json' \
-d '{"text":"if aqi goes above 250 set the bedroom air purifier on high"}'
The API returns the command, valid_json, catalog issues, confidence,
latency_ms and a single ok flag. Act only on ok: true.
How it is meant to be used (examples/)
The model is a parser. model/devices-catalog.json is the law. Every example
below runs the output through catalog validation before anything moves.
| Example | What it shows |
|---|---|
examples/quickstart.py |
text in, validated command out, off-domain input rejected |
examples/rule_engine.py |
what to do with the when clause: store it, watch sensors, fire once per crossing |
examples/mqtt_bridge.py |
publish to home/<subject>/set for ESP32 / Tasmota / Zigbee2MQTT devices |
examples/home_assistant.py |
map subjects to Home Assistant entities and call services locally |
Run any of them from the repo root, all have a --dry-run or simulated mode:
python examples/quickstart.py
python examples/rule_engine.py
python examples/mqtt_bridge.py --dry-run "bedroom ka pankha high pe chala do"
python examples/home_assistant.py --dry-run "garden ki light band kar do"
The pipeline this slots into
mic -> wake word -> Whisper/Vosk (Hindi+English) -> THIS MODEL -> catalog check -> actuator
~100 ms ~0 ms GPIO / MQTT / ROS
Plain commands (no when) go straight to the actuator. Conditional commands are
stored as rules and evaluated on every sensor tick. rule_engine.py output on a
simulated day:
"agar temperature 35 se upar jaye to bedroom ka ac high kar do"
RULE stored: bedroom_ac -> HIGH when temperature > 35
"roz subah 6 baje garden sprinkler chalu kar do"
RULE stored: garden_sprinkler -> ON when time == 06:00
[sensor] temperature = 36
ACTUATE bedroom_ac HIGH (temperature=36 > 35)
[sensor] time = 06:00
ACTUATE garden_sprinkler ON (time=06:00 == 06:00)
Robotics use
For a mobile robot or a robotic arm the same pattern applies: swap the device
catalog for your robot's capabilities (arm_gripper: OPEN/CLOSE,
base_motor: START/STOP/LOW/MEDIUM/HIGH, head_camera: ON/OFF/STATUS),
regenerate the data with the scripts in the dataset repo, retrain for about an
hour on two CPU cores, and you have a command parser that fits in the L2 cache.
The when clause maps directly to sensor-gated behaviours
("agar battery 15 se neeche jaye to charging dock pe jao").
Output schema
{"activity": "...", "subject": "...", "action": "...",
"when": {"metric": "...", "op": ">|<|==", "value": 42}}
- 1,004 subjects in
model/devices-catalog.json, each with itsactivity(18 of them: light, power, air, access, security, air_quality, cooling, heating, entertainment, music, garden, kitchen, office, network, water, laundry, cleaning, aquarium) and its allowed actions. - 12 actions: ON, OFF, START, STOP, LOW, MEDIUM, HIGH, OPEN, CLOSE, LOCK, UNLOCK, STATUS.
- 7 metrics for
when: temperature, humidity, aqi, battery, water_level, soil_moisture, time (value"HH:MM", op==). whenis optional. Rules in the training data are sensible by construction: high temperature turns cooling ON or heating OFF, never the reverse.
Confidence guard
parse_scored returns the mean log-probability of the generated characters.
In-domain commands decode at about 0.0; gibberish or off-domain sentences
score clearly lower and the server flags them (default threshold -0.006, env
HCAI_CONF_THRESHOLD). Example: "what is the capital of france" scores -0.022
and is rejected. This is a useful guard, not a guarantee: command-shaped
off-domain input can still slip through.
Evaluation
1,500 held-out rows from the same generator (900 plain, 600 conditional), greedy decode, exact string match on the JSON:
| Metric | Value |
|---|---|
| exact match, all | 96.73 % |
| exact match, plain | 96.89 % |
| exact match, conditional | 96.50 % |
| valid JSON | 100 % |
Full numbers and 12 sample failures are in model/metrics_v3.json. The
held-out rows themselves were not preserved in this release; the dataset repo
ships the generators and the deterministic prep_mix.py (seed 108) that
describes the train/eval construction. The preserved 340,000-row training mix differs from the shipped script's configured counts; exact reconstruction of the missing historical evaluation split is unverified. The new Pi diagnostic uses a frozen sample from the preserved training mix and does not re-estimate the historical held-out score.
Honest limitations
- Synthetic, templated training data. The accuracy above measures interpolation inside the template space, not the wild. Novel or indirect phrasing ("andhera kar do", "thoda tez") fails.
- Observed misses that pass validation. "kitchen ki chimney high pe karo"
returned
STATUSin our tests, and that is an allowed action for the chimney, so the catalog check passed it. Treat STATUS on an imperative sentence with suspicion in your own guard. - Char-level spelling errors in subjects happen (see sample failures); the catalog check exists exactly for that. Reject anything with issues.
- No negation, no reject class. "fan mat chalao" can become
ON. - One instruction, one command. No multi-intent, no context or pronouns.
- Time rules are one-shot
==triggers; recurrence ("roz") is not encoded separately, your scheduler decides that. - Romanised Hindi only. Devanagari input is not in the vocabulary and is silently dropped by the tokenizer.
Adapting to your devices
- Edit the device tables in
generators/generate_dataset_v2.pyand the rule tables ingenerators/generate_dataset_v3_conditional.py(dataset repo). - Regenerate and mix:
python generate_dataset_v2.py && python generate_dataset_v3_conditional.py && python prep_mix.py - Retrain (needs torch; about 80 minutes on two CPU cores, minutes on a GPU):
python training/train_v3.py train --seconds 4200 && python training/train_v3.py export - Drop
model_v3_fp16.npz,model_v3_config.jsonand the newdevices-catalog.jsonintomodel/.
Files
| Path | What |
|---|---|
config.json |
release metadata; also the file Hugging Face uses to count downloads |
model/model_v3_fp16.npz |
weights, fp16, 2.1 MB |
model/model_v3_config.json |
vocab (itos) + architecture config |
model/devices-catalog.json |
the 1,004 subjects with allowed actions |
model/metrics_v3.json |
eval numbers and sample failures |
engine.py |
numpy inference, kv-cached greedy decode |
server.py, static/index.html |
Flask API + offline web UI |
examples/ |
quickstart, rule engine, MQTT bridge, Home Assistant bridge |
training/train_v3.py |
full PyTorch training and export source |
Dockerfile, docker-compose.yml |
150 MB image, healthcheck included |
License
Apache-2.0 for the weights, code and catalog. The training data is fully synthetic and released under the same license in the dataset repo. No warranty: this model suggests commands, your code decides whether to obey.
Copyright (c) 2026 sraivante applies only to original material and selection/arrangement contributed by the user; third-party ownership and licenses remain unchanged.
- Downloads last month
- 128
Dataset used to train sraivante/superfast-tiny-home-robotics-json-1m-v3
Evaluation results
- exact match (all) on home-commands-json-v3 (1,500 held-out rows, greedy decode)self-reported96.730
- exact match (plain commands, n=900) on home-commands-json-v3 (1,500 held-out rows, greedy decode)self-reported96.890
- exact match (conditional rules, n=600) on home-commands-json-v3 (1,500 held-out rows, greedy decode)self-reported96.500