Superfast Tiny Home Robotics JSON 1M v3

A 1,143,328-parameter character-level transformer, trained from scratch, that turns an English or Hinglish sentence into one JSON device command, including conditional rules ("if temperature goes above 42...", "roz shaam 7 baje..."). Weights are 2.1 MB (fp16). Inference is numpy only: no PyTorch, no tokenizer library, no internet, no GPU.

"bhai bedroom ka pankha high pe chala do"
  -> {"activity":"air","subject":"bedroom_ceiling_fan","action":"HIGH"}

"agar room ka temperature 42 se upar jaye to bedroom ka heater band kar do"
  -> {"activity":"heating","subject":"bedroom_room_heater","action":"OFF",
      "when":{"metric":"temperature","op":">","value":42}}

"roz shaam 7 baje garden ki batti chala do"
  -> {"activity":"light","subject":"garden_light","action":"ON",
      "when":{"metric":"time","op":"==","value":"19:00"}}

"when soil moisture drops below 20 start the garden sprinkler"
  -> {"activity":"garden","subject":"garden_sprinkler","action":"ON",
      "when":{"metric":"soil_moisture","op":"<","value":20}}

This is the v3 of the series. Compared with v1 (30 commands, 10 device types, transformers runtime) it covers 1,004 devices, 12 actions, 7 sensor metrics for rules, and ships with a numpy engine plus a web UI and Docker image.

Release Value
Publisher sraivante
Release date 2026-09-26
Version v3.0.0
Training data home-commands-json-v3 (synthetic, templated)
Architecture GPT-style causal decoder, char-level vocab (78 symbols), d=160, 4 layers, 4 heads, FFN 512, context 360
Parameters 1,143,328
Weights model/model_v3_fp16.npz, 2,146,783 bytes
Runtime deps numpy (plus flask for the optional server)
License Apache-2.0

Raspberry Pi 5 edge-device performance

Measured 2026-09-26 on a Raspberry Pi 5 Model B, 16 GB RAM, Debian 13.2/aarch64, Python 3.13.5. Measurement update: edge-rpi5-20260926.1; weights and training data are unchanged.

Threads Median ms p95 ms p99 ms Serial requests/s Process RSS MiB CPU °C min–mean–max Timed correct
1 410.38 553.20 558.65 2.53 43.77 59.0–60.9–63.4 500/500
4 393.10 550.26 553.48 2.55 44.67 60.0–62.0–64.5 500/500

Batch 1; 20 warmups and 500 timed calls per row, cycling over 10 frozen model-specific inputs. Each timed reading has a matching CPU-temperature sample, captured immediately afterward outside the latency timer; the table shows min–mean–max temperature. RSS includes Python/runtime overhead. Timings include local text-to-output inference and exclude loading, SSH/HTTP, speech recognition and actuation. Throughput is serial reciprocal mean. Runs used one and four CPU/BLAS threads sequentially. These synthetic workloads do not establish cross-model superiority or robot-control reliability.

On-device quality check: 192/200 on training-set diagnostic, NOT held-out accuracy. For v3 this is explicitly a training-set diagnostic, not held-out accuracy; the earlier 96.73% result was not rerun.

Full Pi methodology, temperature plot and reproduction · Every latency/temperature reading · Raw measurements · Versioned evaluation data.

Why a tiny model and not an LLM

A home robot or a Raspberry Pi behind a microphone does not need to know the capital of France. It needs to turn "pankha tez kar do" into a command in a few milliseconds, offline, on a CPU that is also doing wake-word detection. This model is that one component: the parser between speech-to-text and the actuator. Everything else (safety, permissions, device state) stays in deterministic code you control.

Measured on this release, single request, Python 3.13, numpy 2.3, Intel Core i7-1360P, default threads:

Input Latency
plain command (about 30 chars) 70 to 90 ms
conditional rule (about 70 chars) 130 to 150 ms
process RSS, engine loaded + warm (Python + numpy included) 73 MB

Latency is dominated by Python overhead per generated character, not by FLOPs. Raspberry Pi 5 measurements are now recorded above; Jetson Orin Nano remains unmeasured. The table here retains the earlier desktop measurements.

Quick start

pip install numpy huggingface_hub
python -c "from huggingface_hub import snapshot_download; snapshot_download('sraivante/superfast-tiny-home-robotics-json-1m-v3', local_dir='superfast-tiny-home-robotics-json-1m-v3')"
cd superfast-tiny-home-robotics-json-1m-v3
python engine.py "kitchen ki light band kar do"
# {"activity":"light","subject":"kitchen_light","action":"OFF"}

git clone https://huggingface.co/sraivante/superfast-tiny-home-robotics-json-1m-v3 also works.

In Python:

import engine
engine.init("model")                        # loads the npz once
raw, confidence = engine.parse_scored("turn off the porch light")
# raw        -> '{"activity":"light","subject":"porch_light","action":"OFF"}'
# confidence -> avg log-prob per generated char; near 0.0 = in-domain

Web UI + REST API (Docker, no GPU):

docker compose up --build         # open http://localhost:8080
curl -s localhost:8080/api/parse -H 'Content-Type: application/json' \
     -d '{"text":"if aqi goes above 250 set the bedroom air purifier on high"}'

The API returns the command, valid_json, catalog issues, confidence, latency_ms and a single ok flag. Act only on ok: true.

How it is meant to be used (examples/)

The model is a parser. model/devices-catalog.json is the law. Every example below runs the output through catalog validation before anything moves.

Example What it shows
examples/quickstart.py text in, validated command out, off-domain input rejected
examples/rule_engine.py what to do with the when clause: store it, watch sensors, fire once per crossing
examples/mqtt_bridge.py publish to home/<subject>/set for ESP32 / Tasmota / Zigbee2MQTT devices
examples/home_assistant.py map subjects to Home Assistant entities and call services locally

Run any of them from the repo root, all have a --dry-run or simulated mode:

python examples/quickstart.py
python examples/rule_engine.py
python examples/mqtt_bridge.py --dry-run "bedroom ka pankha high pe chala do"
python examples/home_assistant.py --dry-run "garden ki light band kar do"

The pipeline this slots into

mic -> wake word -> Whisper/Vosk (Hindi+English) -> THIS MODEL -> catalog check -> actuator
                                                      ~100 ms        ~0 ms       GPIO / MQTT / ROS

Plain commands (no when) go straight to the actuator. Conditional commands are stored as rules and evaluated on every sensor tick. rule_engine.py output on a simulated day:

"agar temperature 35 se upar jaye to bedroom ka ac high kar do"
  RULE stored: bedroom_ac -> HIGH when temperature > 35
"roz subah 6 baje garden sprinkler chalu kar do"
  RULE stored: garden_sprinkler -> ON when time == 06:00

[sensor] temperature = 36
  ACTUATE bedroom_ac                   HIGH    (temperature=36 > 35)
[sensor] time = 06:00
  ACTUATE garden_sprinkler             ON      (time=06:00 == 06:00)

Robotics use

For a mobile robot or a robotic arm the same pattern applies: swap the device catalog for your robot's capabilities (arm_gripper: OPEN/CLOSE, base_motor: START/STOP/LOW/MEDIUM/HIGH, head_camera: ON/OFF/STATUS), regenerate the data with the scripts in the dataset repo, retrain for about an hour on two CPU cores, and you have a command parser that fits in the L2 cache. The when clause maps directly to sensor-gated behaviours ("agar battery 15 se neeche jaye to charging dock pe jao").

Output schema

{"activity": "...", "subject": "...", "action": "...",
 "when": {"metric": "...", "op": ">|<|==", "value": 42}}
  • 1,004 subjects in model/devices-catalog.json, each with its activity (18 of them: light, power, air, access, security, air_quality, cooling, heating, entertainment, music, garden, kitchen, office, network, water, laundry, cleaning, aquarium) and its allowed actions.
  • 12 actions: ON, OFF, START, STOP, LOW, MEDIUM, HIGH, OPEN, CLOSE, LOCK, UNLOCK, STATUS.
  • 7 metrics for when: temperature, humidity, aqi, battery, water_level, soil_moisture, time (value "HH:MM", op ==).
  • when is optional. Rules in the training data are sensible by construction: high temperature turns cooling ON or heating OFF, never the reverse.

Confidence guard

parse_scored returns the mean log-probability of the generated characters. In-domain commands decode at about 0.0; gibberish or off-domain sentences score clearly lower and the server flags them (default threshold -0.006, env HCAI_CONF_THRESHOLD). Example: "what is the capital of france" scores -0.022 and is rejected. This is a useful guard, not a guarantee: command-shaped off-domain input can still slip through.

Evaluation

1,500 held-out rows from the same generator (900 plain, 600 conditional), greedy decode, exact string match on the JSON:

Metric Value
exact match, all 96.73 %
exact match, plain 96.89 %
exact match, conditional 96.50 %
valid JSON 100 %

Full numbers and 12 sample failures are in model/metrics_v3.json. The held-out rows themselves were not preserved in this release; the dataset repo ships the generators and the deterministic prep_mix.py (seed 108) that describes the train/eval construction. The preserved 340,000-row training mix differs from the shipped script's configured counts; exact reconstruction of the missing historical evaluation split is unverified. The new Pi diagnostic uses a frozen sample from the preserved training mix and does not re-estimate the historical held-out score.

Honest limitations

  • Synthetic, templated training data. The accuracy above measures interpolation inside the template space, not the wild. Novel or indirect phrasing ("andhera kar do", "thoda tez") fails.
  • Observed misses that pass validation. "kitchen ki chimney high pe karo" returned STATUS in our tests, and that is an allowed action for the chimney, so the catalog check passed it. Treat STATUS on an imperative sentence with suspicion in your own guard.
  • Char-level spelling errors in subjects happen (see sample failures); the catalog check exists exactly for that. Reject anything with issues.
  • No negation, no reject class. "fan mat chalao" can become ON.
  • One instruction, one command. No multi-intent, no context or pronouns.
  • Time rules are one-shot == triggers; recurrence ("roz") is not encoded separately, your scheduler decides that.
  • Romanised Hindi only. Devanagari input is not in the vocabulary and is silently dropped by the tokenizer.

Adapting to your devices

  1. Edit the device tables in generators/generate_dataset_v2.py and the rule tables in generators/generate_dataset_v3_conditional.py (dataset repo).
  2. Regenerate and mix: python generate_dataset_v2.py && python generate_dataset_v3_conditional.py && python prep_mix.py
  3. Retrain (needs torch; about 80 minutes on two CPU cores, minutes on a GPU): python training/train_v3.py train --seconds 4200 && python training/train_v3.py export
  4. Drop model_v3_fp16.npz, model_v3_config.json and the new devices-catalog.json into model/.

Files

Path What
config.json release metadata; also the file Hugging Face uses to count downloads
model/model_v3_fp16.npz weights, fp16, 2.1 MB
model/model_v3_config.json vocab (itos) + architecture config
model/devices-catalog.json the 1,004 subjects with allowed actions
model/metrics_v3.json eval numbers and sample failures
engine.py numpy inference, kv-cached greedy decode
server.py, static/index.html Flask API + offline web UI
examples/ quickstart, rule engine, MQTT bridge, Home Assistant bridge
training/train_v3.py full PyTorch training and export source
Dockerfile, docker-compose.yml 150 MB image, healthcheck included

License

Apache-2.0 for the weights, code and catalog. The training data is fully synthetic and released under the same license in the dataset repo. No warranty: this model suggests commands, your code decides whether to obey.

Copyright (c) 2026 sraivante applies only to original material and selection/arrangement contributed by the user; third-party ownership and licenses remain unchanged.

Downloads last month
128
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train sraivante/superfast-tiny-home-robotics-json-1m-v3

Evaluation results

  • exact match (all) on home-commands-json-v3 (1,500 held-out rows, greedy decode)
    self-reported
    96.730
  • exact match (plain commands, n=900) on home-commands-json-v3 (1,500 held-out rows, greedy decode)
    self-reported
    96.890
  • exact match (conditional rules, n=600) on home-commands-json-v3 (1,500 held-out rows, greedy decode)
    self-reported
    96.500