---
license: apache-2.0
language:
- en
- hi
pipeline_tag: text-generation
tags:
- function-calling
- tool-use
- agent
- intent-classification
- slot-filling
- seq2seq
- hinglish
- code-mixed
- voice-assistant
- raspberry-pi
- edge
- on-device
- cpu
- numpy
---
# tiny-superfast-agentic-MLM-6.5m-v1: a 6.4M-parameter voice-command parser for laptops and Raspberry Pi
**tiny-superfast-agentic-MLM-6.5m-v1** turns one short spoken or typed command into a structured command that a
program can execute. It understands English and Hinglish (Roman-script Hindi-English), with some Devanagari Hindi, and covers **371 actions**
(audio, display, Wi-Fi, apps, files, timers, power, network, services, Raspberry Pi GPIO/I2C and more), and runs on
a plain CPU with **NumPy only**: no GPU, no PyTorch and no internet needed at inference.
```text
ENGLISH "set the volume to 40" -> set_volume {"value": 40}
HINGLISH "volume 40 kar do" -> set_volume {"value": 40}
ENGLISH "connect to wifi Redmi Note 12 password hello@123" -> connect_wifi {"ssid": "Redmi Note 12", "password": "hello@123"}
HINGLISH "pwd Tiger@2025 daalo aur Hostel Wing C se jud jao" -> connect_wifi {"ssid": "Hostel Wing C", "password": "Tiger@2025"}
HINGLISH "10 min baad shutdown kar dena" -> shutdown {"amount": 10, "unit": "min"}
HINGLISH "board pin 11 pe led jalao" -> gpio_on {"pin": 11, "numbering": "board"}
HINDI "5 मिनट का टाइमर लगाओ" -> set_timer {"amount": 5, "unit": "min"}
```
| | |
|---|---|
| Parameters | **6,441,472** (≈6.4M; "6.5m" in the name is rounded up) |
| Model file | `model/model.npz`, 23.9 MB, float32. `pytorch/model.pt` is included for fine-tuning |
| Runs on | CPU only, fully offline: Windows, Linux, Raspberry Pi 5. Inference needs only `numpy` |
| Accuracy | **97.53% exact** on 6,611 held-out-template commands (`test_strict`); **99.10%** on validation; 100% on 461 unseen Wi-Fi phrasings |
| Speed | median **18 ms** per command on a laptop i7-1360P, **30 ms** on a Raspberry Pi 5 (10 min sustained load) |
| Load time | 0.13 s (laptop), 0.20 s (Raspberry Pi 5) |
| Bigger sibling | [tiny-agentic-home-robotic-for-edge-device-v10](https://huggingface.co/sraivante/tiny-agentic-home-robotic-for-edge-device-v10) (23.9M params, MiniLM-based) |
> **About the name.** "MLM" here means *micro language model*. It is **not** a masked language model like BERT:
> the architecture is a small Transformer encoder-decoder (seq2seq) that writes the command one token at a time.
> **Executor notice.** The model only *parses* commands; it never runs anything. The executor in this repository is a
> **sample for testing, not a finished product**. It gates every action by risk (safe / caution / critical) and asks
> for confirmation before anything risky. Use it on a test machine and read the command before you run it.
## Best use: a personal voice companion for small commands
The model is built to be the "understanding" step of a small, private, always-available assistant: you **speak** a
short command, and your computer or Raspberry Pi **does it**. It replaces a large LLM for common, well-defined device
commands, so the answer is instant (tens of milliseconds), private (nothing leaves the device) and free.
Attach any speech recognizer (ASR) in front of it and an executor behind it:
```mermaid
flowchart LR
U(["User speaks
'volume 40 kar do'"]) --> A["ASR
speech → text
(Whisper, Vosk, ...)"]
A -->|"text"| M["tiny-superfast-agentic-MLM-6.5m-v1
text → structured command
6.4M params · ~20–70 ms on CPU"]
M -->|"set_volume {value: 40}
+ confidence"| E["Executor
validate → risk gate → OS command"]
E --> O(["Output action
volume set to 40%"])
E -. "clarify / unknown /
low confidence" .-> U
```
If your viewer does not render Mermaid:
```text
user voice ==> [ ASR ] --text--> [ tiny-superfast-agentic-MLM-6.5m-v1 ] --command JSON--> [ executor ] ==> output action
|
ask the user again <---- clarify / unknown / low confidence
```
| Stage | What it does | Example |
|---|---|---|
| **1. ASR** (you bring it) | Turns speech into text. Any engine works: Whisper / faster-whisper, Vosk, a phone keyboard, or a chat box. | audio → `volume 40 kar do` |
| **2. This model** | Turns text into one action from the 371-action catalog, plus its typed arguments and a confidence score. | → `{"action": "set_volume", "value": 40}`, confidence 0.997 |
| **3. Executor** (sample included) | Checks the arguments, applies the risk gate, renders a whitelisted OS command for Windows or Raspberry Pi, and runs it or asks first. | → Pi: `wpctl set-volume @DEFAULT_AUDIO_SINK@ 40%` |
| **4. Output action** | The device acts, and the result can be read back to the user (TTS). | volume is now 40% |
A good agent pattern: execute only when confidence is high and the risk gate allows it, ask the user when the result
is `clarify`, and hand `unknown` or low-confidence text to a bigger model or back to the user.
### Minimal voice loop (example)
```python
# pip install numpy faster-whisper (ASR is your choice; this is one option)
import os; os.environ.setdefault("OPENBLAS_NUM_THREADS", "4") # "2" on a Raspberry Pi 5, see Speed
import sys; sys.path[:0] = ["model", "executor"]
from runtime import CommandModel
import executor as EX
from faster_whisper import WhisperModel
asr = WhisperModel("small", device="cpu", compute_type="int8")
parser = CommandModel("model")
def handle(wav_path):
text = " ".join(s.text for s in asr.transcribe(wav_path)[0]).strip() # 1. speech -> text
out = parser.predict(text); meta = out.pop("_meta") # 2. text -> command
plan = EX.plan(out, None, text) # 3. command -> gated OS command
if plan.action == "clarify" or meta["confidence"] < 0.8:
return f"Sorry, can you say that again? ({plan.note or text})"
ok, why = EX.gate(plan, allow_caution=False) # safe actions only in this demo
return EX.run(plan, out, dry_run=not ok) # 4. act (dry run if gated)
```
The ASR part of this snippet is an illustration. The model was trained on typed/synthetic text, not on real ASR
transcripts, so test it with your own ASR output (see Limitations).
## What it understands
Every example is shown in ENGLISH, then the same command in HINGLISH. The outputs, confidences and executor plans are
real outputs of this model and the included sample executor.
| Use case | Language | Command | Model output | Confidence | Risk | Raspberry Pi command |
|---|---|---|---|---|---|---|
| Audio | ENGLISH | `set the volume to 40` | `set_volume {"value": 40}` | 0.996 | safe | `wpctl set-volume … 40%` |
| Audio | HINGLISH | `volume 40 kar do` | `set_volume {"value": 40}` | 0.997 | safe | `wpctl set-volume … 40%` |
| Wi-Fi | ENGLISH | `connect to wifi Redmi Note 12 password hello@123` | `connect_wifi {"ssid": "Redmi Note 12", "password": "hello@123"}` | 1.000 | caution | `nmcli dev wifi connect 'Redmi Note 12' password hello@123` |
| Wi-Fi | HINGLISH | `Redmi Note 12 wifi se connect karo password hello@123` | same | 1.000 | caution | same |
| Wi-Fi, password first | ENGLISH | `password Tiger@2025 then join Hostel Wing C` | `connect_wifi {"ssid": "Hostel Wing C", "password": "Tiger@2025"}` | 1.000 | caution | `nmcli dev wifi connect 'Hostel Wing C' …` |
| Wi-Fi, password first | HINGLISH | `pwd Tiger@2025 daalo aur Hostel Wing C se jud jao` | same | 1.000 | caution | same |
| Apps | ENGLISH | `open chrome` | `open_app {"app": "chrome"}` | 0.990 | caution | `chromium-browser &` |
| Apps | HINGLISH | `chrome kholo` | `open_app {"app": "chrome"}` | 0.986 | caution | `chromium-browser &` |
| Timers | ENGLISH | `set a timer for 5 minutes` | `set_timer {"amount": 5, "unit": "min"}` | 0.998 | safe | built-in timer |
| Timers | HINGLISH | `5 minute ka timer lagao` | `set_timer {"amount": 5, "unit": "min"}` | 0.998 | safe | built-in timer |
| Power | ENGLISH | `shut down in 10 minutes` | `shutdown {"amount": 10, "unit": "min"}` | 0.998 | critical | `sudo shutdown -h +10` |
| Power | HINGLISH | `10 min baad shutdown kar dena` | `shutdown {"amount": 10, "unit": "min"}` | 0.998 | critical | `sudo shutdown -h +10` |
| Negation | ENGLISH | `don't shut down the laptop` | `cancel_shutdown {}` | 0.995 | caution | `shutdown -c` |
| Negation | HINGLISH | `shutdown mat karo` | `cancel_shutdown {}` | 0.995 | caution | `shutdown -c` |
| Pi GPIO | ENGLISH | `turn on board pin 11` | `gpio_on {"pin": 11, "numbering": "board"}` | 0.999 | critical | `sudo pinctrl set 17 op dh` |
| Pi GPIO | HINGLISH | `board pin 11 pe led jalao` | `gpio_on {"pin": 11, "numbering": "board"}` | 0.998 | critical | `sudo pinctrl set 17 op dh` |
| Pi GPIO, ambiguous pin | HINGLISH | `pin no 11 high karo` | `gpio_on {"pin": 11}` | 0.993 | → clarify | executor asks: *board pin 11 (= GPIO 17) or GPIO 11 (= board pin 23)?* |
| Diagnostics | ENGLISH | `what is the cpu temperature` | `get_temperature {}` | 0.995 | safe | `vcgencmd measure_temp` |
| Diagnostics | HINGLISH | `cpu ka temperature batao` | `get_temperature {}` | 0.996 | safe | `vcgencmd measure_temp` |
| Files | HINGLISH | `reports naam ka folder banao` | `create_folder {"path": "reports"}` | 0.999 | caution | `mkdir -p -- reports` |
| Out of scope | ENGLISH | `tell me a joke` | `unknown {}` | 0.993 | safe | nothing |
| Out of scope (miss) | HINGLISH | `kya haal hai bhai` | `recent_files {}` | **0.604** | safe | (low confidence: ask or hand off) |
The last row is a real miss: small talk went to a harmless action, but with low confidence. This is why the agent
should check the confidence before acting.
## Test results
### Accuracy
Held-out evaluation after training (raw model output, before the runtime "snap" repair). "Exact" means the action
and every argument match; "Action" means only the action matches.
| Set | n | Exact | Action | What it tests |
|---|---|---|---|---|
| validation | 10,413 | 99.10% | 99.29% | same templates as training, unseen rows |
| test | 10,618 | 98.37% | 98.52% | held-out templates |
| **test_strict** | **6,611** | **97.53%** | **97.70%** | held-out templates with no paraphrase sibling in training (the honest generalisation number) |
| curated | 371 | 99.46% | 99.73% | one held-out prompt per action |
| B_typo | 3,000 | 99.17% | 99.50% | typos |
| wifi_holdout | 461 | 100.00% | 100.00% | unseen Wi-Fi connect phrasings |
| gap11_holdout | 696 | 99.28% | 100.00% | unseen templates: password-first Wi-Fi, CPU clock vs usage, board↔BCM pins, fan vs mouse speed |
| A_challenge | 1,152 | 97.40% | 98.44% | external challenge set |
| B_golden_final | 384 | 97.40% | 97.40% | external golden set, including out-of-scope prompts |
| B_golden_dev | 383 | 95.04% | 95.82% | external golden set |
| A_practical-dev | 1,321 | 95.31% | 96.44% | external practical set |
| A_practical-holdout | 428 | 92.06% | 94.16% | external practical set; some misses are label-convention differences |
The same model was also run on the laptop and on the Raspberry Pi 5 (2,340 prompts, with the runtime snap). **Both
devices produce identical predictions: 97.82% exact, 98.46% action.**
### Speed and temperature
10 minutes of continuous single-stream inference on each device. Full per-pass and per-minute tables are in
[`eval/BENCHMARK.md`](eval/BENCHMARK.md).
| | Laptop (Intel i7-1360P) | Raspberry Pi 5 (16 GB) |
|---|---|---|
| Latency p50 / p95 / p99 | 18.4 / 168.6 / 239.6 ms | 30.2 / 243.9 / 281.6 ms |
| Throughput, cool → after 10 min | 29.6 → 18.5 commands/s | 13.9 → 13.9 commands/s |
| Temperature idle → peak | 52.6 → 91.9 °C | 38.5 → 59.8 °C |
| Throttling | slows down ~39% once hot | none (2400 MHz throughout, `get_throttled=0x0`) |
| Best BLAS threads | 4 | 2 |
- Latency grows with the length of the command the model writes: short commands take ~10–35 ms, Wi-Fi commands with
an SSID and a password take ~100–250 ms.
- **Set the BLAS thread count.** The matrices are small, so NumPy's default (one thread per core) is up to 4x slower.
Use `OPENBLAS_NUM_THREADS=4` on x86 laptops and `OPENBLAS_NUM_THREADS=2` on the Raspberry Pi 5 (`quickstart.py`
does this for you).
### Test devices
| | Laptop | Raspberry Pi 5 |
|---|---|---|
| CPU | Intel Core i7-1360P (4P + 8E cores) | Broadcom BCM2712, 4× Cortex-A76 @ 2.4 GHz |
| RAM | 32 GB | 16 GB |
| OS | Windows 11 | Raspberry Pi OS (Debian 13), kernel 6.12 |
| Python / NumPy | 3.13 / 2.5.3 | 3.13 / 2.2.4 |
| Cooling | built-in fan, AC power | active cooler (PWM fan) |
## Download and test
### 1. Get the files
```bash
pip install -U huggingface_hub
hf download sraivante/tiny-superfast-agentic-MLM-6.5m-v1 --local-dir tiny-agentic-6.5m
cd tiny-agentic-6.5m
```
### 2. Parse commands (nothing is executed)
```bash
pip install numpy
python quickstart.py "volume 40 kar do" "connect to wifi Redmi Note 12 password hello@123"
python quickstart.py # interactive
```
Each line prints the command, the confidence, the latency, the risk level and the OS command the sample executor
*would* run on this machine.
### 3. Test UI and checks (Python 3.9+)
| | Windows | Raspberry Pi / Linux |
|---|---|---|
| Test UI in the browser (http://localhost:8000) | `run.bat` | `./run.sh` |
| Terminal agent | `run.bat cli` | `./run.sh cli` |
| Smoke test (model + executor, nothing executed) | `run.bat test` / `run.bat test --full` | `./run.sh test` / `./run.sh test --full` |
| HTML report of all 371 actions | `run.bat report` | `./run.sh report` |
The first run creates `.venv` and installs `numpy`, `fastapi` and `uvicorn`. `report` really runs the *safe,
read-only* actions on your machine (for example reading the volume or the IP address) and dry-runs everything else.
In the UI, *Run all* is always a dry run. **Execute** on a row runs that one command: safe actions run directly,
caution actions need the "allow caution" box, and critical actions ask you to type the action name. The UI listens on
127.0.0.1 only and has no login, so do not expose it to a network.
### 4. Reproduce the speed test
```bash
cd eval
sh bench_run_all.sh python ../model