# Smartwatch Integration Guide How to run **Smartwatch LM v0.1** on a wrist device and wire it to sensors, timers, and apps. The model is a **12M-parameter** GPT exported as ONNX. It does not execute device actions itself — it emits **intent tags** and **slot placeholders** that your firmware parses and handles. --- ## What you ship | File | Size (approx.) | Purpose | |------|----------------|---------| | `smartwatch_lm_merged.onnx` | ~52 MB | ONNX Runtime inference | | `tokenizer.json` | ~200 KB | Text ↔ token ids | | `tokenizer_config.json` | small | Tokenizer metadata | | `config.json` | small | Architecture and I/O names | | `reply_utils.py` | small | Cleanup, intent parse, slot fill | | `onnx_sample.py` | small | Reference ONNX generate loop | Optional on PC: `checkpoint.pt`, `chat.py`, `model.py` for PyTorch chat and fine-tuning. **ONNX I/O:** - Input: `input_ids` — int64, shape `[batch, seq]`, max seq **256** - Output: `logits` — float, shape `[batch, seq, vocab_size]` (vocab **3524**) Sample the **last position** logits autoregressively until EOS or max tokens. --- ## Architecture ``` User input (touch / voice-to-text) → build_prompt() → tokenizer.encode() → ONNX generate loop → tokenizer.decode (new tokens only) → reply_utils.clean + extract_intent_reply → intent router → sensor handlers → fill_slots() → display / TTS ``` | Component | Role | |-----------|------| | LM | Intent + reply template with `` tokens | | Router | Map `` to a handler | | Handlers | Read/write device state | | Slot map | Live values for ``, ``, etc. | See [Intent reference](./intent-reference.md) for all 35 intents. --- ## Step 1 — Inference runtime | Platform | Runtime | |----------|---------| | Wear OS | ONNX Runtime Mobile (NNAPI / XNNPACK) | | watchOS | ORT Mobile or Core ML (consider INT8 quant) | | Companion phone | ORT on phone, BLE to watch for display | | Prototype | `python onnx_sample.py` on desktop | Budget ~52 MB weights + activations. Quantization or phone-side inference helps on tight RAM. --- ## Step 2 — Generation loop Recommended settings (see [Avoiding gibberish](./avoiding-gibberish.md)): ```text temperature = 0.5 top_k = 40 max_new_tokens = 40 block_size = 256 eos_token_id = 0 ``` Reference implementations: - ONNX: [`onnx_sample.py`](../onnx_sample.py) - PyTorch: [`model.py`](../model.py) `GPT.generate()` + [`chat.py`](../chat.py) --- ## Step 3 — Prompt and history ``` user: How many steps today? bot: You're at of — keep going! user: Set a 10 minute timer bot: ``` Use `build_prompt()` from [`reply_utils.py`](../reply_utils.py). - History stores **raw** bot lines with unfilled slots - Never store display text with real numbers in history - Drop oldest turns when encoded length nears 256 tokens --- ## Step 4 — Cleanup and slots Always run the cleanup pipeline from [`reply_utils.py`](../reply_utils.py): 1. `extract_bot_reply(prompt, generated)` — one line, no fake `\nuser:` tail 2. `extract_intent_reply(raw)` → intent + template 3. `fill_slots(template, slot_map)` → user-visible string Example slot map: ```python { "STEPS_TODAY": "4,231", "STEP_GOAL": "10,000", "STEPS_REMAINING": "5,769", "BATTERY_PCT": "67%", "TIME": "2:15 PM", } ``` Character cleanup (`Ġ`, mojibake, etc.) is documented in [Avoiding gibberish](./avoiding-gibberish.md). --- ## Step 5 — Intent router ```python from reply_utils import fill_slots def dispatch(intent: str, template: str, slots: dict[str, str]) -> str: filled = fill_slots(template, slots) if intent == "GET_STEPS": refresh_step_cache(slots) elif intent == "START_TIMER": timer.start(slots.get("DURATION", "5 minutes")) elif intent == "NONE": pass else: log.warning("unsupported intent: %s", intent) return filled ``` Validate intents against your allowlist before side effects. --- ## Example trace **User:** “How many steps to hit my goal?” 1. Prompt ends with `bot:` 2. Model raw output: ` You need more to reach .` 3. Handler refreshes step slots from pedometer 4. Display: `You need 5,769 more to reach 10,000.` 5. History stores unfilled raw bot line --- ## Testing on desktop ```bash pip install torch tokenizers python chat.py ``` ```bash pip install numpy onnxruntime tokenizers python onnx_sample.py "How many steps today?" ``` ```bash python reply_utils.py # cleanup demo, no model ``` --- ## Production checklist - [ ] `smartwatch_lm_merged.onnx` + `tokenizer.json` bundled or downloaded once - [ ] Generation: temp 0.5, top_k 40, max 40 tokens - [ ] `reply_utils` cleanup on every decode path - [ ] History uses unfilled slot templates only - [ ] Intent allowlist matches implemented handlers - [ ] Slot map from real sensors before `fill_slots` - [ ] Fallback when parse fails or intent is unknown --- ## Related docs - [Avoiding gibberish](./avoiding-gibberish.md) — special characters, truncation, sample scripts - [Intent reference](./intent-reference.md) — all intents and slots