prathamkode commited on
Commit
7629151
Β·
verified Β·
1 Parent(s): af872ae

Add integration and output quality documentation

Browse files
README.md CHANGED
@@ -89,12 +89,19 @@ The model predicts the next token. Slot placeholders like `<STEPS_TODAY>` are fi
89
 
90
  On-device smartwatch / wearable assistant demos. Not intended for general-purpose chat or safety-critical applications.
91
 
 
 
 
 
 
 
 
 
 
 
92
  ## Limitations
93
 
94
  - Trained on synthetic data; behavior on real user phrasing may vary
95
  - Does not output real metric values β€” only intents and slot tokens
96
  - Small model with limited reasoning capability
97
 
98
- ## Training
99
-
100
- Trained with the `electron-v1` project (`collab-run-1`). See the parent repo for training instructions and intent reference.
 
89
 
90
  On-device smartwatch / wearable assistant demos. Not intended for general-purpose chat or safety-critical applications.
91
 
92
+ ## Documentation
93
+
94
+ | Guide | Description |
95
+ |-------|-------------|
96
+ | [Intent reference](docs/intent-reference.md) | All 35 intents and slot placeholders |
97
+ | [Web output quality](docs/web-output-quality.md) | Sampling, cleanup, and anti-gibberish techniques for browser/on-device runtimes |
98
+ | [Smartwatch integration](docs/smartwatch-integration.md) | End-to-end guide to wire the model into wrist hardware |
99
+
100
+ Live browser demo: see the `collab-run-1/web` app in the electron-v1 training project (links to this model on Hugging Face).
101
+
102
  ## Limitations
103
 
104
  - Trained on synthetic data; behavior on real user phrasing may vary
105
  - Does not output real metric values β€” only intents and slot tokens
106
  - Small model with limited reasoning capability
107
 
 
 
 
docs/intent-reference.md ADDED
@@ -0,0 +1,163 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Wrist Wearable Assistant β€” Intent Reference
2
+
3
+ This document describes every intent in the [`tinydata/schema.py`](../tinydata/schema.py) vocabulary. Intents are machine-readable action tags the on-device model emits; your app parses them and runs the matching handler.
4
+
5
+ ## How intents work
6
+
7
+ Bot lines in training data and at runtime use this format:
8
+
9
+ ```
10
+ bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> β€” keep going!
11
+ ```
12
+
13
+ | Part | Role |
14
+ |------|------|
15
+ | `<INTENT:NAME>` | Tells the app **what to do** (fetch steps, start timer, etc.) |
16
+ | `<SLOT>` placeholders | Names **which values** the app injects from sensors/settings |
17
+ | Text after the tag | What the user **sees or hears** on the watch |
18
+
19
+ The model never outputs real metric values β€” only intents and slot tokens. Your runtime replaces slots with live data.
20
+
21
+ `NONE` is used when the bot only acknowledges or chats with no device action.
22
+
23
+ ---
24
+
25
+ ## Steps
26
+
27
+ | Intent | What it does | Typical slots |
28
+ |--------|--------------|---------------|
29
+ | `GET_STEPS` | Read step metrics from the pedometer: today’s count, goals, streaks, week/month totals, comparisons. | `<STEPS_TODAY>`, `<STEPS_YESTERDAY>`, `<STEP_GOAL>`, `<STEPS_REMAINING>`, `<STEP_GOAL_PCT>`, `<STEP_STREAK_DAYS>`, `<STEPS_WEEK_TOTAL>`, `<STEPS_LAST_WEEK_TOTAL>`, `<STEPS_MONTH_TOTAL>`, `<TIME>`, `<DATE>`, `<PERCENT>` |
30
+ | `SET_STEP_GOAL` | Change the user’s daily step target. | `<STEP_GOAL>` |
31
+
32
+ ---
33
+
34
+ ## Distance
35
+
36
+ | Intent | What it does | Typical slots |
37
+ |--------|--------------|---------------|
38
+ | `GET_DISTANCE` | Read distance metrics: today’s total, last session, units, weekly progress, lifetime mileage, hourly average. | `<DISTANCE_TODAY>`, `<DISTANCE_LAST_SESSION>`, `<DISTANCE_UNIT>`, `<DISTANCE_WEEK_REMAINING>`, `<DISTANCE_WEEK_BEST>`, `<DISTANCE_LIFETIME>`, `<DISTANCE_HOURLY_AVG>`, `<TIME>`, `<DATE>` |
39
+
40
+ ---
41
+
42
+ ## Heart rate
43
+
44
+ | Intent | What it does | Typical slots |
45
+ |--------|--------------|---------------|
46
+ | `GET_HEART_RATE` | Read heart rate data: current BPM, resting rate, overnight low, session average, daily peak, elevated minutes, effort zone. | `<HR_CURRENT_BPM>`, `<HR_RESTING_BPM>`, `<HR_RESTING_OVERNIGHT_LOW>`, `<HR_AVG_SESSION>`, `<HR_PEAK_TODAY>`, `<HR_ELEVATED_MINUTES>`, `<HR_ZONE>`, `<DURATION>` |
47
+ | `MEASURE_HEART_RATE` | Trigger a live on-demand pulse reading from the optical sensor. | `<HR_CURRENT_BPM>` |
48
+
49
+ ---
50
+
51
+ ## Sleep
52
+
53
+ | Intent | What it does | Typical slots |
54
+ |--------|--------------|---------------|
55
+ | `GET_SLEEP` | Read sleep data: duration last night, sleep/wake times, weekly average, awake minutes, goal gap, weekend total. | `<SLEEP_HOURS_LAST_NIGHT>`, `<SLEEP_START_TIME>`, `<SLEEP_WAKE_TIME>`, `<SLEEP_AVG_WEEK>`, `<SLEEP_AWAKE_MINUTES>`, `<SLEEP_GOAL_HOURS>`, `<SLEEP_WEEKEND_TOTAL>`, `<DATE>`, `<DURATION>` |
56
+ | `LOG_NAP` | Record or confirm a short nap session. | `<DURATION>`, `<SLEEP_HOURS_LAST_NIGHT>`, `<TIME>` |
57
+
58
+ ---
59
+
60
+ ## Calories
61
+
62
+ | Intent | What it does | Typical slots |
63
+ |--------|--------------|---------------|
64
+ | `GET_CALORIES` | Read calorie burn: daily total, active burn, remaining to goal, yesterday comparison, per-mile estimate. | `<CALORIES_TODAY>`, `<CALORIES_ACTIVE>`, `<CALORIES_REMAINING>`, `<CALORIE_GOAL>`, `<CALORIE_GOAL_PCT>`, `<CALORIES_YESTERDAY>`, `<CALORIES_PER_MILE>`, `<DISTANCE_UNIT>`, `<DURATION>` |
65
+ | `SET_CALORIE_GOAL` | Change the user’s daily active calorie burn target. | `<CALORIE_GOAL>` |
66
+
67
+ ---
68
+
69
+ ## Active minutes
70
+
71
+ | Intent | What it does | Typical slots |
72
+ |--------|--------------|---------------|
73
+ | `GET_ACTIVE_MINUTES` | Read movement time: today’s active minutes, remaining to goal, weekly total, most active day. | `<ACTIVE_MINUTES_TODAY>`, `<ACTIVE_MINUTES_REMAINING>`, `<ACTIVE_MINUTES_GOAL>`, `<ACTIVE_MINUTES_WEEK>`, `<MOST_ACTIVE_DAY>`, `<DURATION>`, `<TIME>`, `<DATE>` |
74
+ | `SET_ACTIVE_GOAL` | Change the user’s daily active minutes target. | `<ACTIVE_MINUTES_GOAL>` |
75
+
76
+ ---
77
+
78
+ ## Battery
79
+
80
+ | Intent | What it does | Typical slots |
81
+ |--------|--------------|---------------|
82
+ | `GET_BATTERY` | Read power state: percentage, estimated days left, days since charge, charge complete status. | `<BATTERY_PCT>`, `<BATTERY_DAYS_LEFT>`, `<DAYS_SINCE_CHARGE>`, `<CHARGE_COMPLETE>`, `<DURATION>`, `<DATE>` |
83
+ | `ENABLE_POWER_SAVE` | Turn on low-power / battery saver mode on the device. | `<BATTERY_PCT>` |
84
+ | `DISABLE_AOD` | Turn off always-on display to reduce power draw. | `<BATTERY_PCT>` |
85
+
86
+ ---
87
+
88
+ ## Alarms and reminders
89
+
90
+ | Intent | What it does | Typical slots |
91
+ |--------|--------------|---------------|
92
+ | `SET_ALARM` | Create or update a wake alarm at a given time. | `<ALARM_TIME>`, `<ALARM_LABEL>`, `<DATE>`, `<TIME>` |
93
+ | `LIST_ALARMS` | Show all alarms currently stored on the device. | `<ALARM_TIME>`, `<ALARM_LABEL>` |
94
+ | `DELETE_ALARM` | Remove one or all alarms. | `<ALARM_LABEL>`, `<ALARM_TIME>` |
95
+ | `SNOOZE_ALARM` | Snooze the currently firing alarm for a short period. | `<ALARM_TIME>`, `<DURATION>`, `<ALARM_LABEL>` |
96
+ | `SET_REMINDER` | Schedule a recurring reminder (water, bedtime wind-down, move prompts). | `<ALARM_TIME>`, `<ALARM_LABEL>`, `<REMINDER_INTERVAL>`, `<TIME>` |
97
+ | `MUTE_REMINDERS` | Disable hourly move reminders or similar nudges. | *(none)* |
98
+ | `EXPLAIN_NUDGE` | Explain why the watch just vibrated (e.g. sedentary move nudge). | `<NUDGE_REASON>`, `<DURATION>`, `<TIME>` |
99
+
100
+ ---
101
+
102
+ ## Timers and stopwatch
103
+
104
+ | Intent | What it does | Typical slots |
105
+ |--------|--------------|---------------|
106
+ | `START_TIMER` | Start a countdown timer (cooking, stretch, etc.). | `<DURATION>`, `<TIME>` |
107
+ | `PAUSE_TIMER` | Pause an active countdown. | `<TIMER_REMAINING>`, `<DURATION>` |
108
+ | `CANCEL_TIMER` | Stop and clear a running timer. | `<TIMER_REMAINING>`, `<DURATION>` |
109
+ | `GET_TIMER_REMAINING` | Report time left on a timer or whether it has finished. | `<TIMER_REMAINING>`, `<DURATION>` |
110
+ | `START_STOPWATCH` | Open or start the stopwatch app. | *(none)* |
111
+ | `LAP_STOPWATCH` | Record a lap split on the stopwatch. | `<LAP_TIME>`, `<DURATION>` |
112
+ | `RESET_STOPWATCH` | Reset stopwatch to zero. | `<DURATION>`, `<LAP_TIME>` |
113
+
114
+ ---
115
+
116
+ ## Workout toggles
117
+
118
+ | Intent | What it does | Typical slots |
119
+ |--------|--------------|---------------|
120
+ | `START_WORKOUT` | Begin tracking a workout (walk, run, cycle, indoor cardio). | `<WORKOUT_TYPE>` |
121
+ | `PAUSE_WORKOUT` | Pause the active workout session. | `<WORKOUT_ELAPSED>`, `<WORKOUT_STATE>`, `<WORKOUT_TYPE>` |
122
+ | `RESUME_WORKOUT` | Resume a paused workout. | `<WORKOUT_ELAPSED>`, `<WORKOUT_STATE>`, `<WORKOUT_TYPE>` |
123
+ | `STOP_WORKOUT` | End the workout and save the session. | `<WORKOUT_ELAPSED>`, `<WORKOUT_TYPE>`, `<WORKOUT_STATE>` |
124
+ | `DISCARD_WORKOUT` | End the workout without saving data. | `<WORKOUT_TYPE>`, `<WORKOUT_STATE>` |
125
+ | `GET_WORKOUT_STATUS` | Check if a workout is running, paused, or left on by mistake; report elapsed time. | `<WORKOUT_ELAPSED>`, `<WORKOUT_STATE>`, `<WORKOUT_TYPE>` |
126
+ | `GET_WORKOUT_SUMMARY` | Show the post-workout summary screen (duration, distance, calories, HR). | `<WORKOUT_TYPE>`, `<WORKOUT_ELAPSED>`, `<DISTANCE_TODAY>`, `<DISTANCE_UNIT>`, `<CALORIES_ACTIVE>`, `<HR_AVG_SESSION>` |
127
+
128
+ ---
129
+
130
+ ## Non-action
131
+
132
+ | Intent | What it does | Typical slots |
133
+ |--------|--------------|---------------|
134
+ | `NONE` | Pure conversational reply β€” thanks, encouragement, clarification β€” with no sensor fetch or device command. | *(none)* |
135
+
136
+ ---
137
+
138
+ ## Summary
139
+
140
+ | Category | Intents | Count |
141
+ |----------|---------|------:|
142
+ | Steps | `GET_STEPS`, `SET_STEP_GOAL` | 2 |
143
+ | Distance | `GET_DISTANCE` | 1 |
144
+ | Heart rate | `GET_HEART_RATE`, `MEASURE_HEART_RATE` | 2 |
145
+ | Sleep | `GET_SLEEP`, `LOG_NAP` | 2 |
146
+ | Calories | `GET_CALORIES`, `SET_CALORIE_GOAL` | 2 |
147
+ | Active minutes | `GET_ACTIVE_MINUTES`, `SET_ACTIVE_GOAL` | 2 |
148
+ | Battery | `GET_BATTERY`, `ENABLE_POWER_SAVE`, `DISABLE_AOD` | 3 |
149
+ | Alarms / reminders | `SET_ALARM`, `LIST_ALARMS`, `DELETE_ALARM`, `SNOOZE_ALARM`, `SET_REMINDER`, `MUTE_REMINDERS`, `EXPLAIN_NUDGE` | 7 |
150
+ | Timers / stopwatch | `START_TIMER`, `PAUSE_TIMER`, `CANCEL_TIMER`, `GET_TIMER_REMAINING`, `START_STOPWATCH`, `LAP_STOPWATCH`, `RESET_STOPWATCH` | 7 |
151
+ | Workout | `START_WORKOUT`, `PAUSE_WORKOUT`, `RESUME_WORKOUT`, `STOP_WORKOUT`, `DISCARD_WORKOUT`, `GET_WORKOUT_STATUS`, `GET_WORKOUT_SUMMARY` | 7 |
152
+ | Non-action | `NONE` | 1 |
153
+ | **Total** | | **35** |
154
+
155
+ ---
156
+
157
+ ## Related files
158
+
159
+ - [`tinydata/schema.py`](../tinydata/schema.py) β€” canonical intent and placeholder definitions
160
+ - [`tinydata/topics.py`](../tinydata/topics.py) β€” 100 training scenarios mapped to intents and slots
161
+ - [`tinydata/generate.py`](../tinydata/generate.py) β€” synthetic conversation generator
162
+
163
+ Each wearable product may implement only a subset of these intents depending on hardware and firmware. Train and deploy against the intents your device actually supports.
docs/smartwatch-integration.md ADDED
@@ -0,0 +1,299 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Smartwatch Integration Guide β€” Export-0.1
2
+
3
+ How to run **Smartwatch LM v0.1** (`Export-0.1`) on a real wrist device and wire it to sensors, timers, and apps.
4
+
5
+ The model is a ~28M-parameter GPT exported as ONNX. It does **not** execute device actions itself β€” it emits **intent tags** and **slot placeholders** that your firmware or companion app parses and handles.
6
+
7
+ ---
8
+
9
+ ## What you ship
10
+
11
+ From this repo (repo root):
12
+
13
+ | Artifact | Size (approx.) | Use on watch |
14
+ |----------|----------------|--------------|
15
+ | `smartwatch_lm_merged.onnx` | ~52 MB | ONNX Runtime inference |
16
+ | `tokenizer.json` | ~200 KB | Text ↔ token ids |
17
+ | `tokenizer_config.json` | small | Tokenizer metadata |
18
+ | `config.json` | small | Architecture constants, I/O names |
19
+
20
+ Optional for training / debugging on a PC:
21
+
22
+ - `checkpoint.pt` β€” PyTorch weights
23
+ - `chat.py` β€” reference REPL with the same prompt format as production
24
+
25
+ **ONNX I/O** (from `config.json`):
26
+
27
+ - **Input:** `input_ids` β€” `int64`, shape `[batch, seq]`, max seq length **256**
28
+ - **Output:** `logits` β€” `float`, shape `[batch, seq, vocab_size]` (vocab **3524**)
29
+
30
+ Use the **last position** logits (`logits[batch, seq-1, :]`) to sample the next token, autoregressively, until EOS or max tokens.
31
+
32
+ ---
33
+
34
+ ## Architecture on the watch
35
+
36
+ ```mermaid
37
+ flowchart LR
38
+ subgraph input
39
+ MIC[Mic / touch]
40
+ STT[Speech-to-text optional]
41
+ end
42
+
43
+ subgraph lm [Smartwatch LM]
44
+ PROMPT[buildPrompt]
45
+ TOK[Tokenizer]
46
+ ONNX[ONNX Runtime]
47
+ PARSE[extractIntentReply]
48
+ FILL[fillSlots]
49
+ end
50
+
51
+ subgraph device
52
+ ROUTER[Intent router]
53
+ HANDLERS[Sensor / app handlers]
54
+ UI[Display / TTS]
55
+ end
56
+
57
+ MIC --> STT --> PROMPT
58
+ PROMPT --> TOK --> ONNX
59
+ ONNX --> TOK --> PARSE --> FILL
60
+ PARSE --> ROUTER --> HANDLERS
61
+ HANDLERS --> FILL
62
+ FILL --> UI
63
+ ```
64
+
65
+ **Separation of concerns:**
66
+
67
+ | Component | Responsibility |
68
+ |-----------|------------------|
69
+ | **LM** | Pick intent + reply template with `<SLOT>` tokens |
70
+ | **Router** | Map `<INTENT:NAME>` to a function / service |
71
+ | **Handlers** | Read/write real device state (steps, HR, alarms, …) |
72
+ | **Slot map** | `{"STEPS_TODAY": "4,231", "STEP_GOAL": "10,000", …}` from live data |
73
+ | **UI** | Show or speak the filled template string |
74
+
75
+ The model **never outputs real step counts or heart rates** β€” only placeholders. Your app must inject values before display.
76
+
77
+ See the full intent list in [Intent reference](./readme.md).
78
+
79
+ ---
80
+
81
+ ## Integration steps
82
+
83
+ ### Step 1 β€” Choose an inference runtime
84
+
85
+ | Platform | Suggested runtime | Notes |
86
+ |----------|-------------------|-------|
87
+ | **Wear OS** (Kotlin) | [ONNX Runtime Mobile](https://onnxruntime.ai/docs/get-started/with-mobile.html) | NNAPI / XNNPACK EP; quantize later for size |
88
+ | **watchOS** | ORT Mobile or Core ML conversion | 52 MB is heavy β€” consider INT8 quant or smaller context |
89
+ | **Samsung / Fitbit / Garmin** | Vendor SDK + ORT or TFLite | May need custom build; check RAM limits |
90
+ | **Companion phone** | Same model on phone, BLE to watch | Offloads compute; watch shows final text + runs handlers |
91
+ | **Browser / WebView** | Reference: [`collab-run-1/web`](../collab-run-1/web) | ONNX Runtime Web (WASM); good prototype |
92
+
93
+ Minimum RAM budget: model weights (~52 MB) + activations + tokenizer + app heap. Many watches need **quantization** (INT8) or **phone-side inference** to fit comfortably.
94
+
95
+ ### Step 2 β€” Port the generation loop
96
+
97
+ Mirror [`collab-run-1/web/src/model.ts`](../collab-run-1/web/src/model.ts) and [`config.ts`](../collab-run-1/web/src/config.ts):
98
+
99
+ ```text
100
+ TEMPERATURE = 0.5
101
+ TOP_K = 40
102
+ MAX_NEW_TOKENS = 40
103
+ BLOCK_SIZE = 256
104
+ ```
105
+
106
+ Loop:
107
+
108
+ 1. Encode prompt β†’ `input_ids`
109
+ 2. While under max tokens:
110
+ - Run ONNX with last ≀256 ids
111
+ - Sample next token (top-k + temperature)
112
+ - Append; stop on EOS (token id 0) after a few steps
113
+ 3. Decode **only new** token ids
114
+ 4. Post-process (see [Web output quality](./web-output-quality.md))
115
+
116
+ Reference Python loop: [`model.py`](../model.py) `GPT.generate()` and [`chat.py`](../chat.py) `ChatSession.say()`.
117
+
118
+ ### Step 3 β€” Implement prompt + history
119
+
120
+ ```text
121
+ user: How many steps today?
122
+ bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> β€” keep going!
123
+ user: Set a 10 minute timer
124
+ bot:
125
+ ```
126
+
127
+ Rules:
128
+
129
+ - Store **`rawBot`** in history: `<INTENT:GET_STEPS> You're at <STEPS_TODAY> of …` (slots unfilled)
130
+ - Do **not** store filled strings like `"You're at 4,231 of 10,000"` in history
131
+ - Trim history if encoded length nears 256 tokens (drop oldest turns)
132
+
133
+ Port [`buildPrompt()`](../collab-run-1/web/src/reply.ts) verbatim.
134
+
135
+ ### Step 4 β€” Parse intent and fill slots
136
+
137
+ Port from [`reply.ts`](../collab-run-1/web/src/reply.ts):
138
+
139
+ 1. `cleanReply(rawText)` β€” fix BPE / encoding glitches
140
+ 2. `extractIntentReply(cleaned)` β†’ `{ intent, template }`
141
+ 3. `fillSlots(template, slotMap)` β†’ user-visible string
142
+
143
+ Example slot map (refresh before each reply):
144
+
145
+ ```json
146
+ {
147
+ "STEPS_TODAY": "4,231",
148
+ "STEP_GOAL": "10,000",
149
+ "STEPS_REMAINING": "5,769",
150
+ "STEP_GOAL_PCT": "42%",
151
+ "TIME": "2:15 PM",
152
+ "DATE": "Sunday"
153
+ }
154
+ ```
155
+
156
+ Format numbers and units to match your locale β€” the model was trained on human-readable strings in synthetic data.
157
+
158
+ ### Step 5 β€” Intent router
159
+
160
+ Register one handler per intent your hardware supports (subset of 35 is fine):
161
+
162
+ ```kotlin
163
+ // Pseudocode β€” Wear OS style
164
+ fun dispatch(intent: String, template: String, slots: SlotMap): String {
165
+ val filled = fillSlots(template, slots)
166
+ when (intent) {
167
+ "GET_STEPS" -> { /* optional: refresh step cache */ }
168
+ "START_TIMER" -> timerService.start(parseDuration(slots))
169
+ "START_WORKOUT" -> workoutService.start(slots["WORKOUT_TYPE"])
170
+ "NONE" -> { /* no op */ }
171
+ else -> log.warn("Unsupported intent: $intent")
172
+ }
173
+ return filled
174
+ }
175
+ ```
176
+
177
+ **Validate intents** against an allowlist before calling handlers. If the model emits an unknown or unsupported intent, show the filled template anyway but skip the side effect, or ask the user to repeat.
178
+
179
+ Handler β†’ slot refresh pattern:
180
+
181
+ | Intent | Handler reads | Typical slots updated |
182
+ |--------|---------------|------------------------|
183
+ | `GET_STEPS` | Pedometer API | `STEPS_TODAY`, `STEP_GOAL`, `STEPS_REMAINING`, … |
184
+ | `GET_HEART_RATE` | HR sensor | `HR_CURRENT_BPM`, `HR_ZONE`, … |
185
+ | `GET_BATTERY` | Power manager | `BATTERY_PCT`, `BATTERY_DAYS_LEFT` |
186
+ | `START_TIMER` | Timer app | `TIMER_REMAINING`, `DURATION` |
187
+ | `SET_ALARM` | Alarm storage | `ALARM_TIME`, `ALARM_LABEL` |
188
+
189
+ Full mapping: [Intent reference](./readme.md).
190
+
191
+ ### Step 6 β€” User input path
192
+
193
+ | Input mode | Flow |
194
+ |------------|------|
195
+ | **Touch** | Typed or quick-reply chips β†’ `userMessage` string |
196
+ | **Voice** | On-watch or phone STT β†’ text β†’ same pipeline |
197
+ | **Hybrid** | STT on phone, LM on phone, send `{intent, displayText}` to watch via BLE |
198
+
199
+ Keep utterances short β€” training data mimics wrist-scale queries (β€œsteps today”, β€œstart a walk”, β€œbattery level”).
200
+
201
+ ### Step 7 β€” Output path
202
+
203
+ - **Display:** one or two lines on watch face / assistant overlay
204
+ - **TTS:** speak `fillSlots` result only (not raw `<INTENT:…>` tags unless you want β€œOpening timer” style debug)
205
+ - **Haptics:** optional pulse when `intent != NONE` and handler succeeds
206
+
207
+ Show intent in debug builds only (the web demo appends `[GET_STEPS]` for testing).
208
+
209
+ ---
210
+
211
+ ## Example end-to-end trace
212
+
213
+ **User:** β€œHow many steps do I need to hit my goal?”
214
+
215
+ 1. **Prompt** (with prior history if any):
216
+
217
+ ```text
218
+ user: How many steps do I need to hit my goal?
219
+ bot:
220
+ ```
221
+
222
+ 2. **Model generates** (raw decode):
223
+
224
+ ```text
225
+ <INTENT:GET_STEPS> You need <STEPS_REMAINING> more to reach <STEP_GOAL>.
226
+ ```
227
+
228
+ 3. **Parse:** intent = `GET_STEPS`, template = `You need <STEPS_REMAINING> more to reach <STEP_GOAL>.`
229
+
230
+ 4. **Handler:** `GET_STEPS` β†’ read pedometer + goals β†’ update slot map.
231
+
232
+ 5. **Fill slots:**
233
+
234
+ ```text
235
+ You need 5,769 more to reach 10,000.
236
+ ```
237
+
238
+ 6. **History append:**
239
+
240
+ ```text
241
+ rawBot = "<INTENT:GET_STEPS> You need <STEPS_REMAINING> more to reach <STEP_GOAL>."
242
+ ```
243
+
244
+ ---
245
+
246
+ ## Performance and size tips
247
+
248
+ | Technique | Benefit |
249
+ |-----------|---------|
250
+ | **INT8 quantization** | Cuts model size ~4Γ—; test intent accuracy after quant |
251
+ | **Phone-side inference** | Watch stays thin; model updates without OTA to watch |
252
+ | **Cache ONNX session** | Amortize 30–60s web load cost β€” on watch, load once at boot |
253
+ | **Limit history** | 2–4 turns usually enough; saves context window |
254
+ | **Lower max tokens** | 25–40 tokens keeps latency down on CPU |
255
+ | **Warm-up run** | One dummy forward pass after load avoids first-query stall |
256
+
257
+ Web reference timings: first WASM session init can take 30–60 seconds; generation is sequential token-by-token. Plan UX (spinner, β€œthinking…” glyph) accordingly.
258
+
259
+ ---
260
+
261
+ ## Testing without hardware
262
+
263
+ 1. **Browser demo** β€” [`collab-run-1/web`](../collab-run-1/web): `bun run copy-models && bun run dev`
264
+ 2. **Python REPL** β€” `python chat.py` (requires `torch`, `tokenizers`)
265
+ 3. **Dummy slot panel** β€” web Settings sidebar edits the same slot map your watch firmware should expose
266
+
267
+ Compare browser vs Python replies for the same prompts before flashing firmware.
268
+
269
+ ---
270
+
271
+ ## Production checklist
272
+
273
+ - [ ] ONNX + tokenizer bundled in app storage or downloaded once
274
+ - [ ] Tokenizer matches `Export-0.1/tokenizer.json` exactly
275
+ - [ ] Generation params match web reference (temp 0.5, top_k 40, max 40)
276
+ - [ ] `buildPrompt` / `extractIntentReply` / `fillSlots` ported or shared
277
+ - [ ] History stores unfilled bot lines only
278
+ - [ ] Intent allowlist matches implemented handlers
279
+ - [ ] Slot map refreshed from real sensors before `fillSlots`
280
+ - [ ] Graceful fallback for parse failures and unsupported intents
281
+ - [ ] Privacy: on-device inference β€” no cloud required for core loop
282
+ - [ ] Power: defer inference off critical paths; avoid running LM on every tick
283
+
284
+ ---
285
+
286
+ ## Limitations
287
+
288
+ - Trained on **synthetic** dialogs β€” real speech recognition errors and slang may reduce intent accuracy
289
+ - **No safety layer** β€” not for medical advice or safety-critical control
290
+ - **35 intents** β€” extend by fine-tuning on new data, not by prompt hacking alone
291
+ - **English-centric** BPE vocab β€” other languages need retraining
292
+
293
+ ---
294
+
295
+ ## Related docs
296
+
297
+ - [Web output quality](./web-output-quality.md) β€” sampling, cleanup, anti-gibberish techniques
298
+ - [Intent reference](./readme.md) β€” all intents and slots
299
+ - [Model README](../README.md) β€” model card and Hugging Face deployment
docs/web-output-quality.md ADDED
@@ -0,0 +1,201 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Web Output Quality β€” Avoiding Gibberish and Bad Replies
2
+
3
+ This document describes every technique the **Export-0.1 / browser demo** (`collab-run-1/web`) uses to keep model output readable, on-format, and useful. The web app is the reference runtime for the exported model; most of these patterns should be replicated on-device.
4
+
5
+ Related code lives in:
6
+
7
+ - [`collab-run-1/web/src/config.ts`](../collab-run-1/web/src/config.ts) β€” generation defaults
8
+ - [`collab-run-1/web/src/model.ts`](../collab-run-1/web/src/model.ts) β€” ONNX inference loop
9
+ - [`collab-run-1/web/src/reply.ts`](../collab-run-1/web/src/reply.ts) β€” post-processing
10
+ - [`collab-run-1/web/src/app.ts`](../collab-run-1/web/src/app.ts) β€” prompt assembly and slot fill
11
+ - [`chat.py`](../chat.py) β€” Python reference (`extract_bot_reply`)
12
+
13
+ ---
14
+
15
+ ## Overview
16
+
17
+ Small language models on constrained domains still drift: they ramble, invent fake numbers, corrupt BPE tokens, or start a new user turn mid-reply. The web stack handles this in four layers:
18
+
19
+ | Layer | Goal |
20
+ |-------|------|
21
+ | **Training & data** | Teach a narrow, structured format (intents + slots, no raw metrics) |
22
+ | **Sampling** | Keep generation short and low-randomness |
23
+ | **Prompt & history** | Anchor the model in `user:/bot:` turns |
24
+ | **Post-processing** | Clean tokenizer artifacts, parse one line, inject real sensor values |
25
+
26
+ The browser demo is stricter than the Python REPL in a few places (lower temperature, fewer tokens). That is intentional β€” WASM inference is slower, and shorter, more deterministic replies feel better on a wrist UI.
27
+
28
+ ---
29
+
30
+ ## 1. Constrained output format (training + runtime)
31
+
32
+ The model is not trained for open-ended chat. Every bot line follows:
33
+
34
+ ```
35
+ bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> β€” keep going!
36
+ ```
37
+
38
+ **Why this reduces gibberish:**
39
+
40
+ - **Intent tags** (`<INTENT:NAME>`) give the app a machine-readable action even if the natural-language tail is slightly off.
41
+ - **Slot placeholders** (`<STEPS_TODAY>`, `<HR_CURRENT_BPM>`, …) replace raw numbers. The model never learns to emit `"8432 steps"` β€” it learns token names your runtime fills from sensors. That prevents hallucinated metrics.
42
+ - **Closed vocabulary** β€” only 35 intents and a fixed placeholder set (see [`tinydata/schema.py`](../tinydata/schema.py)). Training data is validated so bot lines cannot invent tags like `<step_count>` or nest intents.
43
+
44
+ Training data generation ([`tinydata/generate.py`](../tinydata/generate.py)) rejects samples that:
45
+
46
+ - Miss the leading `<INTENT:…>` on bot lines
47
+ - Use unknown intents or placeholders
48
+ - Include raw numeric metrics (2+ digit numbers) in bot text
49
+ - Break the strict `user:` / `bot:` line format
50
+
51
+ Your runtime should **never show slot tokens raw to the user** β€” always run `fillSlots()` with live values (see [`reply.ts`](../collab-run-1/web/src/reply.ts)).
52
+
53
+ ---
54
+
55
+ ## 2. Sampling controls (inference)
56
+
57
+ Web defaults in [`config.ts`](../collab-run-1/web/src/config.ts):
58
+
59
+ | Parameter | Web value | Python REPL (Export-0.1) | Effect |
60
+ |-----------|-----------|--------------------------|--------|
61
+ | `TEMPERATURE` | **0.5** | 0.8 | Lower = less random token choice, fewer weird word combinations |
62
+ | `TOP_K` | 40 | 40 | Sample only from the 40 most likely next tokens |
63
+ | `MAX_NEW_TOKENS` | **40** | 120 | Hard cap on reply length β€” stops rambling early |
64
+ | `BLOCK_SIZE` | 256 | 256 | Context window; older tokens are dropped from the forward pass |
65
+
66
+ Implementation in [`model.ts`](../collab-run-1/web/src/model.ts):
67
+
68
+ - **Top-k sampling** β€” sort vocab by probability, renormalize over top 40, then sample. Cuts off the long tail of nonsense tokens.
69
+ - **Temperature scaling** β€” divide logits by `max(temperature, 1e-8)` before softmax.
70
+ - **EOS early stop** β€” after step 2, if the sampled token id is `0`, stop generating. Prevents run-on after an end-of-sequence signal.
71
+ - **Context trimming** β€” only the last 256 token ids are passed to ONNX each step (`inputIds.slice(-seqLen)`), matching training context limits.
72
+
73
+ For a smartwatch, prefer **web-style settings** (temp ~0.5, max tokens ~40) unless you need longer replies for voice.
74
+
75
+ ---
76
+
77
+ ## 3. Prompt format and conversation history
78
+
79
+ The model expects a fixed transcript shape:
80
+
81
+ ```
82
+ user: How many steps today?
83
+ bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> β€” keep going!
84
+ user: thanks
85
+ bot:
86
+ ```
87
+
88
+ [`buildPrompt()`](../collab-run-1/web/src/reply.ts) builds this string. Important details:
89
+
90
+ - Every turn uses lowercase `user:` and `bot:` prefixes with a trailing space after the colon.
91
+ - The current turn ends with `bot:` and **no** reply yet β€” the model continues from there.
92
+ - **History stores `rawBot`** β€” the intent tag and unfilled slot template, not the display string with dummy sensor values. If you put filled text like `"You're at 1,000 steps"` back into history, the model sees numbers it was never trained on and quality drops.
93
+
94
+ Only **new** tokens are decoded (`ids.slice(promptLen)`), so the reply extraction starts clean after the prompt.
95
+
96
+ ---
97
+
98
+ ## 4. Post-processing (`reply.ts`)
99
+
100
+ Raw decode output can still contain tokenizer glitches or extra lines. The web app runs a pipeline before display.
101
+
102
+ ### 4.1 `cleanReply(text)`
103
+
104
+ Fixes common BPE / UTF-8 artifacts:
105
+
106
+ - `Δ ` β†’ space (GPT-style space marker)
107
+ - `Ċ` β†’ newline
108
+ - Mojibake sequences for em dash and curly quotes β†’ `β€”`, `'`
109
+ - Collapse repeated spaces
110
+ - Normalize spaces **inside** angle brackets: `< INTENT : GET_STEPS >` β†’ `<INTENT:GET_STEPS>`
111
+ - Trim stray space-before-apostrophe: ` '` β†’ `'`
112
+
113
+ ### 4.2 `extractIntentReply(text)`
114
+
115
+ Parses the structured reply:
116
+
117
+ 1. Run `cleanReply`.
118
+ 2. Find the first `<INTENT:…>` tag (case-insensitive, tolerant of internal spaces).
119
+ 3. **Truncate at `\nuser:`** β€” if the model starts a fake next turn, discard everything after.
120
+ 4. Take **only the first line** β€” one utterance per reply, matching Python `extract_bot_reply`.
121
+ 5. Return `{ intent, template }` where `template` is the text after the intent tag.
122
+
123
+ If no intent tag is found, intent defaults to `NONE` and the first line is used as template.
124
+
125
+ ### 4.3 `fillSlots(template, data)`
126
+
127
+ Replace `<SLOT_NAME>` with values from your sensor map. Unknown slots stay as-is (useful for debugging).
128
+
129
+ ### 4.4 Python parity β€” `extract_bot_reply`
130
+
131
+ The Colab / Export-0.1 Python helper does the same job with prompt-relative slicing:
132
+
133
+ - Strip the prompt prefix from the full decode
134
+ - Stop at `\nuser:`, `\n\n`, or first newline
135
+ - Return a single-line bot utterance
136
+
137
+ The web README notes: *"Generation uses the same prompt format and cleanup as Colab `ask()`"* β€” with the extra `cleanReply` and intent parsing layers on top for browser display.
138
+
139
+ ---
140
+
141
+ ## 5. Infrastructure choices that affect quality
142
+
143
+ These are not sampling tricks, but bad infrastructure produces garbage logits or wrong token ids.
144
+
145
+ | Choice | Location | Why |
146
+ |--------|----------|-----|
147
+ | **Legacy ONNX export** (opset 17, `dynamo=False`) | [`export_onnx.py`](../collab-run-1/export_onnx.py) | Colab’s default exporter can produce invalid graphs; re-export fixes nonsense outputs in WASM |
148
+ | **Single merged ONNX file** | `smartwatch_lm_merged.onnx` | Browser fetch + ORT session needs one blob; external weights often fail in static hosting |
149
+ | **`@huggingface/tokenizers` (JS)** | [`tokenizer.ts`](../collab-run-1/web/src/tokenizer.ts) | Pure JS tokenizer matches training `tokenizer.json`; avoids ORT tokenizer conflicts |
150
+ | **`graphOptimizationLevel: "disabled"`** | [`model.ts`](../collab-run-1/web/src/model.ts) | Some WASM builds misbehave with aggressive graph opts on this model |
151
+ | **Decode only new tokens** | [`model.ts`](../collab-run-1/web/src/model.ts) | Prevents prompt text from leaking into the β€œreply” string |
152
+
153
+ ---
154
+
155
+ ## 6. Recommended runtime checklist
156
+
157
+ When porting off the web demo, implement this loop:
158
+
159
+ ```
160
+ 1. buildPrompt(history with raw bot lines, userMessage)
161
+ 2. encode β†’ generate (temp 0.5, top_k 40, max 40 tokens, EOS stop)
162
+ 3. decode new tokens only
163
+ 4. cleanReply β†’ extractIntentReply
164
+ 5. route intent to device handler (optional: refresh slot map from sensors)
165
+ 6. fillSlots(template, slotMap) β†’ show / speak to user
166
+ 7. append (userMessage, rawBotLine) to history
167
+ ```
168
+
169
+ **Do not:**
170
+
171
+ - Feed display strings with real numbers back into history
172
+ - Skip truncation at `\nuser:` or first newline
173
+ - Use high temperature on a 28M-param domain model
174
+ - Let the model invent metric values β€” always use slots
175
+
176
+ **Do:**
177
+
178
+ - Reset or trim history when context approaches 256 tokens
179
+ - Fall back to a safe message if intent is unknown or template is empty
180
+ - Keep generation caps tight for wrist UX (one short sentence)
181
+
182
+ ---
183
+
184
+ ## 7. What this does *not* fix
185
+
186
+ These techniques improve **format adherence and readability**, not general reasoning:
187
+
188
+ - Off-domain questions may still get weak or wrong intents
189
+ - Typos in user input are not corrected
190
+ - No content moderation or safety filter is applied
191
+ - Synthetic training data may not match real user phrasing
192
+
193
+ For production wearables, combine this model with **intent validation** (reject unknown intents), **handler guards** (don’t start a workout if one is already active), and optional **cloud fallback** for out-of-scope queries.
194
+
195
+ ---
196
+
197
+ ## Related docs
198
+
199
+ - [Intent reference](./readme.md) β€” all 35 intents and slot names
200
+ - [Smartwatch integration guide](./smartwatch-integration.md) β€” end-to-end device wiring
201
+ - [Export-0.1 README](../README.md) β€” model card and ONNX I/O