Add integration and output quality documentation
Browse files- README.md +10 -3
- docs/intent-reference.md +163 -0
- docs/smartwatch-integration.md +299 -0
- docs/web-output-quality.md +201 -0
README.md
CHANGED
|
@@ -89,12 +89,19 @@ The model predicts the next token. Slot placeholders like `<STEPS_TODAY>` are fi
|
|
| 89 |
|
| 90 |
On-device smartwatch / wearable assistant demos. Not intended for general-purpose chat or safety-critical applications.
|
| 91 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 92 |
## Limitations
|
| 93 |
|
| 94 |
- Trained on synthetic data; behavior on real user phrasing may vary
|
| 95 |
- Does not output real metric values β only intents and slot tokens
|
| 96 |
- Small model with limited reasoning capability
|
| 97 |
|
| 98 |
-
## Training
|
| 99 |
-
|
| 100 |
-
Trained with the `electron-v1` project (`collab-run-1`). See the parent repo for training instructions and intent reference.
|
|
|
|
| 89 |
|
| 90 |
On-device smartwatch / wearable assistant demos. Not intended for general-purpose chat or safety-critical applications.
|
| 91 |
|
| 92 |
+
## Documentation
|
| 93 |
+
|
| 94 |
+
| Guide | Description |
|
| 95 |
+
|-------|-------------|
|
| 96 |
+
| [Intent reference](docs/intent-reference.md) | All 35 intents and slot placeholders |
|
| 97 |
+
| [Web output quality](docs/web-output-quality.md) | Sampling, cleanup, and anti-gibberish techniques for browser/on-device runtimes |
|
| 98 |
+
| [Smartwatch integration](docs/smartwatch-integration.md) | End-to-end guide to wire the model into wrist hardware |
|
| 99 |
+
|
| 100 |
+
Live browser demo: see the `collab-run-1/web` app in the electron-v1 training project (links to this model on Hugging Face).
|
| 101 |
+
|
| 102 |
## Limitations
|
| 103 |
|
| 104 |
- Trained on synthetic data; behavior on real user phrasing may vary
|
| 105 |
- Does not output real metric values β only intents and slot tokens
|
| 106 |
- Small model with limited reasoning capability
|
| 107 |
|
|
|
|
|
|
|
|
|
docs/intent-reference.md
ADDED
|
@@ -0,0 +1,163 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Wrist Wearable Assistant β Intent Reference
|
| 2 |
+
|
| 3 |
+
This document describes every intent in the [`tinydata/schema.py`](../tinydata/schema.py) vocabulary. Intents are machine-readable action tags the on-device model emits; your app parses them and runs the matching handler.
|
| 4 |
+
|
| 5 |
+
## How intents work
|
| 6 |
+
|
| 7 |
+
Bot lines in training data and at runtime use this format:
|
| 8 |
+
|
| 9 |
+
```
|
| 10 |
+
bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> β keep going!
|
| 11 |
+
```
|
| 12 |
+
|
| 13 |
+
| Part | Role |
|
| 14 |
+
|------|------|
|
| 15 |
+
| `<INTENT:NAME>` | Tells the app **what to do** (fetch steps, start timer, etc.) |
|
| 16 |
+
| `<SLOT>` placeholders | Names **which values** the app injects from sensors/settings |
|
| 17 |
+
| Text after the tag | What the user **sees or hears** on the watch |
|
| 18 |
+
|
| 19 |
+
The model never outputs real metric values β only intents and slot tokens. Your runtime replaces slots with live data.
|
| 20 |
+
|
| 21 |
+
`NONE` is used when the bot only acknowledges or chats with no device action.
|
| 22 |
+
|
| 23 |
+
---
|
| 24 |
+
|
| 25 |
+
## Steps
|
| 26 |
+
|
| 27 |
+
| Intent | What it does | Typical slots |
|
| 28 |
+
|--------|--------------|---------------|
|
| 29 |
+
| `GET_STEPS` | Read step metrics from the pedometer: todayβs count, goals, streaks, week/month totals, comparisons. | `<STEPS_TODAY>`, `<STEPS_YESTERDAY>`, `<STEP_GOAL>`, `<STEPS_REMAINING>`, `<STEP_GOAL_PCT>`, `<STEP_STREAK_DAYS>`, `<STEPS_WEEK_TOTAL>`, `<STEPS_LAST_WEEK_TOTAL>`, `<STEPS_MONTH_TOTAL>`, `<TIME>`, `<DATE>`, `<PERCENT>` |
|
| 30 |
+
| `SET_STEP_GOAL` | Change the userβs daily step target. | `<STEP_GOAL>` |
|
| 31 |
+
|
| 32 |
+
---
|
| 33 |
+
|
| 34 |
+
## Distance
|
| 35 |
+
|
| 36 |
+
| Intent | What it does | Typical slots |
|
| 37 |
+
|--------|--------------|---------------|
|
| 38 |
+
| `GET_DISTANCE` | Read distance metrics: todayβs total, last session, units, weekly progress, lifetime mileage, hourly average. | `<DISTANCE_TODAY>`, `<DISTANCE_LAST_SESSION>`, `<DISTANCE_UNIT>`, `<DISTANCE_WEEK_REMAINING>`, `<DISTANCE_WEEK_BEST>`, `<DISTANCE_LIFETIME>`, `<DISTANCE_HOURLY_AVG>`, `<TIME>`, `<DATE>` |
|
| 39 |
+
|
| 40 |
+
---
|
| 41 |
+
|
| 42 |
+
## Heart rate
|
| 43 |
+
|
| 44 |
+
| Intent | What it does | Typical slots |
|
| 45 |
+
|--------|--------------|---------------|
|
| 46 |
+
| `GET_HEART_RATE` | Read heart rate data: current BPM, resting rate, overnight low, session average, daily peak, elevated minutes, effort zone. | `<HR_CURRENT_BPM>`, `<HR_RESTING_BPM>`, `<HR_RESTING_OVERNIGHT_LOW>`, `<HR_AVG_SESSION>`, `<HR_PEAK_TODAY>`, `<HR_ELEVATED_MINUTES>`, `<HR_ZONE>`, `<DURATION>` |
|
| 47 |
+
| `MEASURE_HEART_RATE` | Trigger a live on-demand pulse reading from the optical sensor. | `<HR_CURRENT_BPM>` |
|
| 48 |
+
|
| 49 |
+
---
|
| 50 |
+
|
| 51 |
+
## Sleep
|
| 52 |
+
|
| 53 |
+
| Intent | What it does | Typical slots |
|
| 54 |
+
|--------|--------------|---------------|
|
| 55 |
+
| `GET_SLEEP` | Read sleep data: duration last night, sleep/wake times, weekly average, awake minutes, goal gap, weekend total. | `<SLEEP_HOURS_LAST_NIGHT>`, `<SLEEP_START_TIME>`, `<SLEEP_WAKE_TIME>`, `<SLEEP_AVG_WEEK>`, `<SLEEP_AWAKE_MINUTES>`, `<SLEEP_GOAL_HOURS>`, `<SLEEP_WEEKEND_TOTAL>`, `<DATE>`, `<DURATION>` |
|
| 56 |
+
| `LOG_NAP` | Record or confirm a short nap session. | `<DURATION>`, `<SLEEP_HOURS_LAST_NIGHT>`, `<TIME>` |
|
| 57 |
+
|
| 58 |
+
---
|
| 59 |
+
|
| 60 |
+
## Calories
|
| 61 |
+
|
| 62 |
+
| Intent | What it does | Typical slots |
|
| 63 |
+
|--------|--------------|---------------|
|
| 64 |
+
| `GET_CALORIES` | Read calorie burn: daily total, active burn, remaining to goal, yesterday comparison, per-mile estimate. | `<CALORIES_TODAY>`, `<CALORIES_ACTIVE>`, `<CALORIES_REMAINING>`, `<CALORIE_GOAL>`, `<CALORIE_GOAL_PCT>`, `<CALORIES_YESTERDAY>`, `<CALORIES_PER_MILE>`, `<DISTANCE_UNIT>`, `<DURATION>` |
|
| 65 |
+
| `SET_CALORIE_GOAL` | Change the userβs daily active calorie burn target. | `<CALORIE_GOAL>` |
|
| 66 |
+
|
| 67 |
+
---
|
| 68 |
+
|
| 69 |
+
## Active minutes
|
| 70 |
+
|
| 71 |
+
| Intent | What it does | Typical slots |
|
| 72 |
+
|--------|--------------|---------------|
|
| 73 |
+
| `GET_ACTIVE_MINUTES` | Read movement time: todayβs active minutes, remaining to goal, weekly total, most active day. | `<ACTIVE_MINUTES_TODAY>`, `<ACTIVE_MINUTES_REMAINING>`, `<ACTIVE_MINUTES_GOAL>`, `<ACTIVE_MINUTES_WEEK>`, `<MOST_ACTIVE_DAY>`, `<DURATION>`, `<TIME>`, `<DATE>` |
|
| 74 |
+
| `SET_ACTIVE_GOAL` | Change the userβs daily active minutes target. | `<ACTIVE_MINUTES_GOAL>` |
|
| 75 |
+
|
| 76 |
+
---
|
| 77 |
+
|
| 78 |
+
## Battery
|
| 79 |
+
|
| 80 |
+
| Intent | What it does | Typical slots |
|
| 81 |
+
|--------|--------------|---------------|
|
| 82 |
+
| `GET_BATTERY` | Read power state: percentage, estimated days left, days since charge, charge complete status. | `<BATTERY_PCT>`, `<BATTERY_DAYS_LEFT>`, `<DAYS_SINCE_CHARGE>`, `<CHARGE_COMPLETE>`, `<DURATION>`, `<DATE>` |
|
| 83 |
+
| `ENABLE_POWER_SAVE` | Turn on low-power / battery saver mode on the device. | `<BATTERY_PCT>` |
|
| 84 |
+
| `DISABLE_AOD` | Turn off always-on display to reduce power draw. | `<BATTERY_PCT>` |
|
| 85 |
+
|
| 86 |
+
---
|
| 87 |
+
|
| 88 |
+
## Alarms and reminders
|
| 89 |
+
|
| 90 |
+
| Intent | What it does | Typical slots |
|
| 91 |
+
|--------|--------------|---------------|
|
| 92 |
+
| `SET_ALARM` | Create or update a wake alarm at a given time. | `<ALARM_TIME>`, `<ALARM_LABEL>`, `<DATE>`, `<TIME>` |
|
| 93 |
+
| `LIST_ALARMS` | Show all alarms currently stored on the device. | `<ALARM_TIME>`, `<ALARM_LABEL>` |
|
| 94 |
+
| `DELETE_ALARM` | Remove one or all alarms. | `<ALARM_LABEL>`, `<ALARM_TIME>` |
|
| 95 |
+
| `SNOOZE_ALARM` | Snooze the currently firing alarm for a short period. | `<ALARM_TIME>`, `<DURATION>`, `<ALARM_LABEL>` |
|
| 96 |
+
| `SET_REMINDER` | Schedule a recurring reminder (water, bedtime wind-down, move prompts). | `<ALARM_TIME>`, `<ALARM_LABEL>`, `<REMINDER_INTERVAL>`, `<TIME>` |
|
| 97 |
+
| `MUTE_REMINDERS` | Disable hourly move reminders or similar nudges. | *(none)* |
|
| 98 |
+
| `EXPLAIN_NUDGE` | Explain why the watch just vibrated (e.g. sedentary move nudge). | `<NUDGE_REASON>`, `<DURATION>`, `<TIME>` |
|
| 99 |
+
|
| 100 |
+
---
|
| 101 |
+
|
| 102 |
+
## Timers and stopwatch
|
| 103 |
+
|
| 104 |
+
| Intent | What it does | Typical slots |
|
| 105 |
+
|--------|--------------|---------------|
|
| 106 |
+
| `START_TIMER` | Start a countdown timer (cooking, stretch, etc.). | `<DURATION>`, `<TIME>` |
|
| 107 |
+
| `PAUSE_TIMER` | Pause an active countdown. | `<TIMER_REMAINING>`, `<DURATION>` |
|
| 108 |
+
| `CANCEL_TIMER` | Stop and clear a running timer. | `<TIMER_REMAINING>`, `<DURATION>` |
|
| 109 |
+
| `GET_TIMER_REMAINING` | Report time left on a timer or whether it has finished. | `<TIMER_REMAINING>`, `<DURATION>` |
|
| 110 |
+
| `START_STOPWATCH` | Open or start the stopwatch app. | *(none)* |
|
| 111 |
+
| `LAP_STOPWATCH` | Record a lap split on the stopwatch. | `<LAP_TIME>`, `<DURATION>` |
|
| 112 |
+
| `RESET_STOPWATCH` | Reset stopwatch to zero. | `<DURATION>`, `<LAP_TIME>` |
|
| 113 |
+
|
| 114 |
+
---
|
| 115 |
+
|
| 116 |
+
## Workout toggles
|
| 117 |
+
|
| 118 |
+
| Intent | What it does | Typical slots |
|
| 119 |
+
|--------|--------------|---------------|
|
| 120 |
+
| `START_WORKOUT` | Begin tracking a workout (walk, run, cycle, indoor cardio). | `<WORKOUT_TYPE>` |
|
| 121 |
+
| `PAUSE_WORKOUT` | Pause the active workout session. | `<WORKOUT_ELAPSED>`, `<WORKOUT_STATE>`, `<WORKOUT_TYPE>` |
|
| 122 |
+
| `RESUME_WORKOUT` | Resume a paused workout. | `<WORKOUT_ELAPSED>`, `<WORKOUT_STATE>`, `<WORKOUT_TYPE>` |
|
| 123 |
+
| `STOP_WORKOUT` | End the workout and save the session. | `<WORKOUT_ELAPSED>`, `<WORKOUT_TYPE>`, `<WORKOUT_STATE>` |
|
| 124 |
+
| `DISCARD_WORKOUT` | End the workout without saving data. | `<WORKOUT_TYPE>`, `<WORKOUT_STATE>` |
|
| 125 |
+
| `GET_WORKOUT_STATUS` | Check if a workout is running, paused, or left on by mistake; report elapsed time. | `<WORKOUT_ELAPSED>`, `<WORKOUT_STATE>`, `<WORKOUT_TYPE>` |
|
| 126 |
+
| `GET_WORKOUT_SUMMARY` | Show the post-workout summary screen (duration, distance, calories, HR). | `<WORKOUT_TYPE>`, `<WORKOUT_ELAPSED>`, `<DISTANCE_TODAY>`, `<DISTANCE_UNIT>`, `<CALORIES_ACTIVE>`, `<HR_AVG_SESSION>` |
|
| 127 |
+
|
| 128 |
+
---
|
| 129 |
+
|
| 130 |
+
## Non-action
|
| 131 |
+
|
| 132 |
+
| Intent | What it does | Typical slots |
|
| 133 |
+
|--------|--------------|---------------|
|
| 134 |
+
| `NONE` | Pure conversational reply β thanks, encouragement, clarification β with no sensor fetch or device command. | *(none)* |
|
| 135 |
+
|
| 136 |
+
---
|
| 137 |
+
|
| 138 |
+
## Summary
|
| 139 |
+
|
| 140 |
+
| Category | Intents | Count |
|
| 141 |
+
|----------|---------|------:|
|
| 142 |
+
| Steps | `GET_STEPS`, `SET_STEP_GOAL` | 2 |
|
| 143 |
+
| Distance | `GET_DISTANCE` | 1 |
|
| 144 |
+
| Heart rate | `GET_HEART_RATE`, `MEASURE_HEART_RATE` | 2 |
|
| 145 |
+
| Sleep | `GET_SLEEP`, `LOG_NAP` | 2 |
|
| 146 |
+
| Calories | `GET_CALORIES`, `SET_CALORIE_GOAL` | 2 |
|
| 147 |
+
| Active minutes | `GET_ACTIVE_MINUTES`, `SET_ACTIVE_GOAL` | 2 |
|
| 148 |
+
| Battery | `GET_BATTERY`, `ENABLE_POWER_SAVE`, `DISABLE_AOD` | 3 |
|
| 149 |
+
| Alarms / reminders | `SET_ALARM`, `LIST_ALARMS`, `DELETE_ALARM`, `SNOOZE_ALARM`, `SET_REMINDER`, `MUTE_REMINDERS`, `EXPLAIN_NUDGE` | 7 |
|
| 150 |
+
| Timers / stopwatch | `START_TIMER`, `PAUSE_TIMER`, `CANCEL_TIMER`, `GET_TIMER_REMAINING`, `START_STOPWATCH`, `LAP_STOPWATCH`, `RESET_STOPWATCH` | 7 |
|
| 151 |
+
| Workout | `START_WORKOUT`, `PAUSE_WORKOUT`, `RESUME_WORKOUT`, `STOP_WORKOUT`, `DISCARD_WORKOUT`, `GET_WORKOUT_STATUS`, `GET_WORKOUT_SUMMARY` | 7 |
|
| 152 |
+
| Non-action | `NONE` | 1 |
|
| 153 |
+
| **Total** | | **35** |
|
| 154 |
+
|
| 155 |
+
---
|
| 156 |
+
|
| 157 |
+
## Related files
|
| 158 |
+
|
| 159 |
+
- [`tinydata/schema.py`](../tinydata/schema.py) β canonical intent and placeholder definitions
|
| 160 |
+
- [`tinydata/topics.py`](../tinydata/topics.py) β 100 training scenarios mapped to intents and slots
|
| 161 |
+
- [`tinydata/generate.py`](../tinydata/generate.py) β synthetic conversation generator
|
| 162 |
+
|
| 163 |
+
Each wearable product may implement only a subset of these intents depending on hardware and firmware. Train and deploy against the intents your device actually supports.
|
docs/smartwatch-integration.md
ADDED
|
@@ -0,0 +1,299 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Smartwatch Integration Guide β Export-0.1
|
| 2 |
+
|
| 3 |
+
How to run **Smartwatch LM v0.1** (`Export-0.1`) on a real wrist device and wire it to sensors, timers, and apps.
|
| 4 |
+
|
| 5 |
+
The model is a ~28M-parameter GPT exported as ONNX. It does **not** execute device actions itself β it emits **intent tags** and **slot placeholders** that your firmware or companion app parses and handles.
|
| 6 |
+
|
| 7 |
+
---
|
| 8 |
+
|
| 9 |
+
## What you ship
|
| 10 |
+
|
| 11 |
+
From this repo (repo root):
|
| 12 |
+
|
| 13 |
+
| Artifact | Size (approx.) | Use on watch |
|
| 14 |
+
|----------|----------------|--------------|
|
| 15 |
+
| `smartwatch_lm_merged.onnx` | ~52 MB | ONNX Runtime inference |
|
| 16 |
+
| `tokenizer.json` | ~200 KB | Text β token ids |
|
| 17 |
+
| `tokenizer_config.json` | small | Tokenizer metadata |
|
| 18 |
+
| `config.json` | small | Architecture constants, I/O names |
|
| 19 |
+
|
| 20 |
+
Optional for training / debugging on a PC:
|
| 21 |
+
|
| 22 |
+
- `checkpoint.pt` β PyTorch weights
|
| 23 |
+
- `chat.py` β reference REPL with the same prompt format as production
|
| 24 |
+
|
| 25 |
+
**ONNX I/O** (from `config.json`):
|
| 26 |
+
|
| 27 |
+
- **Input:** `input_ids` β `int64`, shape `[batch, seq]`, max seq length **256**
|
| 28 |
+
- **Output:** `logits` β `float`, shape `[batch, seq, vocab_size]` (vocab **3524**)
|
| 29 |
+
|
| 30 |
+
Use the **last position** logits (`logits[batch, seq-1, :]`) to sample the next token, autoregressively, until EOS or max tokens.
|
| 31 |
+
|
| 32 |
+
---
|
| 33 |
+
|
| 34 |
+
## Architecture on the watch
|
| 35 |
+
|
| 36 |
+
```mermaid
|
| 37 |
+
flowchart LR
|
| 38 |
+
subgraph input
|
| 39 |
+
MIC[Mic / touch]
|
| 40 |
+
STT[Speech-to-text optional]
|
| 41 |
+
end
|
| 42 |
+
|
| 43 |
+
subgraph lm [Smartwatch LM]
|
| 44 |
+
PROMPT[buildPrompt]
|
| 45 |
+
TOK[Tokenizer]
|
| 46 |
+
ONNX[ONNX Runtime]
|
| 47 |
+
PARSE[extractIntentReply]
|
| 48 |
+
FILL[fillSlots]
|
| 49 |
+
end
|
| 50 |
+
|
| 51 |
+
subgraph device
|
| 52 |
+
ROUTER[Intent router]
|
| 53 |
+
HANDLERS[Sensor / app handlers]
|
| 54 |
+
UI[Display / TTS]
|
| 55 |
+
end
|
| 56 |
+
|
| 57 |
+
MIC --> STT --> PROMPT
|
| 58 |
+
PROMPT --> TOK --> ONNX
|
| 59 |
+
ONNX --> TOK --> PARSE --> FILL
|
| 60 |
+
PARSE --> ROUTER --> HANDLERS
|
| 61 |
+
HANDLERS --> FILL
|
| 62 |
+
FILL --> UI
|
| 63 |
+
```
|
| 64 |
+
|
| 65 |
+
**Separation of concerns:**
|
| 66 |
+
|
| 67 |
+
| Component | Responsibility |
|
| 68 |
+
|-----------|------------------|
|
| 69 |
+
| **LM** | Pick intent + reply template with `<SLOT>` tokens |
|
| 70 |
+
| **Router** | Map `<INTENT:NAME>` to a function / service |
|
| 71 |
+
| **Handlers** | Read/write real device state (steps, HR, alarms, β¦) |
|
| 72 |
+
| **Slot map** | `{"STEPS_TODAY": "4,231", "STEP_GOAL": "10,000", β¦}` from live data |
|
| 73 |
+
| **UI** | Show or speak the filled template string |
|
| 74 |
+
|
| 75 |
+
The model **never outputs real step counts or heart rates** β only placeholders. Your app must inject values before display.
|
| 76 |
+
|
| 77 |
+
See the full intent list in [Intent reference](./readme.md).
|
| 78 |
+
|
| 79 |
+
---
|
| 80 |
+
|
| 81 |
+
## Integration steps
|
| 82 |
+
|
| 83 |
+
### Step 1 β Choose an inference runtime
|
| 84 |
+
|
| 85 |
+
| Platform | Suggested runtime | Notes |
|
| 86 |
+
|----------|-------------------|-------|
|
| 87 |
+
| **Wear OS** (Kotlin) | [ONNX Runtime Mobile](https://onnxruntime.ai/docs/get-started/with-mobile.html) | NNAPI / XNNPACK EP; quantize later for size |
|
| 88 |
+
| **watchOS** | ORT Mobile or Core ML conversion | 52 MB is heavy β consider INT8 quant or smaller context |
|
| 89 |
+
| **Samsung / Fitbit / Garmin** | Vendor SDK + ORT or TFLite | May need custom build; check RAM limits |
|
| 90 |
+
| **Companion phone** | Same model on phone, BLE to watch | Offloads compute; watch shows final text + runs handlers |
|
| 91 |
+
| **Browser / WebView** | Reference: [`collab-run-1/web`](../collab-run-1/web) | ONNX Runtime Web (WASM); good prototype |
|
| 92 |
+
|
| 93 |
+
Minimum RAM budget: model weights (~52 MB) + activations + tokenizer + app heap. Many watches need **quantization** (INT8) or **phone-side inference** to fit comfortably.
|
| 94 |
+
|
| 95 |
+
### Step 2 β Port the generation loop
|
| 96 |
+
|
| 97 |
+
Mirror [`collab-run-1/web/src/model.ts`](../collab-run-1/web/src/model.ts) and [`config.ts`](../collab-run-1/web/src/config.ts):
|
| 98 |
+
|
| 99 |
+
```text
|
| 100 |
+
TEMPERATURE = 0.5
|
| 101 |
+
TOP_K = 40
|
| 102 |
+
MAX_NEW_TOKENS = 40
|
| 103 |
+
BLOCK_SIZE = 256
|
| 104 |
+
```
|
| 105 |
+
|
| 106 |
+
Loop:
|
| 107 |
+
|
| 108 |
+
1. Encode prompt β `input_ids`
|
| 109 |
+
2. While under max tokens:
|
| 110 |
+
- Run ONNX with last β€256 ids
|
| 111 |
+
- Sample next token (top-k + temperature)
|
| 112 |
+
- Append; stop on EOS (token id 0) after a few steps
|
| 113 |
+
3. Decode **only new** token ids
|
| 114 |
+
4. Post-process (see [Web output quality](./web-output-quality.md))
|
| 115 |
+
|
| 116 |
+
Reference Python loop: [`model.py`](../model.py) `GPT.generate()` and [`chat.py`](../chat.py) `ChatSession.say()`.
|
| 117 |
+
|
| 118 |
+
### Step 3 β Implement prompt + history
|
| 119 |
+
|
| 120 |
+
```text
|
| 121 |
+
user: How many steps today?
|
| 122 |
+
bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> β keep going!
|
| 123 |
+
user: Set a 10 minute timer
|
| 124 |
+
bot:
|
| 125 |
+
```
|
| 126 |
+
|
| 127 |
+
Rules:
|
| 128 |
+
|
| 129 |
+
- Store **`rawBot`** in history: `<INTENT:GET_STEPS> You're at <STEPS_TODAY> of β¦` (slots unfilled)
|
| 130 |
+
- Do **not** store filled strings like `"You're at 4,231 of 10,000"` in history
|
| 131 |
+
- Trim history if encoded length nears 256 tokens (drop oldest turns)
|
| 132 |
+
|
| 133 |
+
Port [`buildPrompt()`](../collab-run-1/web/src/reply.ts) verbatim.
|
| 134 |
+
|
| 135 |
+
### Step 4 β Parse intent and fill slots
|
| 136 |
+
|
| 137 |
+
Port from [`reply.ts`](../collab-run-1/web/src/reply.ts):
|
| 138 |
+
|
| 139 |
+
1. `cleanReply(rawText)` β fix BPE / encoding glitches
|
| 140 |
+
2. `extractIntentReply(cleaned)` β `{ intent, template }`
|
| 141 |
+
3. `fillSlots(template, slotMap)` β user-visible string
|
| 142 |
+
|
| 143 |
+
Example slot map (refresh before each reply):
|
| 144 |
+
|
| 145 |
+
```json
|
| 146 |
+
{
|
| 147 |
+
"STEPS_TODAY": "4,231",
|
| 148 |
+
"STEP_GOAL": "10,000",
|
| 149 |
+
"STEPS_REMAINING": "5,769",
|
| 150 |
+
"STEP_GOAL_PCT": "42%",
|
| 151 |
+
"TIME": "2:15 PM",
|
| 152 |
+
"DATE": "Sunday"
|
| 153 |
+
}
|
| 154 |
+
```
|
| 155 |
+
|
| 156 |
+
Format numbers and units to match your locale β the model was trained on human-readable strings in synthetic data.
|
| 157 |
+
|
| 158 |
+
### Step 5 β Intent router
|
| 159 |
+
|
| 160 |
+
Register one handler per intent your hardware supports (subset of 35 is fine):
|
| 161 |
+
|
| 162 |
+
```kotlin
|
| 163 |
+
// Pseudocode β Wear OS style
|
| 164 |
+
fun dispatch(intent: String, template: String, slots: SlotMap): String {
|
| 165 |
+
val filled = fillSlots(template, slots)
|
| 166 |
+
when (intent) {
|
| 167 |
+
"GET_STEPS" -> { /* optional: refresh step cache */ }
|
| 168 |
+
"START_TIMER" -> timerService.start(parseDuration(slots))
|
| 169 |
+
"START_WORKOUT" -> workoutService.start(slots["WORKOUT_TYPE"])
|
| 170 |
+
"NONE" -> { /* no op */ }
|
| 171 |
+
else -> log.warn("Unsupported intent: $intent")
|
| 172 |
+
}
|
| 173 |
+
return filled
|
| 174 |
+
}
|
| 175 |
+
```
|
| 176 |
+
|
| 177 |
+
**Validate intents** against an allowlist before calling handlers. If the model emits an unknown or unsupported intent, show the filled template anyway but skip the side effect, or ask the user to repeat.
|
| 178 |
+
|
| 179 |
+
Handler β slot refresh pattern:
|
| 180 |
+
|
| 181 |
+
| Intent | Handler reads | Typical slots updated |
|
| 182 |
+
|--------|---------------|------------------------|
|
| 183 |
+
| `GET_STEPS` | Pedometer API | `STEPS_TODAY`, `STEP_GOAL`, `STEPS_REMAINING`, β¦ |
|
| 184 |
+
| `GET_HEART_RATE` | HR sensor | `HR_CURRENT_BPM`, `HR_ZONE`, β¦ |
|
| 185 |
+
| `GET_BATTERY` | Power manager | `BATTERY_PCT`, `BATTERY_DAYS_LEFT` |
|
| 186 |
+
| `START_TIMER` | Timer app | `TIMER_REMAINING`, `DURATION` |
|
| 187 |
+
| `SET_ALARM` | Alarm storage | `ALARM_TIME`, `ALARM_LABEL` |
|
| 188 |
+
|
| 189 |
+
Full mapping: [Intent reference](./readme.md).
|
| 190 |
+
|
| 191 |
+
### Step 6 β User input path
|
| 192 |
+
|
| 193 |
+
| Input mode | Flow |
|
| 194 |
+
|------------|------|
|
| 195 |
+
| **Touch** | Typed or quick-reply chips β `userMessage` string |
|
| 196 |
+
| **Voice** | On-watch or phone STT β text β same pipeline |
|
| 197 |
+
| **Hybrid** | STT on phone, LM on phone, send `{intent, displayText}` to watch via BLE |
|
| 198 |
+
|
| 199 |
+
Keep utterances short β training data mimics wrist-scale queries (βsteps todayβ, βstart a walkβ, βbattery levelβ).
|
| 200 |
+
|
| 201 |
+
### Step 7 β Output path
|
| 202 |
+
|
| 203 |
+
- **Display:** one or two lines on watch face / assistant overlay
|
| 204 |
+
- **TTS:** speak `fillSlots` result only (not raw `<INTENT:β¦>` tags unless you want βOpening timerβ style debug)
|
| 205 |
+
- **Haptics:** optional pulse when `intent != NONE` and handler succeeds
|
| 206 |
+
|
| 207 |
+
Show intent in debug builds only (the web demo appends `[GET_STEPS]` for testing).
|
| 208 |
+
|
| 209 |
+
---
|
| 210 |
+
|
| 211 |
+
## Example end-to-end trace
|
| 212 |
+
|
| 213 |
+
**User:** βHow many steps do I need to hit my goal?β
|
| 214 |
+
|
| 215 |
+
1. **Prompt** (with prior history if any):
|
| 216 |
+
|
| 217 |
+
```text
|
| 218 |
+
user: How many steps do I need to hit my goal?
|
| 219 |
+
bot:
|
| 220 |
+
```
|
| 221 |
+
|
| 222 |
+
2. **Model generates** (raw decode):
|
| 223 |
+
|
| 224 |
+
```text
|
| 225 |
+
<INTENT:GET_STEPS> You need <STEPS_REMAINING> more to reach <STEP_GOAL>.
|
| 226 |
+
```
|
| 227 |
+
|
| 228 |
+
3. **Parse:** intent = `GET_STEPS`, template = `You need <STEPS_REMAINING> more to reach <STEP_GOAL>.`
|
| 229 |
+
|
| 230 |
+
4. **Handler:** `GET_STEPS` β read pedometer + goals β update slot map.
|
| 231 |
+
|
| 232 |
+
5. **Fill slots:**
|
| 233 |
+
|
| 234 |
+
```text
|
| 235 |
+
You need 5,769 more to reach 10,000.
|
| 236 |
+
```
|
| 237 |
+
|
| 238 |
+
6. **History append:**
|
| 239 |
+
|
| 240 |
+
```text
|
| 241 |
+
rawBot = "<INTENT:GET_STEPS> You need <STEPS_REMAINING> more to reach <STEP_GOAL>."
|
| 242 |
+
```
|
| 243 |
+
|
| 244 |
+
---
|
| 245 |
+
|
| 246 |
+
## Performance and size tips
|
| 247 |
+
|
| 248 |
+
| Technique | Benefit |
|
| 249 |
+
|-----------|---------|
|
| 250 |
+
| **INT8 quantization** | Cuts model size ~4Γ; test intent accuracy after quant |
|
| 251 |
+
| **Phone-side inference** | Watch stays thin; model updates without OTA to watch |
|
| 252 |
+
| **Cache ONNX session** | Amortize 30β60s web load cost β on watch, load once at boot |
|
| 253 |
+
| **Limit history** | 2β4 turns usually enough; saves context window |
|
| 254 |
+
| **Lower max tokens** | 25β40 tokens keeps latency down on CPU |
|
| 255 |
+
| **Warm-up run** | One dummy forward pass after load avoids first-query stall |
|
| 256 |
+
|
| 257 |
+
Web reference timings: first WASM session init can take 30β60 seconds; generation is sequential token-by-token. Plan UX (spinner, βthinkingβ¦β glyph) accordingly.
|
| 258 |
+
|
| 259 |
+
---
|
| 260 |
+
|
| 261 |
+
## Testing without hardware
|
| 262 |
+
|
| 263 |
+
1. **Browser demo** β [`collab-run-1/web`](../collab-run-1/web): `bun run copy-models && bun run dev`
|
| 264 |
+
2. **Python REPL** β `python chat.py` (requires `torch`, `tokenizers`)
|
| 265 |
+
3. **Dummy slot panel** β web Settings sidebar edits the same slot map your watch firmware should expose
|
| 266 |
+
|
| 267 |
+
Compare browser vs Python replies for the same prompts before flashing firmware.
|
| 268 |
+
|
| 269 |
+
---
|
| 270 |
+
|
| 271 |
+
## Production checklist
|
| 272 |
+
|
| 273 |
+
- [ ] ONNX + tokenizer bundled in app storage or downloaded once
|
| 274 |
+
- [ ] Tokenizer matches `Export-0.1/tokenizer.json` exactly
|
| 275 |
+
- [ ] Generation params match web reference (temp 0.5, top_k 40, max 40)
|
| 276 |
+
- [ ] `buildPrompt` / `extractIntentReply` / `fillSlots` ported or shared
|
| 277 |
+
- [ ] History stores unfilled bot lines only
|
| 278 |
+
- [ ] Intent allowlist matches implemented handlers
|
| 279 |
+
- [ ] Slot map refreshed from real sensors before `fillSlots`
|
| 280 |
+
- [ ] Graceful fallback for parse failures and unsupported intents
|
| 281 |
+
- [ ] Privacy: on-device inference β no cloud required for core loop
|
| 282 |
+
- [ ] Power: defer inference off critical paths; avoid running LM on every tick
|
| 283 |
+
|
| 284 |
+
---
|
| 285 |
+
|
| 286 |
+
## Limitations
|
| 287 |
+
|
| 288 |
+
- Trained on **synthetic** dialogs β real speech recognition errors and slang may reduce intent accuracy
|
| 289 |
+
- **No safety layer** β not for medical advice or safety-critical control
|
| 290 |
+
- **35 intents** β extend by fine-tuning on new data, not by prompt hacking alone
|
| 291 |
+
- **English-centric** BPE vocab β other languages need retraining
|
| 292 |
+
|
| 293 |
+
---
|
| 294 |
+
|
| 295 |
+
## Related docs
|
| 296 |
+
|
| 297 |
+
- [Web output quality](./web-output-quality.md) β sampling, cleanup, anti-gibberish techniques
|
| 298 |
+
- [Intent reference](./readme.md) β all intents and slots
|
| 299 |
+
- [Model README](../README.md) β model card and Hugging Face deployment
|
docs/web-output-quality.md
ADDED
|
@@ -0,0 +1,201 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Web Output Quality β Avoiding Gibberish and Bad Replies
|
| 2 |
+
|
| 3 |
+
This document describes every technique the **Export-0.1 / browser demo** (`collab-run-1/web`) uses to keep model output readable, on-format, and useful. The web app is the reference runtime for the exported model; most of these patterns should be replicated on-device.
|
| 4 |
+
|
| 5 |
+
Related code lives in:
|
| 6 |
+
|
| 7 |
+
- [`collab-run-1/web/src/config.ts`](../collab-run-1/web/src/config.ts) β generation defaults
|
| 8 |
+
- [`collab-run-1/web/src/model.ts`](../collab-run-1/web/src/model.ts) β ONNX inference loop
|
| 9 |
+
- [`collab-run-1/web/src/reply.ts`](../collab-run-1/web/src/reply.ts) β post-processing
|
| 10 |
+
- [`collab-run-1/web/src/app.ts`](../collab-run-1/web/src/app.ts) β prompt assembly and slot fill
|
| 11 |
+
- [`chat.py`](../chat.py) β Python reference (`extract_bot_reply`)
|
| 12 |
+
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
## Overview
|
| 16 |
+
|
| 17 |
+
Small language models on constrained domains still drift: they ramble, invent fake numbers, corrupt BPE tokens, or start a new user turn mid-reply. The web stack handles this in four layers:
|
| 18 |
+
|
| 19 |
+
| Layer | Goal |
|
| 20 |
+
|-------|------|
|
| 21 |
+
| **Training & data** | Teach a narrow, structured format (intents + slots, no raw metrics) |
|
| 22 |
+
| **Sampling** | Keep generation short and low-randomness |
|
| 23 |
+
| **Prompt & history** | Anchor the model in `user:/bot:` turns |
|
| 24 |
+
| **Post-processing** | Clean tokenizer artifacts, parse one line, inject real sensor values |
|
| 25 |
+
|
| 26 |
+
The browser demo is stricter than the Python REPL in a few places (lower temperature, fewer tokens). That is intentional β WASM inference is slower, and shorter, more deterministic replies feel better on a wrist UI.
|
| 27 |
+
|
| 28 |
+
---
|
| 29 |
+
|
| 30 |
+
## 1. Constrained output format (training + runtime)
|
| 31 |
+
|
| 32 |
+
The model is not trained for open-ended chat. Every bot line follows:
|
| 33 |
+
|
| 34 |
+
```
|
| 35 |
+
bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> β keep going!
|
| 36 |
+
```
|
| 37 |
+
|
| 38 |
+
**Why this reduces gibberish:**
|
| 39 |
+
|
| 40 |
+
- **Intent tags** (`<INTENT:NAME>`) give the app a machine-readable action even if the natural-language tail is slightly off.
|
| 41 |
+
- **Slot placeholders** (`<STEPS_TODAY>`, `<HR_CURRENT_BPM>`, β¦) replace raw numbers. The model never learns to emit `"8432 steps"` β it learns token names your runtime fills from sensors. That prevents hallucinated metrics.
|
| 42 |
+
- **Closed vocabulary** β only 35 intents and a fixed placeholder set (see [`tinydata/schema.py`](../tinydata/schema.py)). Training data is validated so bot lines cannot invent tags like `<step_count>` or nest intents.
|
| 43 |
+
|
| 44 |
+
Training data generation ([`tinydata/generate.py`](../tinydata/generate.py)) rejects samples that:
|
| 45 |
+
|
| 46 |
+
- Miss the leading `<INTENT:β¦>` on bot lines
|
| 47 |
+
- Use unknown intents or placeholders
|
| 48 |
+
- Include raw numeric metrics (2+ digit numbers) in bot text
|
| 49 |
+
- Break the strict `user:` / `bot:` line format
|
| 50 |
+
|
| 51 |
+
Your runtime should **never show slot tokens raw to the user** β always run `fillSlots()` with live values (see [`reply.ts`](../collab-run-1/web/src/reply.ts)).
|
| 52 |
+
|
| 53 |
+
---
|
| 54 |
+
|
| 55 |
+
## 2. Sampling controls (inference)
|
| 56 |
+
|
| 57 |
+
Web defaults in [`config.ts`](../collab-run-1/web/src/config.ts):
|
| 58 |
+
|
| 59 |
+
| Parameter | Web value | Python REPL (Export-0.1) | Effect |
|
| 60 |
+
|-----------|-----------|--------------------------|--------|
|
| 61 |
+
| `TEMPERATURE` | **0.5** | 0.8 | Lower = less random token choice, fewer weird word combinations |
|
| 62 |
+
| `TOP_K` | 40 | 40 | Sample only from the 40 most likely next tokens |
|
| 63 |
+
| `MAX_NEW_TOKENS` | **40** | 120 | Hard cap on reply length β stops rambling early |
|
| 64 |
+
| `BLOCK_SIZE` | 256 | 256 | Context window; older tokens are dropped from the forward pass |
|
| 65 |
+
|
| 66 |
+
Implementation in [`model.ts`](../collab-run-1/web/src/model.ts):
|
| 67 |
+
|
| 68 |
+
- **Top-k sampling** β sort vocab by probability, renormalize over top 40, then sample. Cuts off the long tail of nonsense tokens.
|
| 69 |
+
- **Temperature scaling** β divide logits by `max(temperature, 1e-8)` before softmax.
|
| 70 |
+
- **EOS early stop** β after step 2, if the sampled token id is `0`, stop generating. Prevents run-on after an end-of-sequence signal.
|
| 71 |
+
- **Context trimming** β only the last 256 token ids are passed to ONNX each step (`inputIds.slice(-seqLen)`), matching training context limits.
|
| 72 |
+
|
| 73 |
+
For a smartwatch, prefer **web-style settings** (temp ~0.5, max tokens ~40) unless you need longer replies for voice.
|
| 74 |
+
|
| 75 |
+
---
|
| 76 |
+
|
| 77 |
+
## 3. Prompt format and conversation history
|
| 78 |
+
|
| 79 |
+
The model expects a fixed transcript shape:
|
| 80 |
+
|
| 81 |
+
```
|
| 82 |
+
user: How many steps today?
|
| 83 |
+
bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> β keep going!
|
| 84 |
+
user: thanks
|
| 85 |
+
bot:
|
| 86 |
+
```
|
| 87 |
+
|
| 88 |
+
[`buildPrompt()`](../collab-run-1/web/src/reply.ts) builds this string. Important details:
|
| 89 |
+
|
| 90 |
+
- Every turn uses lowercase `user:` and `bot:` prefixes with a trailing space after the colon.
|
| 91 |
+
- The current turn ends with `bot:` and **no** reply yet β the model continues from there.
|
| 92 |
+
- **History stores `rawBot`** β the intent tag and unfilled slot template, not the display string with dummy sensor values. If you put filled text like `"You're at 1,000 steps"` back into history, the model sees numbers it was never trained on and quality drops.
|
| 93 |
+
|
| 94 |
+
Only **new** tokens are decoded (`ids.slice(promptLen)`), so the reply extraction starts clean after the prompt.
|
| 95 |
+
|
| 96 |
+
---
|
| 97 |
+
|
| 98 |
+
## 4. Post-processing (`reply.ts`)
|
| 99 |
+
|
| 100 |
+
Raw decode output can still contain tokenizer glitches or extra lines. The web app runs a pipeline before display.
|
| 101 |
+
|
| 102 |
+
### 4.1 `cleanReply(text)`
|
| 103 |
+
|
| 104 |
+
Fixes common BPE / UTF-8 artifacts:
|
| 105 |
+
|
| 106 |
+
- `Δ ` β space (GPT-style space marker)
|
| 107 |
+
- `Δ` β newline
|
| 108 |
+
- Mojibake sequences for em dash and curly quotes β `β`, `'`
|
| 109 |
+
- Collapse repeated spaces
|
| 110 |
+
- Normalize spaces **inside** angle brackets: `< INTENT : GET_STEPS >` β `<INTENT:GET_STEPS>`
|
| 111 |
+
- Trim stray space-before-apostrophe: ` '` β `'`
|
| 112 |
+
|
| 113 |
+
### 4.2 `extractIntentReply(text)`
|
| 114 |
+
|
| 115 |
+
Parses the structured reply:
|
| 116 |
+
|
| 117 |
+
1. Run `cleanReply`.
|
| 118 |
+
2. Find the first `<INTENT:β¦>` tag (case-insensitive, tolerant of internal spaces).
|
| 119 |
+
3. **Truncate at `\nuser:`** β if the model starts a fake next turn, discard everything after.
|
| 120 |
+
4. Take **only the first line** β one utterance per reply, matching Python `extract_bot_reply`.
|
| 121 |
+
5. Return `{ intent, template }` where `template` is the text after the intent tag.
|
| 122 |
+
|
| 123 |
+
If no intent tag is found, intent defaults to `NONE` and the first line is used as template.
|
| 124 |
+
|
| 125 |
+
### 4.3 `fillSlots(template, data)`
|
| 126 |
+
|
| 127 |
+
Replace `<SLOT_NAME>` with values from your sensor map. Unknown slots stay as-is (useful for debugging).
|
| 128 |
+
|
| 129 |
+
### 4.4 Python parity β `extract_bot_reply`
|
| 130 |
+
|
| 131 |
+
The Colab / Export-0.1 Python helper does the same job with prompt-relative slicing:
|
| 132 |
+
|
| 133 |
+
- Strip the prompt prefix from the full decode
|
| 134 |
+
- Stop at `\nuser:`, `\n\n`, or first newline
|
| 135 |
+
- Return a single-line bot utterance
|
| 136 |
+
|
| 137 |
+
The web README notes: *"Generation uses the same prompt format and cleanup as Colab `ask()`"* β with the extra `cleanReply` and intent parsing layers on top for browser display.
|
| 138 |
+
|
| 139 |
+
---
|
| 140 |
+
|
| 141 |
+
## 5. Infrastructure choices that affect quality
|
| 142 |
+
|
| 143 |
+
These are not sampling tricks, but bad infrastructure produces garbage logits or wrong token ids.
|
| 144 |
+
|
| 145 |
+
| Choice | Location | Why |
|
| 146 |
+
|--------|----------|-----|
|
| 147 |
+
| **Legacy ONNX export** (opset 17, `dynamo=False`) | [`export_onnx.py`](../collab-run-1/export_onnx.py) | Colabβs default exporter can produce invalid graphs; re-export fixes nonsense outputs in WASM |
|
| 148 |
+
| **Single merged ONNX file** | `smartwatch_lm_merged.onnx` | Browser fetch + ORT session needs one blob; external weights often fail in static hosting |
|
| 149 |
+
| **`@huggingface/tokenizers` (JS)** | [`tokenizer.ts`](../collab-run-1/web/src/tokenizer.ts) | Pure JS tokenizer matches training `tokenizer.json`; avoids ORT tokenizer conflicts |
|
| 150 |
+
| **`graphOptimizationLevel: "disabled"`** | [`model.ts`](../collab-run-1/web/src/model.ts) | Some WASM builds misbehave with aggressive graph opts on this model |
|
| 151 |
+
| **Decode only new tokens** | [`model.ts`](../collab-run-1/web/src/model.ts) | Prevents prompt text from leaking into the βreplyβ string |
|
| 152 |
+
|
| 153 |
+
---
|
| 154 |
+
|
| 155 |
+
## 6. Recommended runtime checklist
|
| 156 |
+
|
| 157 |
+
When porting off the web demo, implement this loop:
|
| 158 |
+
|
| 159 |
+
```
|
| 160 |
+
1. buildPrompt(history with raw bot lines, userMessage)
|
| 161 |
+
2. encode β generate (temp 0.5, top_k 40, max 40 tokens, EOS stop)
|
| 162 |
+
3. decode new tokens only
|
| 163 |
+
4. cleanReply β extractIntentReply
|
| 164 |
+
5. route intent to device handler (optional: refresh slot map from sensors)
|
| 165 |
+
6. fillSlots(template, slotMap) β show / speak to user
|
| 166 |
+
7. append (userMessage, rawBotLine) to history
|
| 167 |
+
```
|
| 168 |
+
|
| 169 |
+
**Do not:**
|
| 170 |
+
|
| 171 |
+
- Feed display strings with real numbers back into history
|
| 172 |
+
- Skip truncation at `\nuser:` or first newline
|
| 173 |
+
- Use high temperature on a 28M-param domain model
|
| 174 |
+
- Let the model invent metric values β always use slots
|
| 175 |
+
|
| 176 |
+
**Do:**
|
| 177 |
+
|
| 178 |
+
- Reset or trim history when context approaches 256 tokens
|
| 179 |
+
- Fall back to a safe message if intent is unknown or template is empty
|
| 180 |
+
- Keep generation caps tight for wrist UX (one short sentence)
|
| 181 |
+
|
| 182 |
+
---
|
| 183 |
+
|
| 184 |
+
## 7. What this does *not* fix
|
| 185 |
+
|
| 186 |
+
These techniques improve **format adherence and readability**, not general reasoning:
|
| 187 |
+
|
| 188 |
+
- Off-domain questions may still get weak or wrong intents
|
| 189 |
+
- Typos in user input are not corrected
|
| 190 |
+
- No content moderation or safety filter is applied
|
| 191 |
+
- Synthetic training data may not match real user phrasing
|
| 192 |
+
|
| 193 |
+
For production wearables, combine this model with **intent validation** (reject unknown intents), **handler guards** (donβt start a workout if one is already active), and optional **cloud fallback** for out-of-scope queries.
|
| 194 |
+
|
| 195 |
+
---
|
| 196 |
+
|
| 197 |
+
## Related docs
|
| 198 |
+
|
| 199 |
+
- [Intent reference](./readme.md) β all 35 intents and slot names
|
| 200 |
+
- [Smartwatch integration guide](./smartwatch-integration.md) β end-to-end device wiring
|
| 201 |
+
- [Export-0.1 README](../README.md) β model card and ONNX I/O
|