# Avoiding Gibberish — Output Cleanup Guide Smartwatch LM v0.2 is a **~15.4M-parameter** domain model. It outputs structured replies with intent tags and slot placeholders. Raw tokenizer decode can still contain **BPE artifacts**, **encoding glitches**, or **run-on text**. This guide lists what to strip, how to truncate, and includes copy-paste scripts in this repo. **Scripts in this repo:** | File | Purpose | |------|---------| | [`reply_utils.py`](../reply_utils.py) | `clean_reply`, `extract_bot_reply`, `extract_intent_reply`, `fill_slots` | | [`onnx_sample.py`](../onnx_sample.py) | Full ONNX generate + cleanup pipeline | | [`chat.py`](../chat.py) | PyTorch REPL using the same helpers | --- ## Why gibberish appears | Cause | Example | Fix | |-------|---------|-----| | BPE space marker left in decode | `ĠYou'reĠat` | Replace `Ġ` → space | | BPE newline marker | `line oneĊline two` | Replace `Ċ` → newline | | UTF-8 mojibake | `âĢĶ` instead of `—` | Replace known bad sequences | | Model keeps generating | Fake next turn `\nuser: …` | Truncate at `\nuser:` | | Model rambling | Multiple paragraphs | Take first line only | | Broken intent tags | `< INTENT : GET_STEPS >` | Collapse spaces inside `<…>` | | Hallucinated metrics | `8432 steps` | Use `` slots instead (training constraint) | --- ## Special characters to remove or replace Apply these **in order** after every `tokenizer.decode()`: ### 1. BPE artifacts (always) | Character | Unicode | Replace with | |-----------|---------|--------------| | `Ġ` | U+0120 | space (` `) | | `Ċ` | U+010A | newline (`\n`) | ### 2. Mojibake sequences (when present) | Bad sequence | Replace with | |--------------|--------------| | `âĢĶ` | `—` (em dash) | | `âĢĻ` | `'` | | `âĢĺ` | `'` | | `’` | `'` | | `–` | `—` | ### 3. Whitespace normalization | Pattern | Action | |---------|--------| | Two or more spaces | Collapse to one space | | Space before apostrophe (` '`) | Remove the space → `'` | | Spaces inside angle brackets | Remove: `< STEP_GOAL >` → `` | ### 4. Do **not** remove Keep these — they are part of the protocol: - `` tags - `` placeholders (filled by your app before display) - Normal punctuation: `. , ! ? ' —` --- ## Truncation rules (stop run-on output) After cleaning characters, cut the reply to **one bot utterance**: 1. If `\nuser:` appears → discard everything from that point onward (model started a fake user turn). 2. If `\n\n` appears → keep only the text before the blank line. 3. Keep only the **first line** (everything before the first `\n`). These rules are implemented in `extract_bot_reply()` in [`reply_utils.py`](../reply_utils.py). --- ## Generation settings that reduce gibberish Use conservative sampling — this model is small (~15M params): | Parameter | Recommended | Effect | |-----------|-------------|--------| | `temperature` | **0.5** (max 0.8) | Less random word choice | | `top_k` | **40** | Ignore unlikely tail tokens | | `max_new_tokens` | **40** | Stop before rambling | | `block_size` | **256** | Match model context; trim older history | | EOS token id | **0** | Stop generating after EOS (skip first 2 steps) | Store **unfilled** bot lines in history (`` not `4,231`). Filled numbers in history confuse the model. --- ## Prompt format ``` user: How many steps today? bot: You're at of — keep going! user: thanks bot: ``` Build with `build_prompt()` from [`reply_utils.py`](../reply_utils.py). The model continues after the final `bot:`. --- ## Sample: clean text only (no model) ```bash python reply_utils.py ``` ```python from reply_utils import clean_reply, extract_intent_reply, fill_slots messy = "ĠĠYou'reĠatĠĠâĢĶĠnice!\nuser: more" parsed = extract_intent_reply(messy) display = fill_slots(parsed.template, {"STEPS_TODAY": "4,231"}) print(parsed.intent) # GET_STEPS print(display) # You're at 4,231 — nice! ``` --- ## Sample: PyTorch chat with cleanup ```bash pip install torch tokenizers python chat.py ``` ```python from chat import ChatSession bot = ChatSession(temperature=0.5, max_new_tokens=40, top_k=40) print(bot.say("How many steps today?")) ``` --- ## Sample: ONNX with cleanup ```bash pip install numpy onnxruntime tokenizers python onnx_sample.py "How many steps today?" ``` Pipeline inside `onnx_sample.py`: ``` build_prompt → encode → ONNX loop (top-k, temp 0.5) → decode new tokens only → extract_bot_reply → extract_intent_reply → fill_slots → print ``` --- ## Full runtime checklist ``` 1. build_prompt(history with raw unfilled bot lines, user_message) 2. encode → generate (temp 0.5, top_k 40, max 40, EOS stop) 3. decode ONLY new token ids (not the full prompt) 4. clean_reply → extract_bot_reply → extract_intent_reply 5. fill_slots(template, sensor_map) → show / speak 6. append (user_message, raw_bot_line) to history ``` **Do not:** - Show raw `Ġ` or mojibake to the user - Put real sensor numbers back into conversation history - Skip truncation at `\nuser:` or first newline - Run temperature above 0.8 on this model **Do:** - Run `clean_reply` on every decode path - Validate intent names against your handler allowlist - Trim history when approaching 256 tokens --- ## Related docs - [Intent reference](./intent-reference.md) — all 35 intents and slots - [Smartwatch integration](./smartwatch-integration.md) — device wiring