Make repo self-contained: rewrite docs, single-model benchmark, remove external references
dbb5d78 verified |
Download docs/avoiding-gibberish.md from prathamkode/smartwatch-lm-0.2: direct link, hf CLI and curl.
- Browser
- Download file 5.79 kB
-
https://huggingface.co/prathamkode/smartwatch-lm-0.2/resolve/main/docs/avoiding-gibberish.md
- Command line
-
hf download hf://prathamkode/smartwatch-lm-0.2/docs/avoiding-gibberish.md
-
curl -L -o avoiding-gibberish.md https://huggingface.co/prathamkode/smartwatch-lm-0.2/resolve/main/docs/avoiding-gibberish.md
5.79 kB
| # Avoiding Gibberish — Output Cleanup Guide | |
| Smartwatch LM v0.2 is a **~15.4M-parameter** domain model. It outputs structured replies with intent tags and slot placeholders. Raw tokenizer decode can still contain **BPE artifacts**, **encoding glitches**, or **run-on text**. This guide lists what to strip, how to truncate, and includes copy-paste scripts in this repo. | |
| **Scripts in this repo:** | |
| | File | Purpose | | |
| |------|---------| | |
| | [`reply_utils.py`](../reply_utils.py) | `clean_reply`, `extract_bot_reply`, `extract_intent_reply`, `fill_slots` | | |
| | [`onnx_sample.py`](../onnx_sample.py) | Full ONNX generate + cleanup pipeline | | |
| | [`chat.py`](../chat.py) | PyTorch REPL using the same helpers | | |
| --- | |
| ## Why gibberish appears | |
| | Cause | Example | Fix | | |
| |-------|---------|-----| | |
| | BPE space marker left in decode | `ĠYou'reĠat` | Replace `Ġ` → space | | |
| | BPE newline marker | `line oneĊline two` | Replace `Ċ` → newline | | |
| | UTF-8 mojibake | `âĢĶ` instead of `—` | Replace known bad sequences | | |
| | Model keeps generating | Fake next turn `\nuser: …` | Truncate at `\nuser:` | | |
| | Model rambling | Multiple paragraphs | Take first line only | | |
| | Broken intent tags | `< INTENT : GET_STEPS >` | Collapse spaces inside `<…>` | | |
| | Hallucinated metrics | `8432 steps` | Use `<STEPS_TODAY>` slots instead (training constraint) | | |
| --- | |
| ## Special characters to remove or replace | |
| Apply these **in order** after every `tokenizer.decode()`: | |
| ### 1. BPE artifacts (always) | |
| | Character | Unicode | Replace with | | |
| |-----------|---------|--------------| | |
| | `Ġ` | U+0120 | space (` `) | | |
| | `Ċ` | U+010A | newline (`\n`) | | |
| ### 2. Mojibake sequences (when present) | |
| | Bad sequence | Replace with | | |
| |--------------|--------------| | |
| | `âĢĶ` | `—` (em dash) | | |
| | `âĢĻ` | `'` | | |
| | `âĢĺ` | `'` | | |
| | `’` | `'` | | |
| | `–` | `—` | | |
| ### 3. Whitespace normalization | |
| | Pattern | Action | | |
| |---------|--------| | |
| | Two or more spaces | Collapse to one space | | |
| | Space before apostrophe (` '`) | Remove the space → `'` | | |
| | Spaces inside angle brackets | Remove: `< STEP_GOAL >` → `<STEP_GOAL>` | | |
| ### 4. Do **not** remove | |
| Keep these — they are part of the protocol: | |
| - `<INTENT:NAME>` tags | |
| - `<SLOT_NAME>` placeholders (filled by your app before display) | |
| - Normal punctuation: `. , ! ? ' —` | |
| --- | |
| ## Truncation rules (stop run-on output) | |
| After cleaning characters, cut the reply to **one bot utterance**: | |
| 1. If `\nuser:` appears → discard everything from that point onward (model started a fake user turn). | |
| 2. If `\n\n` appears → keep only the text before the blank line. | |
| 3. Keep only the **first line** (everything before the first `\n`). | |
| These rules are implemented in `extract_bot_reply()` in [`reply_utils.py`](../reply_utils.py). | |
| --- | |
| ## Generation settings that reduce gibberish | |
| Use conservative sampling — this model is small (~15M params): | |
| | Parameter | Recommended | Effect | | |
| |-----------|-------------|--------| | |
| | `temperature` | **0.5** (max 0.8) | Less random word choice | | |
| | `top_k` | **40** | Ignore unlikely tail tokens | | |
| | `max_new_tokens` | **40** | Stop before rambling | | |
| | `block_size` | **256** | Match model context; trim older history | | |
| | EOS token id | **0** | Stop generating after EOS (skip first 2 steps) | | |
| Store **unfilled** bot lines in history (`<STEPS_TODAY>` not `4,231`). Filled numbers in history confuse the model. | |
| --- | |
| ## Prompt format | |
| ``` | |
| user: How many steps today? | |
| bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> — keep going! | |
| user: thanks | |
| bot: | |
| ``` | |
| Build with `build_prompt()` from [`reply_utils.py`](../reply_utils.py). The model continues after the final `bot:`. | |
| --- | |
| ## Sample: clean text only (no model) | |
| ```bash | |
| python reply_utils.py | |
| ``` | |
| ```python | |
| from reply_utils import clean_reply, extract_intent_reply, fill_slots | |
| messy = "Ġ<INTENT:GET_STEPS>ĠYou'reĠatĠ<STEPS_TODAY>ĠâĢĶĠnice!\nuser: more" | |
| parsed = extract_intent_reply(messy) | |
| display = fill_slots(parsed.template, {"STEPS_TODAY": "4,231"}) | |
| print(parsed.intent) # GET_STEPS | |
| print(display) # You're at 4,231 — nice! | |
| ``` | |
| --- | |
| ## Sample: PyTorch chat with cleanup | |
| ```bash | |
| pip install torch tokenizers | |
| python chat.py | |
| ``` | |
| ```python | |
| from chat import ChatSession | |
| bot = ChatSession(temperature=0.5, max_new_tokens=40, top_k=40) | |
| print(bot.say("How many steps today?")) | |
| ``` | |
| --- | |
| ## Sample: ONNX with cleanup | |
| ```bash | |
| pip install numpy onnxruntime tokenizers | |
| python onnx_sample.py "How many steps today?" | |
| ``` | |
| Pipeline inside `onnx_sample.py`: | |
| ``` | |
| build_prompt → encode → ONNX loop (top-k, temp 0.5) → decode new tokens only | |
| → extract_bot_reply → extract_intent_reply → fill_slots → print | |
| ``` | |
| --- | |
| ## Full runtime checklist | |
| ``` | |
| 1. build_prompt(history with raw unfilled bot lines, user_message) | |
| 2. encode → generate (temp 0.5, top_k 40, max 40, EOS stop) | |
| 3. decode ONLY new token ids (not the full prompt) | |
| 4. clean_reply → extract_bot_reply → extract_intent_reply | |
| 5. fill_slots(template, sensor_map) → show / speak | |
| 6. append (user_message, raw_bot_line) to history | |
| ``` | |
| **Do not:** | |
| - Show raw `Ġ` or mojibake to the user | |
| - Put real sensor numbers back into conversation history | |
| - Skip truncation at `\nuser:` or first newline | |
| - Run temperature above 0.8 on this model | |
| **Do:** | |
| - Run `clean_reply` on every decode path | |
| - Validate intent names against your handler allowlist | |
| - Trim history when approaching 256 tokens | |
| --- | |
| ## Related docs | |
| - [Intent reference](./intent-reference.md) — all 35 intents and slots | |
| - [Smartwatch integration](./smartwatch-integration.md) — device wiring | |