smartwatch-lm-0.2 / docs /avoiding-gibberish.md
prathamkode's picture
Make repo self-contained: rewrite docs, single-model benchmark, remove external references
dbb5d78 verified
|
Raw History Blame Contribute Delete
5.79 kB
# Avoiding Gibberish — Output Cleanup Guide
Smartwatch LM v0.2 is a **~15.4M-parameter** domain model. It outputs structured replies with intent tags and slot placeholders. Raw tokenizer decode can still contain **BPE artifacts**, **encoding glitches**, or **run-on text**. This guide lists what to strip, how to truncate, and includes copy-paste scripts in this repo.
**Scripts in this repo:**
| File | Purpose |
|------|---------|
| [`reply_utils.py`](../reply_utils.py) | `clean_reply`, `extract_bot_reply`, `extract_intent_reply`, `fill_slots` |
| [`onnx_sample.py`](../onnx_sample.py) | Full ONNX generate + cleanup pipeline |
| [`chat.py`](../chat.py) | PyTorch REPL using the same helpers |
---
## Why gibberish appears
| Cause | Example | Fix |
|-------|---------|-----|
| BPE space marker left in decode | `ĠYou'reĠat` | Replace `Ġ` → space |
| BPE newline marker | `line oneĊline two` | Replace `Ċ` → newline |
| UTF-8 mojibake | `âĢĶ` instead of `—` | Replace known bad sequences |
| Model keeps generating | Fake next turn `\nuser: …` | Truncate at `\nuser:` |
| Model rambling | Multiple paragraphs | Take first line only |
| Broken intent tags | `< INTENT : GET_STEPS >` | Collapse spaces inside `<…>` |
| Hallucinated metrics | `8432 steps` | Use `<STEPS_TODAY>` slots instead (training constraint) |
---
## Special characters to remove or replace
Apply these **in order** after every `tokenizer.decode()`:
### 1. BPE artifacts (always)
| Character | Unicode | Replace with |
|-----------|---------|--------------|
| `Ġ` | U+0120 | space (` `) |
| `Ċ` | U+010A | newline (`\n`) |
### 2. Mojibake sequences (when present)
| Bad sequence | Replace with |
|--------------|--------------|
| `âĢĶ` | `—` (em dash) |
| `âĢĻ` | `'` |
| `âĢĺ` | `'` |
| `’` | `'` |
| `–` | `—` |
### 3. Whitespace normalization
| Pattern | Action |
|---------|--------|
| Two or more spaces | Collapse to one space |
| Space before apostrophe (` '`) | Remove the space → `'` |
| Spaces inside angle brackets | Remove: `< STEP_GOAL >` → `<STEP_GOAL>` |
### 4. Do **not** remove
Keep these — they are part of the protocol:
- `<INTENT:NAME>` tags
- `<SLOT_NAME>` placeholders (filled by your app before display)
- Normal punctuation: `. , ! ? ' —`
---
## Truncation rules (stop run-on output)
After cleaning characters, cut the reply to **one bot utterance**:
1. If `\nuser:` appears → discard everything from that point onward (model started a fake user turn).
2. If `\n\n` appears → keep only the text before the blank line.
3. Keep only the **first line** (everything before the first `\n`).
These rules are implemented in `extract_bot_reply()` in [`reply_utils.py`](../reply_utils.py).
---
## Generation settings that reduce gibberish
Use conservative sampling — this model is small (~15M params):
| Parameter | Recommended | Effect |
|-----------|-------------|--------|
| `temperature` | **0.5** (max 0.8) | Less random word choice |
| `top_k` | **40** | Ignore unlikely tail tokens |
| `max_new_tokens` | **40** | Stop before rambling |
| `block_size` | **256** | Match model context; trim older history |
| EOS token id | **0** | Stop generating after EOS (skip first 2 steps) |
Store **unfilled** bot lines in history (`<STEPS_TODAY>` not `4,231`). Filled numbers in history confuse the model.
---
## Prompt format
```
user: How many steps today?
bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> — keep going!
user: thanks
bot:
```
Build with `build_prompt()` from [`reply_utils.py`](../reply_utils.py). The model continues after the final `bot:`.
---
## Sample: clean text only (no model)
```bash
python reply_utils.py
```
```python
from reply_utils import clean_reply, extract_intent_reply, fill_slots
messy = "Ġ<INTENT:GET_STEPS>ĠYou'reĠatĠ<STEPS_TODAY>ĠâĢĶĠnice!\nuser: more"
parsed = extract_intent_reply(messy)
display = fill_slots(parsed.template, {"STEPS_TODAY": "4,231"})
print(parsed.intent) # GET_STEPS
print(display) # You're at 4,231 — nice!
```
---
## Sample: PyTorch chat with cleanup
```bash
pip install torch tokenizers
python chat.py
```
```python
from chat import ChatSession
bot = ChatSession(temperature=0.5, max_new_tokens=40, top_k=40)
print(bot.say("How many steps today?"))
```
---
## Sample: ONNX with cleanup
```bash
pip install numpy onnxruntime tokenizers
python onnx_sample.py "How many steps today?"
```
Pipeline inside `onnx_sample.py`:
```
build_prompt → encode → ONNX loop (top-k, temp 0.5) → decode new tokens only
→ extract_bot_reply → extract_intent_reply → fill_slots → print
```
---
## Full runtime checklist
```
1. build_prompt(history with raw unfilled bot lines, user_message)
2. encode → generate (temp 0.5, top_k 40, max 40, EOS stop)
3. decode ONLY new token ids (not the full prompt)
4. clean_reply → extract_bot_reply → extract_intent_reply
5. fill_slots(template, sensor_map) → show / speak
6. append (user_message, raw_bot_line) to history
```
**Do not:**
- Show raw `Ġ` or mojibake to the user
- Put real sensor numbers back into conversation history
- Skip truncation at `\nuser:` or first newline
- Run temperature above 0.8 on this model
**Do:**
- Run `clean_reply` on every decode path
- Validate intent names against your handler allowlist
- Trim history when approaching 256 tokens
---
## Related docs
- [Intent reference](./intent-reference.md) — all 35 intents and slots
- [Smartwatch integration](./smartwatch-integration.md) — device wiring