Download docs/avoiding-gibberish.md from prathamkode/smartwatch-lm-0.2: direct link, hf CLI and curl.
- Browser
- Download file 5.79 kB
-
https://huggingface.co/prathamkode/smartwatch-lm-0.2/resolve/main/docs/avoiding-gibberish.md
- Command line
-
hf download hf://prathamkode/smartwatch-lm-0.2/docs/avoiding-gibberish.md
-
curl -L -o avoiding-gibberish.md https://huggingface.co/prathamkode/smartwatch-lm-0.2/resolve/main/docs/avoiding-gibberish.md
Avoiding Gibberish — Output Cleanup Guide
Smartwatch LM v0.2 is a ~15.4M-parameter domain model. It outputs structured replies with intent tags and slot placeholders. Raw tokenizer decode can still contain BPE artifacts, encoding glitches, or run-on text. This guide lists what to strip, how to truncate, and includes copy-paste scripts in this repo.
Scripts in this repo:
| File | Purpose |
|---|---|
reply_utils.py |
clean_reply, extract_bot_reply, extract_intent_reply, fill_slots |
onnx_sample.py |
Full ONNX generate + cleanup pipeline |
chat.py |
PyTorch REPL using the same helpers |
Why gibberish appears
| Cause | Example | Fix |
|---|---|---|
| BPE space marker left in decode | ĠYou'reĠat |
Replace Ġ → space |
| BPE newline marker | line oneĊline two |
Replace Ċ → newline |
| UTF-8 mojibake | âĢĶ instead of — |
Replace known bad sequences |
| Model keeps generating | Fake next turn \nuser: … |
Truncate at \nuser: |
| Model rambling | Multiple paragraphs | Take first line only |
| Broken intent tags | < INTENT : GET_STEPS > |
Collapse spaces inside <…> |
| Hallucinated metrics | 8432 steps |
Use <STEPS_TODAY> slots instead (training constraint) |
Special characters to remove or replace
Apply these in order after every tokenizer.decode():
1. BPE artifacts (always)
| Character | Unicode | Replace with |
|---|---|---|
Ġ |
U+0120 | space ( ) |
Ċ |
U+010A | newline (\n) |
2. Mojibake sequences (when present)
| Bad sequence | Replace with |
|---|---|
âĢĶ |
— (em dash) |
âĢĻ |
' |
âĢĺ |
' |
’ |
' |
– |
— |
3. Whitespace normalization
| Pattern | Action |
|---|---|
| Two or more spaces | Collapse to one space |
Space before apostrophe ( ') |
Remove the space → ' |
| Spaces inside angle brackets | Remove: < STEP_GOAL > → <STEP_GOAL> |
4. Do not remove
Keep these — they are part of the protocol:
<INTENT:NAME>tags<SLOT_NAME>placeholders (filled by your app before display)- Normal punctuation:
. , ! ? ' —
Truncation rules (stop run-on output)
After cleaning characters, cut the reply to one bot utterance:
- If
\nuser:appears → discard everything from that point onward (model started a fake user turn). - If
\n\nappears → keep only the text before the blank line. - Keep only the first line (everything before the first
\n).
These rules are implemented in extract_bot_reply() in reply_utils.py.
Generation settings that reduce gibberish
Use conservative sampling — this model is small (~15M params):
| Parameter | Recommended | Effect |
|---|---|---|
temperature |
0.5 (max 0.8) | Less random word choice |
top_k |
40 | Ignore unlikely tail tokens |
max_new_tokens |
40 | Stop before rambling |
block_size |
256 | Match model context; trim older history |
| EOS token id | 0 | Stop generating after EOS (skip first 2 steps) |
Store unfilled bot lines in history (<STEPS_TODAY> not 4,231). Filled numbers in history confuse the model.
Prompt format
user: How many steps today?
bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> — keep going!
user: thanks
bot:
Build with build_prompt() from reply_utils.py. The model continues after the final bot:.
Sample: clean text only (no model)
python reply_utils.py
from reply_utils import clean_reply, extract_intent_reply, fill_slots
messy = "Ġ<INTENT:GET_STEPS>ĠYou'reĠatĠ<STEPS_TODAY>ĠâĢĶĠnice!\nuser: more"
parsed = extract_intent_reply(messy)
display = fill_slots(parsed.template, {"STEPS_TODAY": "4,231"})
print(parsed.intent) # GET_STEPS
print(display) # You're at 4,231 — nice!
Sample: PyTorch chat with cleanup
pip install torch tokenizers
python chat.py
from chat import ChatSession
bot = ChatSession(temperature=0.5, max_new_tokens=40, top_k=40)
print(bot.say("How many steps today?"))
Sample: ONNX with cleanup
pip install numpy onnxruntime tokenizers
python onnx_sample.py "How many steps today?"
Pipeline inside onnx_sample.py:
build_prompt → encode → ONNX loop (top-k, temp 0.5) → decode new tokens only
→ extract_bot_reply → extract_intent_reply → fill_slots → print
Full runtime checklist
1. build_prompt(history with raw unfilled bot lines, user_message)
2. encode → generate (temp 0.5, top_k 40, max 40, EOS stop)
3. decode ONLY new token ids (not the full prompt)
4. clean_reply → extract_bot_reply → extract_intent_reply
5. fill_slots(template, sensor_map) → show / speak
6. append (user_message, raw_bot_line) to history
Do not:
- Show raw
Ġor mojibake to the user - Put real sensor numbers back into conversation history
- Skip truncation at
\nuser:or first newline - Run temperature above 0.8 on this model
Do:
- Run
clean_replyon every decode path - Validate intent names against your handler allowlist
- Trim history when approaching 256 tokens
Related docs
- Intent reference — all 35 intents and slots
- Smartwatch integration — device wiring