smartwatch-lm-0.2 / docs /avoiding-gibberish.md
prathamkode's picture
Make repo self-contained: rewrite docs, single-model benchmark, remove external references
dbb5d78 verified
|
Raw History Blame Contribute Delete
5.79 kB

Avoiding Gibberish — Output Cleanup Guide

Smartwatch LM v0.2 is a ~15.4M-parameter domain model. It outputs structured replies with intent tags and slot placeholders. Raw tokenizer decode can still contain BPE artifacts, encoding glitches, or run-on text. This guide lists what to strip, how to truncate, and includes copy-paste scripts in this repo.

Scripts in this repo:

File Purpose
reply_utils.py clean_reply, extract_bot_reply, extract_intent_reply, fill_slots
onnx_sample.py Full ONNX generate + cleanup pipeline
chat.py PyTorch REPL using the same helpers

Why gibberish appears

Cause Example Fix
BPE space marker left in decode ĠYou'reĠat Replace Ġ → space
BPE newline marker line oneĊline two Replace Ċ → newline
UTF-8 mojibake âĢĶ instead of — Replace known bad sequences
Model keeps generating Fake next turn \nuser: … Truncate at \nuser:
Model rambling Multiple paragraphs Take first line only
Broken intent tags < INTENT : GET_STEPS > Collapse spaces inside <…>
Hallucinated metrics 8432 steps Use <STEPS_TODAY> slots instead (training constraint)

Special characters to remove or replace

Apply these in order after every tokenizer.decode():

1. BPE artifacts (always)

Character Unicode Replace with
Ġ U+0120 space ( )
Ċ U+010A newline (\n)

2. Mojibake sequences (when present)

Bad sequence Replace with
âĢĶ — (em dash)
âĢĻ '
âĢĺ '
’ '
– —

3. Whitespace normalization

Pattern Action
Two or more spaces Collapse to one space
Space before apostrophe ( ') Remove the space → '
Spaces inside angle brackets Remove: < STEP_GOAL > → <STEP_GOAL>

4. Do not remove

Keep these — they are part of the protocol:

  • <INTENT:NAME> tags
  • <SLOT_NAME> placeholders (filled by your app before display)
  • Normal punctuation: . , ! ? ' —

Truncation rules (stop run-on output)

After cleaning characters, cut the reply to one bot utterance:

  1. If \nuser: appears → discard everything from that point onward (model started a fake user turn).
  2. If \n\n appears → keep only the text before the blank line.
  3. Keep only the first line (everything before the first \n).

These rules are implemented in extract_bot_reply() in reply_utils.py.


Generation settings that reduce gibberish

Use conservative sampling — this model is small (~15M params):

Parameter Recommended Effect
temperature 0.5 (max 0.8) Less random word choice
top_k 40 Ignore unlikely tail tokens
max_new_tokens 40 Stop before rambling
block_size 256 Match model context; trim older history
EOS token id 0 Stop generating after EOS (skip first 2 steps)

Store unfilled bot lines in history (<STEPS_TODAY> not 4,231). Filled numbers in history confuse the model.


Prompt format

user: How many steps today?
bot: <INTENT:GET_STEPS> You're at <STEPS_TODAY> of <STEP_GOAL> — keep going!
user: thanks
bot:

Build with build_prompt() from reply_utils.py. The model continues after the final bot:.


Sample: clean text only (no model)

python reply_utils.py
from reply_utils import clean_reply, extract_intent_reply, fill_slots

messy = "Ġ<INTENT:GET_STEPS>ĠYou'reĠatĠ<STEPS_TODAY>ĠâĢĶĠnice!\nuser: more"

parsed = extract_intent_reply(messy)
display = fill_slots(parsed.template, {"STEPS_TODAY": "4,231"})

print(parsed.intent)   # GET_STEPS
print(display)         # You're at 4,231 — nice!

Sample: PyTorch chat with cleanup

pip install torch tokenizers
python chat.py
from chat import ChatSession

bot = ChatSession(temperature=0.5, max_new_tokens=40, top_k=40)
print(bot.say("How many steps today?"))

Sample: ONNX with cleanup

pip install numpy onnxruntime tokenizers
python onnx_sample.py "How many steps today?"

Pipeline inside onnx_sample.py:

build_prompt → encode → ONNX loop (top-k, temp 0.5) → decode new tokens only
→ extract_bot_reply → extract_intent_reply → fill_slots → print

Full runtime checklist

1. build_prompt(history with raw unfilled bot lines, user_message)
2. encode → generate (temp 0.5, top_k 40, max 40, EOS stop)
3. decode ONLY new token ids (not the full prompt)
4. clean_reply → extract_bot_reply → extract_intent_reply
5. fill_slots(template, sensor_map) → show / speak
6. append (user_message, raw_bot_line) to history

Do not:

  • Show raw Ġ or mojibake to the user
  • Put real sensor numbers back into conversation history
  • Skip truncation at \nuser: or first newline
  • Run temperature above 0.8 on this model

Do:

  • Run clean_reply on every decode path
  • Validate intent names against your handler allowlist
  • Trim history when approaching 256 tokens

Related docs