# Heliactis 1 — the orphan-byte defect in detail This file backs the short summary in the [model card](README.md). ## What the defect is `Spark-X2.5-4B` has **502 vocabulary entries that begin with an orphan UTF-8 continuation byte**. When the model emits one, the output contains invalid UTF-8 (`�`) and `llama-server` answers HTTP 500. We reported it upstream: [discussion #25](https://huggingface.co/XHToken/Spark-X2.5-4B/discussions/25). ## Reproduction Reproduced on the unmodified base (transformers, bf16 and fp32), greedy decoding: ``` Μετάφρασε στα Ελληνικά: "The dog followed him to the end of the road." base: Το χάρι του τον�ິດຕ → HTTP 500 Heliactis 1: Το χάριο τον ακολούθησε… → no broken token; "dog" is still mistranslated ``` ## What we changed Heliactis 1 was trained with **on-policy unlikelihood** against those 502 tokens. Probability mass on them at the worst position we found: | | worst position | 20 held-out prompts (max) | |---|---|---| | before (our unreleased intermediate fine-tune) | **12.7%** | 0.25% | | Heliactis 1 | **0.17%** | 0.02% | **It is rarer, not impossible.** Pushing the mass further down also damaged Greek vocabulary. We measured three training depths and kept the one that holds vocabulary inside the noise of our unreleased intermediate fine-tune (the starting point of this training): | training depth | worst-position mass | Greek vocabulary (gate ≥ 0.949) | |---|---|---| | **shipped** | 0.0017 | **0.9553** ✓ | | deeper | 0.0010 | 0.9480 ✗ | | deepest | 0.0001 | 0.9325 ✗ | ## Blocking the ids at inference The 502 tokens occur **0 times** in the correct tokenization of our 374,541-token Greek set, so blocking them outright costs nothing in Greek. **Block them at inference** to close this route completely: ```bash # ban-ids.json ships in this repo: the 502 token ids llama-server -m Heliactis-1-4B-Q4_K_M.gguf --jinja \ $(python -c "import json;print(' '.join(f'-l {i}-inf' for i in json.load(open('ban-ids.json'))))") ``` Via the API, per request: `"logit_bias": [[, false], ...]` for the same ids. A system prompt does **not** fix this: the failing test above already runs with one. ## A second route: Lao Blocking the 502 ids does not close every route to invalid UTF-8. The model can emit a token that **ends** mid-character and then continue with something that does not finish it. Example on Lao (greedy, official `llama-server`): token 21417 ends in bytes `e0 ba`, and the next token is ` Mek` instead of the missing byte. `llama-server` then answers HTTP 500. 1,451 vocabulary entries end with an unfinished character. In our test (one prompt per language) the Lao paragraph failed on Heliactis 1 and on our unreleased intermediate fine-tune, but not on the base model. Thai passed on all three. **Do not use this model for Lao.**