Greedy decoding emits token 128486 (orphan UTF-8 byte + Lao) inside Greek text → invalid UTF-8, llama-server HTTP 500

#25
by apollonlabsai - opened

Hi, and thanks for Spark-X2.5-4B. We are fine-tuning it for Greek, code and tool calling, and it has been a great base.

While running our release checks we found a reproducible generation defect in the base model, and we wanted to report it with a minimal repro.

Prompt (Greek, translation task):

system: Είσαι βοηθός. Απάντησε σύντομα.
user:   Μετάφρασε στα Ελληνικά: "The dog followed him to the end of the road."

Greedy decoding (do_sample=False / temperature 0), enable_thinking=False.

Output (transformers, bf16 and fp32, all files from this repo at commit 0bcb356):

Το σκύρο τον�ິດຕούσε μέχρι το τέλος της δρόμου.

Tokens:

['Το', 'ĠÏĥκ', 'Ïį', 'Ïģο', 'ĠÏĦον', 'ķິàºĶàºķ', 'οÏį', 'Ïĥε', ...]

In the middle of a Greek word the model picks token id 128486 (ķິàºĶàºķ). That token begins with an orphan UTF-8 continuation byte followed by Lao characters, so any sequence containing it decodes to invalid UTF-8 (�).

Minimal transformers repro

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
p = "XHToken/Spark-X2.5-4B"
tok = AutoTokenizer.from_pretrained(p, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(p, torch_dtype=torch.bfloat16, trust_remote_code=True).eval()
msgs = [{"role": "system", "content": "Είσαι βοηθός. Απάντησε σύντομα."},
        {"role": "user", "content": 'Μετάφρασε στα Ελληνικά: "The dog followed him to the end of the road."'}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, enable_thinking=False, return_tensors="pt")
out = model.generate(ids, max_new_tokens=40, do_sample=False)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))

(Tested with transformers 4.57.1, torch 2.14.0 CPU.)

Consequence in llama.cpp: the same output appears with an F16 GGUF converted from this repo (convert_hf_to_gguf.py, build 10909 / commit 329b6160f). llama-server --jinja cannot parse it and returns HTTP 500:

common_chat_peg_parse: unparsed peg-native output: Το σκύρο τον�ິດຕούσε μέχρι το τέλος της δρόμου.

Lower-bit quantizations sometimes pick a different token at that step and do not crash. Rounding hides the defect there; it does not fix it.

A fine-tune of ours on top of the base shows the same token at the same position, so the defect appears to come from the base weights.

Happy to share more prompts or logs if useful.

Apollon Labs

Follow-up: mitigation we applied

Following up on the orphan-byte issue above: this is how we mitigated it on our fine-tune (Heliactis 1, 4B, based on Spark-X2.5-4B).

What worked (weights): on-policy unlikelihood + SFT on a LoRA.

  • We generated the model's own answers to 428 prompts and penalised the probability mass on the 502 orphan-byte token IDs at every generated position. We used the sum of that mass, not the mean, with unlikelihood weight 10.
  • Unlikelihood on the static training set gave almost no signal (mass ~2e-5). The problem only shows up on-policy.
  • Result on the failing prompt ("The dog followed him to the end of the road.", greedy): the max ban mass dropped from 0.127 to 0.0017. There were no orphan tokens in the output and no HTTP 500. On 20 held-out prompts the max mass is 0.0002.

The trade-off we measured: pushing the mass lower (0.0010, then 0.0001) started damaging Greek vocabulary. The model began producing non-words, and our vocabulary gate fell from 0.955 to 0.932. So we stopped at the version that passes every gate.

What we recommend on top (inference): a logit_bias of −∞ on the 502 IDs. The list is here: https://huggingface.co/ApollonLabs/Heliactis-1-4B-GGUF/blob/main/ban-ids.json. A system prompt does not fix it; the failing prompt already runs with one.

A head-only edit without training did not work for us.

Model + card: https://huggingface.co/ApollonLabs/Heliactis-1-4B-GGUF · full reproduction: https://huggingface.co/ApollonLabs/Heliactis-1-4B-GGUF/blob/main/DEFECT.md

Sign up or log in to comment