Instructions to use XHToken/Spark-X2.5-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use XHToken/Spark-X2.5-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="XHToken/Spark-X2.5-4B", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("XHToken/Spark-X2.5-4B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use XHToken/Spark-X2.5-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "XHToken/Spark-X2.5-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XHToken/Spark-X2.5-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/XHToken/Spark-X2.5-4B
- SGLang
How to use XHToken/Spark-X2.5-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "XHToken/Spark-X2.5-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XHToken/Spark-X2.5-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "XHToken/Spark-X2.5-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XHToken/Spark-X2.5-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use XHToken/Spark-X2.5-4B with Docker Model Runner:
docker model run hf.co/XHToken/Spark-X2.5-4B
Greedy decoding emits token 128486 (orphan UTF-8 byte + Lao) inside Greek text → invalid UTF-8, llama-server HTTP 500
Hi, and thanks for Spark-X2.5-4B. We are fine-tuning it for Greek, code and tool calling, and it has been a great base.
While running our release checks we found a reproducible generation defect in the base model, and we wanted to report it with a minimal repro.
Prompt (Greek, translation task):
system: Είσαι βοηθός. Απάντησε σύντομα.
user: Μετάφρασε στα Ελληνικά: "The dog followed him to the end of the road."
Greedy decoding (do_sample=False / temperature 0), enable_thinking=False.
Output (transformers, bf16 and fp32, all files from this repo at commit 0bcb356):
Το σκύρο τον�ິດຕούσε μέχρι το τέλος της δρόμου.
Tokens:
['Το', 'ĠÏĥκ', 'Ïį', 'Ïģο', 'ĠÏĦον', 'ķິàºĶàºķ', 'οÏį', 'Ïĥε', ...]
In the middle of a Greek word the model picks token id 128486 (ķິàºĶàºķ). That token begins with an orphan UTF-8 continuation byte followed by Lao characters, so any sequence containing it decodes to invalid UTF-8 (�).
Minimal transformers repro
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
p = "XHToken/Spark-X2.5-4B"
tok = AutoTokenizer.from_pretrained(p, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(p, torch_dtype=torch.bfloat16, trust_remote_code=True).eval()
msgs = [{"role": "system", "content": "Είσαι βοηθός. Απάντησε σύντομα."},
{"role": "user", "content": 'Μετάφρασε στα Ελληνικά: "The dog followed him to the end of the road."'}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, enable_thinking=False, return_tensors="pt")
out = model.generate(ids, max_new_tokens=40, do_sample=False)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))
(Tested with transformers 4.57.1, torch 2.14.0 CPU.)
Consequence in llama.cpp: the same output appears with an F16 GGUF converted from this repo (convert_hf_to_gguf.py, build 10909 / commit 329b6160f). llama-server --jinja cannot parse it and returns HTTP 500:
common_chat_peg_parse: unparsed peg-native output: Το σκύρο τον�ິດຕούσε μέχρι το τέλος της δρόμου.
Lower-bit quantizations sometimes pick a different token at that step and do not crash. Rounding hides the defect there; it does not fix it.
A fine-tune of ours on top of the base shows the same token at the same position, so the defect appears to come from the base weights.
Happy to share more prompts or logs if useful.
Apollon Labs
Follow-up: mitigation we applied
Following up on the orphan-byte issue above: this is how we mitigated it on our fine-tune (Heliactis 1, 4B, based on Spark-X2.5-4B).
What worked (weights): on-policy unlikelihood + SFT on a LoRA.
- We generated the model's own answers to 428 prompts and penalised the probability mass on the 502 orphan-byte token IDs at every generated position. We used the sum of that mass, not the mean, with unlikelihood weight 10.
- Unlikelihood on the static training set gave almost no signal (mass ~2e-5). The problem only shows up on-policy.
- Result on the failing prompt ("The dog followed him to the end of the road.", greedy): the max ban mass dropped from 0.127 to 0.0017. There were no orphan tokens in the output and no HTTP 500. On 20 held-out prompts the max mass is 0.0002.
The trade-off we measured: pushing the mass lower (0.0010, then 0.0001) started damaging Greek vocabulary. The model began producing non-words, and our vocabulary gate fell from 0.955 to 0.932. So we stopped at the version that passes every gate.
What we recommend on top (inference): a logit_bias of −∞ on the 502 IDs. The list is here: https://huggingface.co/ApollonLabs/Heliactis-1-4B-GGUF/blob/main/ban-ids.json. A system prompt does not fix it; the failing prompt already runs with one.
A head-only edit without training did not work for us.
Model + card: https://huggingface.co/ApollonLabs/Heliactis-1-4B-GGUF · full reproduction: https://huggingface.co/ApollonLabs/Heliactis-1-4B-GGUF/blob/main/DEFECT.md