# Sophia GGUF Sophia's llama.cpp/GGUF build — a ~1B-parameter Chinese chat model (KDA+MLA hybrid, pretrained from scratch on a single RTX 5090). Runs on llama.cpp's `kimi-k3` architecture (Sophia is a fully dense 28-layer instance: [KDA,KDA,KDA,MLA] layer pattern, bounded decay gate, cross-layer attention residuals, SiTU-GLU). ## Files | File | Size | Notes | |---|---|---| | `sophia-bf16.gguf` | 2.1 GB | BF16, best quality | | `sophia-Q4_K_M.gguf` | 685 MB | Q4_K_M; precision-sensitive tensors (KDA scalars, norms, conv, res_score) stay high-precision via fallback | ## Requirements Needs an llama.cpp build with the `sophia` pre-tokenizer patch (~15 lines, see [Sophia repo](https://github.com/Arain119/Sophia) `tools/gguf/llama_cpp_sophia.patch`). The GGUF embeds `tokenizer.ggml.pre = "sophia"`: Sophia's tokenizer.json serializes its Split patterns as literals that never match, so the effective tokenization is ByteLevel+BPE over the whole input — the pre-tokenizer must not split at all. ```bash # chat (embedded chat template applies automatically) llama-cli -m sophia-Q4_K_M.gguf -st -t 8 # server llama-server -m sophia-bf16.gguf -c 4096 -t 8 --port 8180 ``` ## Parity evidence (BF16 / CPU) - Tokenization: **151/151 exact token-id match** vs the HF tokenizer - Teacher-forced logits: **145/150 argmax agreement (96.7%)**, mean |Δlogprob| = 0.0137; every divergence is a top-2 near-tie (gap ≤ 0.06) - Chat-template generation: byte-identical to the HF reference Converter and reproduction scripts: [Sophia repo](https://github.com/Arain119/Sophia) `tools/gguf/`; details in `docs/gguf_llamacpp.md`.