sophia / gguf /README.md
Arain119's picture
Upload gguf/README.md with huggingface_hub
0de0b89 verified
|
Raw History Blame Contribute Delete
1.66 kB
# Sophia GGUF
Sophia's llama.cpp/GGUF build — a ~1B-parameter Chinese chat model (KDA+MLA hybrid, pretrained from scratch on a single RTX 5090).
Runs on llama.cpp's `kimi-k3` architecture (Sophia is a fully dense 28-layer instance: [KDA,KDA,KDA,MLA] layer pattern, bounded decay gate, cross-layer attention residuals, SiTU-GLU).
## Files
| File | Size | Notes |
|---|---|---|
| `sophia-bf16.gguf` | 2.1 GB | BF16, best quality |
| `sophia-Q4_K_M.gguf` | 685 MB | Q4_K_M; precision-sensitive tensors (KDA scalars, norms, conv, res_score) stay high-precision via fallback |
## Requirements
Needs an llama.cpp build with the `sophia` pre-tokenizer patch (~15 lines, see [Sophia repo](https://github.com/Arain119/Sophia) `tools/gguf/llama_cpp_sophia.patch`). The GGUF embeds `tokenizer.ggml.pre = "sophia"`: Sophia's tokenizer.json serializes its Split patterns as literals that never match, so the effective tokenization is ByteLevel+BPE over the whole input — the pre-tokenizer must not split at all.
```bash
# chat (embedded chat template applies automatically)
llama-cli -m sophia-Q4_K_M.gguf -st -t 8
# server
llama-server -m sophia-bf16.gguf -c 4096 -t 8 --port 8180
```
## Parity evidence (BF16 / CPU)
- Tokenization: **151/151 exact token-id match** vs the HF tokenizer
- Teacher-forced logits: **145/150 argmax agreement (96.7%)**, mean |Δlogprob| = 0.0137; every divergence is a top-2 near-tie (gap ≤ 0.06)
- Chat-template generation: byte-identical to the HF reference
Converter and reproduction scripts: [Sophia repo](https://github.com/Arain119/Sophia) `tools/gguf/`; details in `docs/gguf_llamacpp.md`.