Tholos-2B-GGUF

GGUF builds of Tholos-2B, MiniCPM5-2B fine-tuned to be the agent in Tholos. The main card has the step format, the training data and the benchmark results. This page has the files and the commands.

File Quantization Size sha256
Tholos-2B-Q4_K_M.gguf Q4_K_M 1.56 GB 65f4700e0e107d52f5a80c3c513b3387e8cfcef9cb7545f0a2c0a2293d1eadc3
Tholos-2B-Q8_0.gguf Q8_0 2.68 GB 5e57b416fb515260276413a9fadc3a67ef3d81f6fceb427cefe475850df1c476

The repo's SHA256SUMS file lists the same sums. Q4_K_M is the file our benchmark runs use.

llama.cpp

llama-server -hf mertkayacs/Tholos-2B-GGUF:Q4_K_M --jinja -c 16384 -t 4 -a tholos-2b \
  --host 127.0.0.1 --port 8080

Swap Q4_K_M for Q8_0 to load the larger file. Send a json_schema response format with each request so the server constrains decoding; the main card has a complete request.

Ollama

ollama pull hf.co/mertkayacs/Tholos-2B-GGUF:Q4_K_M

The repo carries a template and a params file. The template renders prompts the way the model was trained, with an empty think block before each answer, and matches the training render byte for byte. The params file sets the stop tokens and a 16,384-token context (Ollama's default is 4,096).

Without a template file Ollama picks one automatically. On the base model's official Q4_K_M file that choice left out the empty think block, and the file passed 48 of 160 Tholos-Bench scenarios with it and 96 of 160 with our template.

In Tholos, press Detect in Settings and add the model. Tholos asks Ollama for JSON mode and sends reasoning_effort: "none" on its own.

License

Apache-2.0, the license of the base model MiniCPM5-2B. See the main card for the training data and the terms that apply to it.

Downloads last month
120
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support