How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf mertkayacs/Tholos-2B-GGUF:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf mertkayacs/Tholos-2B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf mertkayacs/Tholos-2B-GGUF:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf mertkayacs/Tholos-2B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf mertkayacs/Tholos-2B-GGUF:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf mertkayacs/Tholos-2B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf mertkayacs/Tholos-2B-GGUF:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf mertkayacs/Tholos-2B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/mertkayacs/Tholos-2B-GGUF:Q4_K_M
Quick Links

Tholos-2B-GGUF

A 2B agent model built by senior AI engineer Mert Kaya: it passes 137 of 160 Tholos-Bench scenarios on one Kaggle T4 GPU (Q4_K_M, llama.cpp with a JSON schema), 25 more than the model it was trained from (evaluation).

GGUF builds of Tholos-2B, MiniCPM5-2B fine-tuned to be the agent in Tholos. The main card has the step format, the training data and the benchmark results. This page has the files and the commands.

File Quantization Size sha256
Tholos-2B-Q4_K_M.gguf Q4_K_M 1.56 GB 65f4700e0e107d52f5a80c3c513b3387e8cfcef9cb7545f0a2c0a2293d1eadc3
Tholos-2B-Q8_0.gguf Q8_0 2.68 GB 5e57b416fb515260276413a9fadc3a67ef3d81f6fceb427cefe475850df1c476

The repo's SHA256SUMS file lists the same sums. Q4_K_M is the file our benchmark runs use.

llama.cpp

llama-server -hf mertkayacs/Tholos-2B-GGUF:Q4_K_M --jinja -c 16384 -t 4 -a tholos-2b \
  --host 127.0.0.1 --port 8080

Swap Q4_K_M for Q8_0 to load the larger file. Send a json_schema response format with each request so the server constrains decoding; the main card has a complete request.

Ollama

ollama pull hf.co/mertkayacs/Tholos-2B-GGUF:Q4_K_M

The repo carries a template and a params file. The template renders prompts the way the model was trained, with an empty think block before each answer, and matches the training render byte for byte. The params file sets the stop tokens and a 16,384-token context (Ollama's default is 4,096).

Without a template file Ollama picks one automatically. On the base model's official Q4_K_M file that choice left out the empty think block, and the file passed 48 of 160 Tholos-Bench scenarios with it and 96 of 160 with our template.

In Tholos, press Detect in Settings and add the model. Tholos asks Ollama for JSON mode and sends reasoning_effort: "none" on its own.

License

Apache-2.0, the license of the base model MiniCPM5-2B. See the main card for the training data and the terms that apply to it.


Eschatia Labs

An Eschatia Labs project. Built by Mert Kaya.

Downloads last month
123
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for mertkayacs/Tholos-2B-GGUF

Quantized
(2)
this model