How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf NotHereNorThere/Coral-2-4b:Q5_K_M
# Run inference directly in the terminal:
llama cli -hf NotHereNorThere/Coral-2-4b:Q5_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf NotHereNorThere/Coral-2-4b:Q5_K_M
# Run inference directly in the terminal:
llama cli -hf NotHereNorThere/Coral-2-4b:Q5_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf NotHereNorThere/Coral-2-4b:Q5_K_M
# Run inference directly in the terminal:
./llama-cli -hf NotHereNorThere/Coral-2-4b:Q5_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf NotHereNorThere/Coral-2-4b:Q5_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf NotHereNorThere/Coral-2-4b:Q5_K_M
Use Docker
docker model run hf.co/NotHereNorThere/Coral-2-4b:Q5_K_M
Quick Links

better late than never?

i made the base model a few months ago and never got around to finishing it. initially i did post train it properly... but i switched to a subsection of dolphin-R1 for Coral 1.6 and 2.0, which seemed like a way better idea than mixing openthoughts and openhermes, to try and buff out dynamic hybrid reasoning since it wasnt actually intentional. only after i finished the training run on the base model did i realize that my entire fine tuning workflow didnt support thinking blocks. that explained a lot. i didn't feel like fixing it then so i kept the base model and training set for later to finish and here we are.

to the like 12 people who downloaded the other Coral models, first of all thanks, i did not expect a single download, actual feedback on the models would be appreciated.

a whole 4 billion paramaters!?

yes i know, what a goliath of a model (at least for the Coral family so far), but seriously it's nothing special

  • a TIES merge from several fine tunes of Qwen 3 4b.
    • including the original and later updated checkpoint.
  • then a QLORA-merge trained on a subsection of dolphin-r1 to cement the CoT.

actual performance

  • it passed rigourorous vibe testing, given you use the correct infrerence/chat settings.
    • keeps qwen 3 4b's chat settings.
  • no refusal, TIES has a habit of removing post training censorship.
  • unfortunatly, thinking got steamrolled (TIES again). it still outputs <think> formatting but theyre always empty.
  • normal 4b model limitations.

quant guide

  • bf16 / dont bother, just use;
  • Q5 / near lossless and fits in a lot of devices
Downloads last month
21
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NotHereNorThere/Coral-2-4b

Finetuned
Qwen/Qwen3-4B
Quantized
(314)
this model

Collection including NotHereNorThere/Coral-2-4b