SAT Tutor - Qwen2.5-7B (GGUF Q4_K_M)

Developed by AGK FIRE INC.

A Qwen2.5-7B-Instruct fine-tune that tutors for the SAT (math + reading/writing), packaged as a Q4_K_M GGUF for local inference with llama.cpp, Ollama, LM Studio, and friends.

What it does

Works through SAT questions step by step and finishes with a clear Answer: X line. Trained on SAT-style math and reading/writing questions with step-by-step solutions.

Held-out eval of the 15k tuned adapter (the GGUF is this adapter quantized; the file itself has not been re-evaluated):

Section Parseable Strict
Math 70/77 (90.9%) 70/100 (70%)
Reading/Writing 45/59 (76.3%) 45/60 (75%)

"Parseable" counts questions where the model emitted the expected Answer: X format; "strict" counts unparseable answers as wrong. Unparseable usually means a formatting miss rather than a wrong answer.

Usage

llama.cpp

./llama-cli -m sat-tutor-qwen2.5-7b.Q4_K_M.gguf -n 512 \
 -p "Solve step by step. If x + 5 = 12, what is x?"

Ollama

# Modelfile
FROM ./sat-tutor-qwen2.5-7b.Q4_K_M.gguf
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>{{ end }}<|im_start|>assistant
"""
PARAMETER stop "<|im_end|>"
ollama create sat-tutor -f Modelfile
ollama run sat-tutor

Running Ollama in Docker

docker run -d --name ollama -p 11434:11434 -v ollama:/root/.ollama ollama/ollama
docker cp sat-tutor-qwen2.5-7b.Q4_K_M.gguf ollama:/root/
docker cp Modelfile ollama:/root/
docker exec -w /root ollama ollama create sat-tutor -f Modelfile
docker exec ollama ollama run sat-tutor

LM Studio / GPT4All / text-generation-webui

Drop the .gguf into the app's models folder - it auto-detects the Qwen2 chat template.

Prompt format

ChatML (Qwen2.5-Instruct):

<|im_start|>system
You are a helpful SAT tutor.<|im_end|>
<|im_start|>user
Solve step by step. If x + 5 = 12, what is x?<|im_end|>
<|im_start|>assistant

The model is trained to reason step by step and end with Answer: X.

Training details

  • Base: Qwen2.5-7B-Instruct, 4-bit QLoRA
  • Data: ~15k SAT-style examples (random sample of a 31k pool), seq len 1024
  • Adapter merged into base weights (fp16), then quantized: fp16 -> Q8_0 -> Q4_K_M
  • Note: the planned 30k run (seq 2048, upsampled reading/writing) has not trained yet

Limitations

  • SAT-focused: off-topic questions get base-model-quality answers at best.
  • 4-bit quantization trades a little accuracy for size; the fp16 merge is the full-quality reference.
  • It can still make mistakes - double-check the math against the steps shown.

License

Apache 2.0 (inherited from Qwen2.5).


© 2026 AGK FIRE INC. Released under Apache 2.0.

Downloads last month
152
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for agk4444/sat-tutor-qwen2.5-7b-gguf

Base model

Qwen/Qwen2.5-7B
Quantized
(428)
this model