SAT Tutor - Qwen3-8B (GGUF Q4_K_M)

by AGK FIRE INC

A Qwen3-8B fine-tune that tutors for the SAT (math + reading/writing), packaged as a Q4_K_M GGUF for local inference with llama.cpp, Ollama, LM Studio, and friends.

What it does

Works through SAT questions step by step and finishes with a clear Answer: X line. Trained on SAT-style math and reading/writing questions with step-by-step solutions.

Held-out eval: 160 questions, identical prompt and decoding for both (greedy, 400 tokens, thinking off), strict scoring = unparseable counted wrong. Reports: eval_base.json, eval_tuned.json.

  • Math: base 57% (57/100) โ†’ tuned 71% (71/100); tuned answered 97/100 vs base 71/100
  • Reading/Writing: base 76.7% (46/60) โ†’ tuned 66.7% (40/60) โ€” a real regression; the adapter overfit toward math and wants more RW data before the teacher is frozen
  • Overall: base 64.4% (103/160) โ†’ tuned 69.4% (111/160)

The tuned adapter was chosen as the quantization target. Q4_K_M on the same suite: 66.9% strict (107/160) โ€” math 67/90 answered (74.4%), reading/writing 40/60 (66.7%), vs tuned fp16 69.4%. Quantization tax โ‰ˆ 2.5 pts, all on math; reading/writing unchanged. Report: eval_gguf_q4km.json.

"Parseable" counts questions where the model emitted the expected Answer: X format; "strict" counts unparseable answers as wrong. Unparseable usually means a formatting miss rather than a wrong answer.

Usage

llama.cpp

./llama-cli -m sat-tutor-qwen3-8b.Q4_K_M.gguf -n 512 \
  -p "Solve step by step. If x + 5 = 12, what is x?"

Ollama

# Modelfile
FROM ./sat-tutor-qwen3-8b.Q4_K_M.gguf
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>{{ end }}<|im_start|>assistant
"""
PARAMETER stop "<|im_end|>"
ollama create sat-tutor -f Modelfile
ollama run sat-tutor

Running Ollama in Docker

docker run -d --name ollama -p 11434:11434 -v ollama:/root/.ollama ollama/ollama
docker cp sat-tutor-qwen3-8b.Q4_K_M.gguf ollama:/root/
docker cp Modelfile ollama:/root/
docker exec -w /root ollama ollama create sat-tutor -f Modelfile
docker exec ollama ollama run sat-tutor

LM Studio / GPT4All / text-generation-webui

Drop the .gguf into the app's models folder - it auto-detects the Qwen3 chat template.

Prompt format

ChatML (Qwen3-Instruct):

<|im_start|>system
You are a helpful SAT tutor.<|im_end|>
<|im_start|>user
Solve step by step. If x + 5 = 12, what is x?<|im_end|>
<|im_start|>assistant

The model is trained to reason step by step and end with Answer: X.

Training details

  • Base: Qwen3-8B, 4-bit QLoRA
  • Data: 20k SAT-style examples per run (~12k math incl. MetaMathQA chain-of-thought, ~8k reading/writing incl. RACE), seq len 2048
  • Adapter merged into base weights (fp16), then quantized: fp16 -> Q8_0 -> Q4_K_M

Limitations

  • SAT-focused: off-topic questions get base-model-quality answers at best.
  • 4-bit quantization trades a little accuracy for size; the fp16 merge is the full-quality reference.
  • It can still make mistakes - double-check the math against the steps shown.

License

ยฉ 2026 AGK FIRE INC. Released under Apache 2.0. (inherited from Qwen3).

Downloads last month
95
GGUF
Model size
8B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for agk4444/sat-tutor-qwen3-8b-gguf

Finetuned
Qwen/Qwen3-8B
Quantized
(450)
this model