SAT Tutor β€” Qwen3-8B adapter

by AGK FIRE INC

A QLoRA fine-tune of Qwen/Qwen3-8B that tutors for the SAT (math + reading & writing). Trained on ~17k SAT-style examples (60% math, 40% reading & writing), one epoch, 2048-token training length, with full step-by-step solutions plus hint-style answers.

Base vs tuned β€” held-out eval (2026-09-26)

160 held-out questions (100 math, 60 reading & writing). Identical prompt and decoding for both: greedy, 400 new tokens, thinking off, same **Answer: X** format instruction. Strict scoring = unparseable counted wrong. Reports: eval_base.json, eval_tuned.json.

Base (strict) Tuned (strict)
Math 57% (57/100) 71% (71/100)
Reading & Writing 76.7% (46/60) 66.7% (40/60)
Overall 64.4% (103/160) 69.4% (111/160)

The tuned adapter also answers far more math questions (97/100 vs 71/100) β€” base refused 29 outright. Known weakness: reading & writing regressed 10 points, the adapter overfit toward math. More RW training data is recommended before the teacher is frozen.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen3-8B", torch_dtype=torch.float16,
    device_map="auto", trust_remote_code=True)
model = PeftModel.from_pretrained(model, "agk4444/sat-tutor-qwen3-8b")
model.eval()

msgs = [{"role": "user", "content": "If x + 5 = 12, what is x?\n\n"
        "End your response with your final answer on its own line in exactly "
        "this format: **Answer: X** where X is A, B, C, or D."}]
prompt = tok.apply_chat_template(msgs, tokenize=False,
                                 add_generation_prompt=True,
                                 enable_thinking=False)
inp = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inp, max_new_tokens=400, do_sample=False)
print(tok.decode(out[0][inp["input_ids"].shape[1]:], skip_special_tokens=True))

Quantized version

For local/phone inference (llama.cpp, Ollama, LM Studio), use the GGUF build:

agk4444/sat-tutor-qwen3-8b-gguf (sat-tutor-qwen3-8b.Q4_K_M.gguf, ~4.9 GB)

Limitations

  • SAT-focused: off-topic questions get base-model-quality answers at best.
  • Reading & writing trails math β€” see the eval table above.
  • It can still make mistakes β€” double-check the math against the steps shown.

Β© 2026 AGK FIRE INC. Released under Apache 2.0.

Downloads last month
43
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for agk4444/sat-tutor-qwen3-8b

Finetuned
Qwen/Qwen3-8B
Adapter
(2207)
this model