sat-tutor-qwen3-8b / README.md
agk4444's picture
Update README.md
7b1924d verified
|
Raw History Blame Contribute Delete
2.8 kB
metadata
license: apache-2.0
library_name: peft
base_model: Qwen/Qwen3-8B
pipeline_tag: text-generation
language:
  - en
tags:
  - sat
  - tutoring
  - math
  - qlora
  - qwen3

SAT Tutor — Qwen3-8B adapter

by AGK FIRE INC

A QLoRA fine-tune of Qwen/Qwen3-8B that tutors for the SAT (math + reading & writing). Trained on ~17k SAT-style examples (60% math, 40% reading & writing), one epoch, 2048-token training length, with full step-by-step solutions plus hint-style answers.

Base vs tuned — held-out eval (2026-09-26)

160 held-out questions (100 math, 60 reading & writing). Identical prompt and decoding for both: greedy, 400 new tokens, thinking off, same **Answer: X** format instruction. Strict scoring = unparseable counted wrong. Reports: eval_base.json, eval_tuned.json.

Base (strict) Tuned (strict)
Math 57% (57/100) 71% (71/100)
Reading & Writing 76.7% (46/60) 66.7% (40/60)
Overall 64.4% (103/160) 69.4% (111/160)

The tuned adapter also answers far more math questions (97/100 vs 71/100) — base refused 29 outright. Known weakness: reading & writing regressed 10 points, the adapter overfit toward math. More RW training data is recommended before the teacher is frozen.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen3-8B", torch_dtype=torch.float16,
    device_map="auto", trust_remote_code=True)
model = PeftModel.from_pretrained(model, "agk4444/sat-tutor-qwen3-8b")
model.eval()

msgs = [{"role": "user", "content": "If x + 5 = 12, what is x?\n\n"
        "End your response with your final answer on its own line in exactly "
        "this format: **Answer: X** where X is A, B, C, or D."}]
prompt = tok.apply_chat_template(msgs, tokenize=False,
                                 add_generation_prompt=True,
                                 enable_thinking=False)
inp = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inp, max_new_tokens=400, do_sample=False)
print(tok.decode(out[0][inp["input_ids"].shape[1]:], skip_special_tokens=True))

Quantized version

For local/phone inference (llama.cpp, Ollama, LM Studio), use the GGUF build:

agk4444/sat-tutor-qwen3-8b-gguf (sat-tutor-qwen3-8b.Q4_K_M.gguf, ~4.9 GB)

Limitations

  • SAT-focused: off-topic questions get base-model-quality answers at best.
  • Reading & writing trails math — see the eval table above.
  • It can still make mistakes — double-check the math against the steps shown.

© 2026 AGK FIRE INC. Released under Apache 2.0.