--- license: apache-2.0 library_name: peft base_model: Qwen/Qwen3-8B pipeline_tag: text-generation language: - en tags: - sat - tutoring - math - qlora - qwen3 --- # SAT Tutor — Qwen3-8B adapter by **AGK FIRE INC** A QLoRA fine-tune of `Qwen/Qwen3-8B` that tutors for the SAT (math + reading & writing). Trained on ~17k SAT-style examples (60% math, 40% reading & writing), one epoch, 2048-token training length, with full step-by-step solutions plus hint-style answers. ## Base vs tuned — held-out eval (2026-09-26) 160 held-out questions (100 math, 60 reading & writing). Identical prompt and decoding for both: greedy, 400 new tokens, thinking off, same `**Answer: X**` format instruction. Strict scoring = unparseable counted wrong. Reports: `eval_base.json`, `eval_tuned.json`. | | Base (strict) | Tuned (strict) | |---|---|---| | Math | 57% (57/100) | **71%** (71/100) | | Reading & Writing | 76.7% (46/60) | 66.7% (40/60) | | Overall | 64.4% (103/160) | **69.4%** (111/160) | The tuned adapter also answers far more math questions (97/100 vs 71/100) — base refused 29 outright. Known weakness: reading & writing regressed 10 points, the adapter overfit toward math. More RW training data is recommended before the teacher is frozen. ## Usage ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer from peft import PeftModel tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( "Qwen/Qwen3-8B", torch_dtype=torch.float16, device_map="auto", trust_remote_code=True) model = PeftModel.from_pretrained(model, "agk4444/sat-tutor-qwen3-8b") model.eval() msgs = [{"role": "user", "content": "If x + 5 = 12, what is x?\n\n" "End your response with your final answer on its own line in exactly " "this format: **Answer: X** where X is A, B, C, or D."}] prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False) inp = tok(prompt, return_tensors="pt").to(model.device) out = model.generate(**inp, max_new_tokens=400, do_sample=False) print(tok.decode(out[0][inp["input_ids"].shape[1]:], skip_special_tokens=True)) ``` ## Quantized version For local/phone inference (llama.cpp, Ollama, LM Studio), use the GGUF build: **[agk4444/sat-tutor-qwen3-8b-gguf](https://huggingface.co/agk4444/sat-tutor-qwen3-8b-gguf)** (`sat-tutor-qwen3-8b.Q4_K_M.gguf`, ~4.9 GB) ## Limitations - SAT-focused: off-topic questions get base-model-quality answers at best. - Reading & writing trails math — see the eval table above. - It can still make mistakes — double-check the math against the steps shown. --- © 2026 AGK FIRE INC. Released under Apache 2.0.