sat-tutor-qwen3-8b / README.md
agk4444's picture
Update README.md
7b1924d verified
|
Raw History Blame Contribute Delete
2.8 kB
---
license: apache-2.0
library_name: peft
base_model: Qwen/Qwen3-8B
pipeline_tag: text-generation
language:
- en
tags:
- sat
- tutoring
- math
- qlora
- qwen3
---
# SAT Tutor — Qwen3-8B adapter
by **AGK FIRE INC**
A QLoRA fine-tune of `Qwen/Qwen3-8B` that tutors for the SAT (math + reading & writing).
Trained on ~17k SAT-style examples (60% math, 40% reading & writing), one epoch,
2048-token training length, with full step-by-step solutions plus hint-style answers.
## Base vs tuned — held-out eval (2026-09-26)
160 held-out questions (100 math, 60 reading & writing). Identical prompt and decoding
for both: greedy, 400 new tokens, thinking off, same `**Answer: X**` format instruction.
Strict scoring = unparseable counted wrong. Reports: `eval_base.json`, `eval_tuned.json`.
| | Base (strict) | Tuned (strict) |
|---|---|---|
| Math | 57% (57/100) | **71%** (71/100) |
| Reading & Writing | 76.7% (46/60) | 66.7% (40/60) |
| Overall | 64.4% (103/160) | **69.4%** (111/160) |
The tuned adapter also answers far more math questions (97/100 vs 71/100) — base
refused 29 outright. Known weakness: reading & writing regressed 10 points, the
adapter overfit toward math. More RW training data is recommended before the
teacher is frozen.
## Usage
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen3-8B", torch_dtype=torch.float16,
device_map="auto", trust_remote_code=True)
model = PeftModel.from_pretrained(model, "agk4444/sat-tutor-qwen3-8b")
model.eval()
msgs = [{"role": "user", "content": "If x + 5 = 12, what is x?\n\n"
"End your response with your final answer on its own line in exactly "
"this format: **Answer: X** where X is A, B, C, or D."}]
prompt = tok.apply_chat_template(msgs, tokenize=False,
add_generation_prompt=True,
enable_thinking=False)
inp = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inp, max_new_tokens=400, do_sample=False)
print(tok.decode(out[0][inp["input_ids"].shape[1]:], skip_special_tokens=True))
```
## Quantized version
For local/phone inference (llama.cpp, Ollama, LM Studio), use the GGUF build:
**[agk4444/sat-tutor-qwen3-8b-gguf](https://huggingface.co/agk4444/sat-tutor-qwen3-8b-gguf)**
(`sat-tutor-qwen3-8b.Q4_K_M.gguf`, ~4.9 GB)
## Limitations
- SAT-focused: off-topic questions get base-model-quality answers at best.
- Reading & writing trails math — see the eval table above.
- It can still make mistakes — double-check the math against the steps shown.
---
© 2026 AGK FIRE INC. Released under Apache 2.0.