Khayyam Math β€” Qwen 2.5-7B v5.1 (LoRA adapter)

A LoRA fine-tune of Qwen2.5-7B-Instruct specialised for generating deterministic SVG figures and learner-facing narrations for math education. v5.1 is the first release trained on production telemetry from khayyammath.com β€” 52 anonymised turns layered on top of the 2,350-example synthetic teacher corpus.

v5.1 ties v4 on the 20-prompt held-out practical-test battery (20 / 20 vs 20 / 20, zero regressions). The qualitative jump over v4 will need a richer eval to demonstrate; v5.1 is the v4 floor + a process upgrade (production telemetry now reaches training).

⚠️ Use khayyam_math >= 0.4.1 to load this model. Earlier releases of the package have a JSON-extractor bug that mis-handles roughly 15 % of v5.1's structured outputs (model emits JSON-escaped SVG that breaks the old fallback path). v0.4.1 ships the unescape fallback.


Quick start (with the khayyam-math package, β‰₯ 0.4.1)

pip install "khayyam-math[qwen] @ git+https://github.com/khayyam-math/khayyam-math"
from khayyam_math import KhayyamMath

client = KhayyamMath(
    provider="qwen",
    model="khayyam-math/khayyam-math-qwen2.5-7b-v5.1",
)
result = client.generate("Solve x^2 - 5x + 6 = 0")
print(result.svg)

Quick start (raw transformers + peft)

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch

base    = "Qwen/Qwen2.5-7B-Instruct"
adapter = "khayyam-math/khayyam-math-qwen2.5-7b-v5.1"

tokenizer = AutoTokenizer.from_pretrained(adapter)
model = AutoModelForCausalLM.from_pretrained(
    base, torch_dtype=torch.bfloat16, device_map="auto",
)
model = PeftModel.from_pretrained(model, adapter)

Serving with vLLM

vllm serve Qwen/Qwen2.5-7B-Instruct \
  --enable-lora \
  --lora-modules khayyam-v5.1=khayyam-math/khayyam-math-qwen2.5-7b-v5.1 \
  --max-lora-rank 8 --dtype bfloat16

(Note --max-lora-rank 8, halved from v4's 16.)


Training

Field Value vs v4
Base Qwen/Qwen2.5-7B-Instruct same
Method LoRA (rank 8, alpha 16, dropout 0.05) rank halved (16 β†’ 8)
Target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj same
Trainable params 20.2 M (0.26 % of base) halved
Optimiser AdamW, lr 2e-4 same
Schedule 2 epochs, batch 1 Γ— grad-accum 4 1 fewer epoch
Precision bf16 same
Max sequence length 6144 tokens same
Training corpus 2,402 (prompt β†’ spec) pairs smaller (was 5,528)
– synthetic (gpt-4o-mini teacher + inspector filter) 2,350
– production sft-clean (khayyammath.com, last 30 days) 39 new for v5.1
– production sft-corrected (repair pairs) 13 new for v5.1
Final training loss 0.0638 higher than v5 (less over-fit)
Final token accuracy 0.987 comparable
Hardware Single RTX 5090 (32 GB) same
Wall-clock 1 h 59 m (7,136 s) ~60 % faster
Trained on 2026-05-23

Why v5.1 and not v5

v5 (rank 16, alpha 32, 3 epochs) over-fit: training loss collapsed to 0.048, and at inference the model emitted truncated or empty SVG on the most element-dense prompts and started breaking JSON structure across raw newlines. v5 failed the practical test (15 / 20) and was archived as qwen_lora_v5_rejected/. v5.1 keeps the same corpus but halves LoRA capacity (rank 8) and drops one epoch β€” the "less is more" pattern PEFT papers recommend for small-corpus fine-tunes.

Why production telemetry matters

The 52 production turns mined under ToS Β§5 (anonymised; salted IP hash; email never linked to training rows) are the first signal the model has seen that wasn't generated by another model. They cover the exact distribution of prompts real learners issue (matrix inverse, DFA, Bayes, vector fields, group theory) rather than what an LLM teacher imagines learners would issue. This is the substrate the model needs to learn to specialise; v5.1 is the start of that loop.


Evaluation

20-prompt held-out battery, scored by xml-parse + drawable-element presence:

Metric v3 v4 v5.1
Valid SVG rate 16 / 20 20 / 20 20 / 20
Regressions vs v4 β€” β€” 0
Recovered v4 known failures n/a n/a both (set_venn, linalg_eigen pass)

Full per-prompt results in runs/v5.1_practical_test/summary_rescore.json of the source repo.

Intended use, limitations, license

Same as v4: generating structured math figure specs (SVG + narration + math_claims) from natural-language prompts. English-only training. Adapter only (base downloads separately ~14 GB). MIT licence.

Citation

A peer-reviewed paper describing this work has not yet been published. If you'd like to cite the model in the meantime, please use the model's Hugging Face URL together with the release version in your bibliography:

khayyam-math/khayyam-math-qwen2.5-7b-v5.1
https://huggingface.co/khayyam-math/khayyam-math-qwen2.5-7b-v5.1
Released 2026-05-23, MIT licence.
Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for khayyam-math/khayyam-math-qwen2.5-7b-v5.1

Base model

Qwen/Qwen2.5-7B
Adapter
(2786)
this model