Instructions to use khayyam-math/khayyam-math-qwen2.5-7b-v5.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use khayyam-math/khayyam-math-qwen2.5-7b-v5.1 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct") model = PeftModel.from_pretrained(base_model, "khayyam-math/khayyam-math-qwen2.5-7b-v5.1") - Notebooks
- Google Colab
- Kaggle
Khayyam Math β Qwen 2.5-7B v5.1 (LoRA adapter)
A LoRA fine-tune of Qwen2.5-7B-Instruct specialised for generating deterministic SVG figures and learner-facing narrations for math education. v5.1 is the first release trained on production telemetry from khayyammath.com β 52 anonymised turns layered on top of the 2,350-example synthetic teacher corpus.
v5.1 ties v4 on the 20-prompt held-out practical-test battery (20 / 20 vs 20 / 20, zero regressions). The qualitative jump over v4 will need a richer eval to demonstrate; v5.1 is the v4 floor + a process upgrade (production telemetry now reaches training).
- π Live demo: khayyammath.com
- π¦ Source code & package: github.com/khayyam-math/khayyam-math
- π Benchmark: khayyam-math/lean-math-1000 (coming)
β οΈ Use khayyam_math >= 0.4.1 to load this model. Earlier
releases of the package have a JSON-extractor bug that mis-handles
roughly 15 % of v5.1's structured outputs (model emits JSON-escaped
SVG that breaks the old fallback path). v0.4.1 ships the unescape
fallback.
Quick start (with the khayyam-math package, β₯ 0.4.1)
pip install "khayyam-math[qwen] @ git+https://github.com/khayyam-math/khayyam-math"
from khayyam_math import KhayyamMath
client = KhayyamMath(
provider="qwen",
model="khayyam-math/khayyam-math-qwen2.5-7b-v5.1",
)
result = client.generate("Solve x^2 - 5x + 6 = 0")
print(result.svg)
Quick start (raw transformers + peft)
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch
base = "Qwen/Qwen2.5-7B-Instruct"
adapter = "khayyam-math/khayyam-math-qwen2.5-7b-v5.1"
tokenizer = AutoTokenizer.from_pretrained(adapter)
model = AutoModelForCausalLM.from_pretrained(
base, torch_dtype=torch.bfloat16, device_map="auto",
)
model = PeftModel.from_pretrained(model, adapter)
Serving with vLLM
vllm serve Qwen/Qwen2.5-7B-Instruct \
--enable-lora \
--lora-modules khayyam-v5.1=khayyam-math/khayyam-math-qwen2.5-7b-v5.1 \
--max-lora-rank 8 --dtype bfloat16
(Note --max-lora-rank 8, halved from v4's 16.)
Training
| Field | Value | vs v4 |
|---|---|---|
| Base | Qwen/Qwen2.5-7B-Instruct |
same |
| Method | LoRA (rank 8, alpha 16, dropout 0.05) | rank halved (16 β 8) |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
same |
| Trainable params | 20.2 M (0.26 % of base) | halved |
| Optimiser | AdamW, lr 2e-4 | same |
| Schedule | 2 epochs, batch 1 Γ grad-accum 4 | 1 fewer epoch |
| Precision | bf16 | same |
| Max sequence length | 6144 tokens | same |
| Training corpus | 2,402 (prompt β spec) pairs | smaller (was 5,528) |
| β synthetic (gpt-4o-mini teacher + inspector filter) | 2,350 | |
| β production sft-clean (khayyammath.com, last 30 days) | 39 | new for v5.1 |
| β production sft-corrected (repair pairs) | 13 | new for v5.1 |
| Final training loss | 0.0638 | higher than v5 (less over-fit) |
| Final token accuracy | 0.987 | comparable |
| Hardware | Single RTX 5090 (32 GB) | same |
| Wall-clock | 1 h 59 m (7,136 s) | ~60 % faster |
| Trained on | 2026-05-23 |
Why v5.1 and not v5
v5 (rank 16, alpha 32, 3 epochs) over-fit: training loss collapsed
to 0.048, and at inference the model emitted truncated or empty SVG
on the most element-dense prompts and started breaking JSON
structure across raw newlines. v5 failed the practical test (15 / 20)
and was archived as qwen_lora_v5_rejected/. v5.1 keeps the same
corpus but halves LoRA capacity (rank 8) and drops one epoch β the
"less is more" pattern PEFT papers recommend for small-corpus
fine-tunes.
Why production telemetry matters
The 52 production turns mined under ToS Β§5 (anonymised; salted IP hash; email never linked to training rows) are the first signal the model has seen that wasn't generated by another model. They cover the exact distribution of prompts real learners issue (matrix inverse, DFA, Bayes, vector fields, group theory) rather than what an LLM teacher imagines learners would issue. This is the substrate the model needs to learn to specialise; v5.1 is the start of that loop.
Evaluation
20-prompt held-out battery, scored by xml-parse + drawable-element presence:
| Metric | v3 | v4 | v5.1 |
|---|---|---|---|
| Valid SVG rate | 16 / 20 | 20 / 20 | 20 / 20 |
| Regressions vs v4 | β | β | 0 |
| Recovered v4 known failures | n/a | n/a | both (set_venn, linalg_eigen pass) |
Full per-prompt results in
runs/v5.1_practical_test/summary_rescore.json
of the source repo.
Intended use, limitations, license
Same as v4: generating structured math figure specs (SVG + narration + math_claims) from natural-language prompts. English-only training. Adapter only (base downloads separately ~14 GB). MIT licence.
Citation
A peer-reviewed paper describing this work has not yet been published. If you'd like to cite the model in the meantime, please use the model's Hugging Face URL together with the release version in your bibliography:
khayyam-math/khayyam-math-qwen2.5-7b-v5.1
https://huggingface.co/khayyam-math/khayyam-math-qwen2.5-7b-v5.1
Released 2026-05-23, MIT licence.
- Downloads last month
- 14