--- license: mit language: - en library_name: peft base_model: Qwen/Qwen2.5-7B-Instruct pipeline_tag: text-generation tags: - lora - peft - math - math-education - svg - figure-generation - diagram - lean4 - mathlib - z3 - visualized-math-teaching - production-telemetry-trained --- # Khayyam Math โ€” Qwen 2.5-7B v5.1 (LoRA adapter) A LoRA fine-tune of **Qwen2.5-7B-Instruct** specialised for generating deterministic SVG figures and learner-facing narrations for math education. v5.1 is the **first release trained on production telemetry** from [**khayyammath.com**](https://khayyammath.com) โ€” 52 anonymised turns layered on top of the 2,350-example synthetic teacher corpus. v5.1 **ties** [v4](https://huggingface.co/khayyam-math/khayyam-math-qwen2.5-7b-v4) on the 20-prompt held-out practical-test battery (20 / 20 vs 20 / 20, zero regressions). The qualitative jump over v4 will need a richer eval to demonstrate; v5.1 is the v4 floor + a process upgrade (production telemetry now reaches training). - ๐ŸŒ **Live demo:** [khayyammath.com](https://khayyammath.com) - ๐Ÿ“ฆ **Source code & package:** [github.com/khayyam-math/khayyam-math](https://github.com/khayyam-math/khayyam-math) - ๐Ÿ“Š **Benchmark:** [khayyam-math/lean-math-1000](https://huggingface.co/datasets/khayyam-math/lean-math-1000) *(coming)* โš ๏ธ **Use `khayyam_math >= 0.4.1`** to load this model. Earlier releases of the package have a JSON-extractor bug that mis-handles roughly 15 % of v5.1's structured outputs (model emits JSON-escaped SVG that breaks the old fallback path). v0.4.1 ships the unescape fallback. --- ## Quick start (with the `khayyam-math` package, โ‰ฅ 0.4.1) ```bash pip install "khayyam-math[qwen] @ git+https://github.com/khayyam-math/khayyam-math" ``` ```python from khayyam_math import KhayyamMath client = KhayyamMath( provider="qwen", model="khayyam-math/khayyam-math-qwen2.5-7b-v5.1", ) result = client.generate("Solve x^2 - 5x + 6 = 0") print(result.svg) ``` ## Quick start (raw `transformers` + `peft`) ```python from transformers import AutoModelForCausalLM, AutoTokenizer from peft import PeftModel import torch base = "Qwen/Qwen2.5-7B-Instruct" adapter = "khayyam-math/khayyam-math-qwen2.5-7b-v5.1" tokenizer = AutoTokenizer.from_pretrained(adapter) model = AutoModelForCausalLM.from_pretrained( base, torch_dtype=torch.bfloat16, device_map="auto", ) model = PeftModel.from_pretrained(model, adapter) ``` ## Serving with vLLM ```bash vllm serve Qwen/Qwen2.5-7B-Instruct \ --enable-lora \ --lora-modules khayyam-v5.1=khayyam-math/khayyam-math-qwen2.5-7b-v5.1 \ --max-lora-rank 8 --dtype bfloat16 ``` (Note `--max-lora-rank 8`, halved from v4's 16.) --- ## Training | Field | Value | vs v4 | |---|---|---| | Base | `Qwen/Qwen2.5-7B-Instruct` | same | | Method | LoRA (rank 8, alpha 16, dropout 0.05) | rank halved (16 โ†’ 8) | | Target modules | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` | same | | Trainable params | **20.2 M** (0.26 % of base) | halved | | Optimiser | AdamW, lr 2e-4 | same | | Schedule | 2 epochs, batch 1 ร— grad-accum 4 | 1 fewer epoch | | Precision | bf16 | same | | Max sequence length | 6144 tokens | same | | Training corpus | **2,402** (prompt โ†’ spec) pairs | smaller (was 5,528) | | โ€“ synthetic (gpt-4o-mini teacher + inspector filter) | 2,350 | | | โ€“ production sft-clean (khayyammath.com, last 30 days) | **39** | new for v5.1 | | โ€“ production sft-corrected (repair pairs) | **13** | new for v5.1 | | Final training loss | **0.0638** | higher than v5 (less over-fit) | | Final token accuracy | **0.987** | comparable | | Hardware | Single RTX 5090 (32 GB) | same | | Wall-clock | 1 h 59 m (7,136 s) | ~60 % faster | | Trained on | 2026-05-23 | | ### Why v5.1 and not v5 v5 (rank 16, alpha 32, 3 epochs) over-fit: training loss collapsed to 0.048, and at inference the model emitted truncated or empty SVG on the most element-dense prompts and started breaking JSON structure across raw newlines. v5 failed the practical test (15 / 20) and was archived as `qwen_lora_v5_rejected/`. v5.1 keeps the same corpus but halves LoRA capacity (rank 8) and drops one epoch โ€” the "less is more" pattern PEFT papers recommend for small-corpus fine-tunes. ### Why production telemetry matters The 52 production turns mined under ToS ยง5 (anonymised; salted IP hash; email never linked to training rows) are the **first signal the model has seen that wasn't generated by another model**. They cover the exact distribution of prompts real learners issue (matrix inverse, DFA, Bayes, vector fields, group theory) rather than what an LLM teacher imagines learners would issue. This is the substrate the model needs to learn to specialise; v5.1 is the start of that loop. --- ## Evaluation 20-prompt held-out battery, scored by xml-parse + drawable-element presence: | Metric | v3 | v4 | **v5.1** | |---|---|---|---| | Valid SVG rate | 16 / 20 | 20 / 20 | **20 / 20** | | Regressions vs v4 | โ€” | โ€” | **0** | | Recovered v4 known failures | n/a | n/a | both (`set_venn`, `linalg_eigen` pass) | Full per-prompt results in [`runs/v5.1_practical_test/summary_rescore.json`](https://github.com/khayyam-math/khayyam-math/blob/main/runs/v5.1_practical_test/summary_rescore.json) of the source repo. ## Intended use, limitations, license Same as [v4](https://huggingface.co/khayyam-math/khayyam-math-qwen2.5-7b-v4): generating structured math figure specs (SVG + narration + math_claims) from natural-language prompts. English-only training. Adapter only (base downloads separately ~14 GB). MIT licence. ## Citation A peer-reviewed paper describing this work has not yet been published. If you'd like to cite the model in the meantime, please use the model's Hugging Face URL together with the release version in your bibliography: ``` khayyam-math/khayyam-math-qwen2.5-7b-v5.1 https://huggingface.co/khayyam-math/khayyam-math-qwen2.5-7b-v5.1 Released 2026-05-23, MIT licence. ```