Qwen2.5-1.5B-GSM8K-CommonsenseQA-SFT
Fine-tuned version of Qwen/Qwen2.5-1.5B produced as part of the research study "What Changes More: Output Accuracy or Reasoning Structure Under Fine-Tuning?"
This is the best-performing SFT checkpoint (Trial SFT-1) selected by BERTScore F1 + BLEU on a hand-curated 54-prompt evaluation set (arithmetic / logical deduction / multi-step word problems).
Training Details
| Parameter | Value |
|---|---|
| Base model | Qwen/Qwen2.5-1.5B |
| Method | Supervised Fine-Tuning (SFT) via LoRA |
| LoRA rank | 4 |
| LoRA alpha | 32 |
| LoRA target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Learning rate | 1e-5 |
| Batch size | 4 |
| Epochs | 1 |
| Training data | GSM8K (math word problems) + CommonsenseQA (commonsense QA) |
| Framework | TRL SFTTrainer + PEFT |
The LoRA adapters have been merged into the base model weights. This checkpoint is fully self-contained — no PEFT dependency required at inference.
Evaluation Results
Evaluated on 54 hand-curated prompts across three categories (18 each): arithmetic, logical deduction, and multi-step word problems. Each prompt has three variants: clean, paraphrased, and distractor.
| Metric | Base | SFT-1 (this model) |
|---|---|---|
| LLM Judge Accuracy | 94.4% | 92.6% |
| Exact Match | 59.3% | 53.7% |
| BERTScore F1 | 0.7645 | 0.7772 |
| BLEU | 0.0307 | 0.0373 |
| Answer Stability Rate | 55.6% | 64.8% |
| Distractor Degradation | 16.7% | 11.1% |
| Reasoning Judge Score (1–5) | 4.69 | 4.85 |
Model selection criterion (per study protocol): BERTScore F1 + BLEU (primary); LLM judge scores are qualitative only and excluded from selection.
Key finding: lexical alignment, reasoning structure, and robustness all improve under SFT, while judge accuracy slightly regresses — consistent with the study hypothesis that structural improvements decouple from accuracy gains under supervised fine-tuning.
How to Use
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"syedtaha22/Qwen2.5-1.5B-GSM8K-CommonsenseQA-SFT",
torch_dtype=torch.bfloat16,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
"syedtaha22/Qwen2.5-1.5B-GSM8K-CommonsenseQA-SFT"
)
prompt = "A baker makes 48 cookies. She sells 1/3 of them in the morning and half of the remainder in the afternoon. How many cookies does she have left?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=256, temperature=0.1, do_sample=True)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Limitations
- Trained for 1 epoch only due to compute constraints. Multi-epoch training may yield qualitatively different accuracy-structure tradeoffs.
- LoRA rank 4 constrains weight updates to a low-rank subspace. Full fine-tuning may produce larger structural changes.
- Logical deduction was absent from training data (OOD category). Accuracy on deductive reasoning tasks does not improve; any structural gains there are generalisation effects, not trained outcomes.
- All DPO preference pairs are derived from arithmetic data only (GSM8K). This model uses SFT only; the DPO variant is a separate checkpoint.
- Single model family (Qwen2.5-1.5B). Findings may not generalise across architectures or parameter scales.
Citation
If you use this model, please cite the underlying base model:
@misc{qwen2025qwen25technicalreport,
title={Qwen2.5 Technical Report},
author={Qwen et al.},
year={2025},
eprint={2412.15115},
archivePrefix={arXiv},
}
- Downloads last month
- 14
Model tree for syedtaha22/Qwen2.5-1.5B-GSM8K-CommonsenseQA-SFT
Base model
Qwen/Qwen2.5-1.5B