Qwen2.5-1.5B-GSM8K-CommonsenseQA-SFT

Fine-tuned version of Qwen/Qwen2.5-1.5B produced as part of the research study "What Changes More: Output Accuracy or Reasoning Structure Under Fine-Tuning?"

This is the best-performing SFT checkpoint (Trial SFT-1) selected by BERTScore F1 + BLEU on a hand-curated 54-prompt evaluation set (arithmetic / logical deduction / multi-step word problems).

Training Details

Parameter Value
Base model Qwen/Qwen2.5-1.5B
Method Supervised Fine-Tuning (SFT) via LoRA
LoRA rank 4
LoRA alpha 32
LoRA target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Learning rate 1e-5
Batch size 4
Epochs 1
Training data GSM8K (math word problems) + CommonsenseQA (commonsense QA)
Framework TRL SFTTrainer + PEFT

The LoRA adapters have been merged into the base model weights. This checkpoint is fully self-contained — no PEFT dependency required at inference.

Evaluation Results

Evaluated on 54 hand-curated prompts across three categories (18 each): arithmetic, logical deduction, and multi-step word problems. Each prompt has three variants: clean, paraphrased, and distractor.

Metric Base SFT-1 (this model)
LLM Judge Accuracy 94.4% 92.6%
Exact Match 59.3% 53.7%
BERTScore F1 0.7645 0.7772
BLEU 0.0307 0.0373
Answer Stability Rate 55.6% 64.8%
Distractor Degradation 16.7% 11.1%
Reasoning Judge Score (1–5) 4.69 4.85

Model selection criterion (per study protocol): BERTScore F1 + BLEU (primary); LLM judge scores are qualitative only and excluded from selection.

Key finding: lexical alignment, reasoning structure, and robustness all improve under SFT, while judge accuracy slightly regresses — consistent with the study hypothesis that structural improvements decouple from accuracy gains under supervised fine-tuning.

How to Use

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "syedtaha22/Qwen2.5-1.5B-GSM8K-CommonsenseQA-SFT",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
    "syedtaha22/Qwen2.5-1.5B-GSM8K-CommonsenseQA-SFT"
)

prompt = "A baker makes 48 cookies. She sells 1/3 of them in the morning and half of the remainder in the afternoon. How many cookies does she have left?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    output = model.generate(**inputs, max_new_tokens=256, temperature=0.1, do_sample=True)

print(tokenizer.decode(output[0], skip_special_tokens=True))

Limitations

  • Trained for 1 epoch only due to compute constraints. Multi-epoch training may yield qualitatively different accuracy-structure tradeoffs.
  • LoRA rank 4 constrains weight updates to a low-rank subspace. Full fine-tuning may produce larger structural changes.
  • Logical deduction was absent from training data (OOD category). Accuracy on deductive reasoning tasks does not improve; any structural gains there are generalisation effects, not trained outcomes.
  • All DPO preference pairs are derived from arithmetic data only (GSM8K). This model uses SFT only; the DPO variant is a separate checkpoint.
  • Single model family (Qwen2.5-1.5B). Findings may not generalise across architectures or parameter scales.

Citation

If you use this model, please cite the underlying base model:

@misc{qwen2025qwen25technicalreport,
  title={Qwen2.5 Technical Report},
  author={Qwen et al.},
  year={2025},
  eprint={2412.15115},
  archivePrefix={arXiv},
}
Downloads last month
14
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for syedtaha22/Qwen2.5-1.5B-GSM8K-CommonsenseQA-SFT

Adapter
(464)
this model

Datasets used to train syedtaha22/Qwen2.5-1.5B-GSM8K-CommonsenseQA-SFT

Paper for syedtaha22/Qwen2.5-1.5B-GSM8K-CommonsenseQA-SFT