YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

se_model: Qwen/Qwen3-8B library_name: peft tags: - qwen3 - lora - peft - commonsense-reasoning - spectral-surgery

Qwen3-8B + Commonsense170K β€” LoRA

LoRA adapter obtained by supervised fine-tuning Qwen3-8B on Commonsense170K.

This checkpoint serves as the source adapter for the corresponding Spectral Surgery experiment.

Training

  • Dataset: Commonsense170K
  • Samples: 170,420
  • Epochs: 2
  • Sequence length: 2048
  • Global batch size: 32
  • Learning rate: 2e-4
  • Scheduler: cosine
  • Warmup ratio: 0.10
  • LoRA rank: 16
  • LoRA alpha: 32
  • LoRA dropout: 0.05
  • Target modules: all linear layers
  • Precision: bf16
  • Seed: 42
  • Chat mode: Qwen3 non-thinking

Evaluation

Eight-task commonsense suite:

  • BoolQ
  • PIQA
  • SocialIQA
  • HellaSwag
  • WinoGrande
  • ARC-Easy
  • ARC-Challenge
  • OpenBookQA
Metric Result
Macro Accuracy 90.7879%
Micro Accuracy 91.8373%
Correct 20,589 / 22,419

Evaluation uses the tokenizer chat template in non-thinking mode, greedy decoding, and a maximum of 8 generated tokens.

Spectral Surgery

This adapter is the source checkpoint for the corresponding Spectral Surgery model.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including tianzl66/Qwen3-8B-CommonSense170K-LoRA