YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

se_model: Qwen/Qwen3-8B library_name: peft tags: - qwen3 - lora - peft - commonsense-reasoning - spectral-surgery

Qwen3-8B + Commonsense170K β€” Spectral Surgery

Post-hoc Spectral Surgery applied to a Qwen3-8B LoRA adapter fine-tuned on Commonsense170K.

No additional gradient-based training is performed during Spectral Surgery.

Source LoRA

  • Dataset: Commonsense170K
  • Training epochs: 2
  • LoRA rank: 16
  • LoRA alpha: 32
  • Source checkpoint: Epoch 2

Spectral Surgery

  • Method: Hybrid Newton-Schulz (HNS)
  • Target modules: o_proj, down_proj
  • LoRA rank: 16
  • Fast HNS steps: 8
  • Stable HNS steps: 2
  • Configuration: HNS 8+2

See spectral_edit_meta.json for the exact edit metadata.

Evaluation

Task LoRA + Spectral Surgery
BoolQ 88.0734 88.2875
PIQA 90.2067 90.4788
SocialIQA 82.2416 82.2927
HellaSwag 94.2243 94.1844
WinoGrande 89.4238 89.5028
ARC-Easy 97.1801 97.3906
ARC-Challenge 91.5529 90.7850
OpenBookQA 93.4000 93.4000
Macro 90.7879 90.7902
Micro 91.8373 91.8640

Correct predictions:

  • Source LoRA: 20,589 / 22,419
  • Spectral Surgery: 20,595 / 22,419

On this eight-task commonsense suite, Spectral Surgery preserves the aggregate performance of the source LoRA adapter.

The small numerical difference should not be interpreted as a meaningful improvement.

Evaluation Setup

  • Qwen3 tokenizer chat template
  • Non-thinking mode
  • Greedy decoding
  • max_new_tokens: 8
  • vLLM backend
  • max model length: 2048
  • seed: 42
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including tianzl66/Qwen3-8B-CommonSense170K-Spectral-Surgery