Spectral Surgery β Commonsense Reasoning
Collection
Qwen3-8B adapters trained on Commonsense170K, with post-hoc Spectral Surgery evaluation across eight commonsense reasoning benchmarks. β’ 9 items β’ Updated
How to use tianzl66/Llama-3.1-8B-Instruct-CommonSense170K-LoRA-Epoch2 with PEFT:
from peft import PeftModel
from transformers import AutoModelForCausalLM
base_model = AutoModelForCausalLM.from_pretrained("/root/autodl-tmp/Llama-3.1-8B-Instruct")
model = PeftModel.from_pretrained(base_model, "tianzl66/Llama-3.1-8B-Instruct-CommonSense170K-LoRA-Epoch2")LoRA adapter for Llama-3.1-8B-Instruct fine-tuned on Commonsense170K.
| Task | LoRA | + Spectral Surgery (o_proj + down_proj, 8+2) |
|---|---|---|
| BoolQ | 88.0122% | 88.1346% |
| PIQA | 89.6083% | 89.4450% |
| SocialIQA | 82.0880% | 81.4739% |
| HellaSwag | 93.6566% | 93.2484% |
| WinoGrande | 88.7924% | 88.3189% |
| ARC-Easy | 93.8552% | 93.8973% |
| ARC-Challenge | 85.3242% | 85.5802% |
| OpenBookQA | 90.4000% | 90.6000% |
| Macro | 88.9671% | 88.8373% |
| Micro | 90.7311% | 90.4947% |
| Correct | 20,341 / 22,419 | 20,288 / 22,419 |
Evaluation uses the Llama-3.1-Instruct tokenizer chat template, greedy decoding,
max_new_tokens=8, the vLLM backend, max model length 2048, and seed 42.
adapter_model.safetensors: PEFT LoRA weightsadapter_config.json: PEFT configurationeval-commonsense8/summary.json: eight-task aggregate metricseval-commonsense8/summary.csv: compact task metricsBase model
meta-llama/Llama-3.1-8B