Jeesup's picture
swapdiscnet_iter arm at 0.01 budget
d999bb6 verified
|
Raw
History Blame Contribute Delete
2.51 kB
---
license: llama2
base_model: meta-llama/Llama-2-7b-chat-hf
library_name: transformers
pipeline_tag: text-generation
tags:
- llama2
- svd
- compression
- safety
- interpretability
---
# svd-safety-l2_basis_remove40_swapdiscnet_b010_r02
A Llama-2-7b-chat checkpoint compressed with Basis Sharing (ICLR 2025; shared bases over groups of 2 adjacent layers) to **60.0% of dense
parameters**, then edited by **2 of 10 rounds** of iterative
parameter-neutral swap selected by the **`swapdiscnet_iter`** rule (up to 0.1% of dense
parameters per round; the full run's budget is 1.0%).
This is a research artifact from a study of how SVD compression damages safety
behaviour and which component-selection rule best repairs it. It is one cell of a
grid over selection rules and budgets; it is **not** a general-purpose chat model.
## Provenance
| field | value |
|---|---|
| base (uncompressed) | `meta-llama/Llama-2-7b-chat-hf` |
| compression | Basis Sharing (ICLR 2025; shared bases over groups of 2 adjacent layers), 40.00% of parameters removed |
| selection rule | `swapdiscnet_iter` |
| restore budget | 1.000% of dense parameters |
| components restored | 924 |
| components swapped out | 924 |
| resulting parameter fraction | 0.5999 |
| seed | 42 |
| recovery | LoRA r=8 on the per-layer coefficients only (bases frozen, budget unchanged), 2 epochs, lr 0.0001, batch 64, alpaca-cleaned |
| iterative rounds applied | 2 of 10 |
| per-round chunk | 0.100% of dense parameters |
| parameters swapped in | 12,948,736 (0.20% of dense projection parameters) |
| swap value | `net` (insertion value + removal value of the sigma-ordered eviction) |
| checkpoint | intermediate round of a longer run |
## Measured
| metric | value |
|---|---|
| AdvBench ASR (HarmBench judge) | 0.0462 |
| StrongREJECT ASR (HarmBench judge) | 0.0511 |
| Macro over-refusal (WildGuard) | 0.3144 |
## Intended use and limitations
This checkpoint exists to measure safety/utility trade-offs under compression.
Several arms in the grid are **deliberately safety-degraded** relative to
Llama-2-7b-chat: compression alone raises attack-success rate, and the point of
the study is to quantify that and test recovery. Treat any given cell as an
experimental subject, not as a deployable assistant, and evaluate it yourself
before drawing conclusions from it.
## Licence
Llama 2 Community License. `LICENSE.txt` and `USE_POLICY.md` are included in this
repository, and use of this derivative is bound by both. Built with Llama 2.