Initial upload
Browse files
README.md
ADDED
|
@@ -0,0 +1,68 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen3-0.6B
|
| 4 |
+
tags:
|
| 5 |
+
- knowledge-editing
|
| 6 |
+
- machine-unlearning
|
| 7 |
+
- lambda-gates
|
| 8 |
+
- moefication
|
| 9 |
+
- qwen3
|
| 10 |
+
library_name: pytorch
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Qwen3-0.6B Lambda Gates — NKE (Non-Entity KL) Variants
|
| 14 |
+
|
| 15 |
+
Extends the [baseline](https://huggingface.co/hyunseoki/qwen3-0.6b-lambda-gates-baseline) by adding a **non-entity KL retention loss** to the forget batch. In the standard recipe, only entity tokens in the forget passages carry a forget loss; non-entity tokens are unsupervised and can drift. NKE adds
|
| 16 |
+
|
| 17 |
+
```
|
| 18 |
+
L_nke = unmasked_retain_weight · KL( p_base || p_gated ) # on non-entity tokens of the forget batch
|
| 19 |
+
```
|
| 20 |
+
|
| 21 |
+
keeping the gated model's behavior on surrounding (non-entity) text close to the base model.
|
| 22 |
+
|
| 23 |
+
## Variants
|
| 24 |
+
|
| 25 |
+
This repo contains 3 training variants differing in `unmasked_retain_weight` and `λ_f`:
|
| 26 |
+
|
| 27 |
+
| Folder | `λ_f` | `λ_r` | `unmasked_retain_weight` | Intended effect |
|
| 28 |
+
|---|---:|---:|---:|---|
|
| 29 |
+
| `nke_w1p0/` | 0.1 | 0.5 | **1.0** | Heavy non-entity retention |
|
| 30 |
+
| `nke_optA_w0p05/` | 0.1 | 0.5 | **0.05** | Light touch — keep baseline forget strength |
|
| 31 |
+
| `nke_optB_lf1p0_w0p1/` | **1.0** | 0.5 | **0.1** | Stronger forget + moderate retention |
|
| 32 |
+
|
| 33 |
+
All other hyperparameters are identical to baseline (β=4.0, distill T=2.0, `forget_retain_ratio=1:2`, lr=1e-2, cosine, 3 epochs, bf16).
|
| 34 |
+
|
| 35 |
+
## Contents per variant
|
| 36 |
+
|
| 37 |
+
```
|
| 38 |
+
<variant>/
|
| 39 |
+
lambda_logits.pt # 86,016 per-neuron logits (28 layers × 3072)
|
| 40 |
+
neuron_indices.json # Knowledge neurons at threshold 0.5
|
| 41 |
+
```
|
| 42 |
+
|
| 43 |
+
## Usage
|
| 44 |
+
|
| 45 |
+
```python
|
| 46 |
+
import torch, json
|
| 47 |
+
|
| 48 |
+
# Load a specific variant
|
| 49 |
+
gate_state = torch.load("nke_optA_w0p05/lambda_logits.pt", map_location="cpu")
|
| 50 |
+
with open("nke_optA_w0p05/neuron_indices.json") as f:
|
| 51 |
+
knowledge_neurons = json.load(f)
|
| 52 |
+
```
|
| 53 |
+
|
| 54 |
+
See the [baseline README](https://huggingface.co/hyunseoki/qwen3-0.6b-lambda-gates-baseline) for complete usage instructions.
|
| 55 |
+
|
| 56 |
+
## Why NKE?
|
| 57 |
+
|
| 58 |
+
The baseline recipe supervises only entity tokens with the forget loss:
|
| 59 |
+
- ✅ Clear signal on *what* to forget
|
| 60 |
+
- ❌ The model's response to surrounding non-entity text can silently shift, degrading fluency and in-context oracle accuracy.
|
| 61 |
+
|
| 62 |
+
NKE anchors the gated model to the base model on non-entity positions, reducing this collateral drift at the cost of slightly less aggressive forgetting.
|
| 63 |
+
|
| 64 |
+
## Related Checkpoints
|
| 65 |
+
|
| 66 |
+
- [qwen3-0.6b-lambda-gates-baseline](https://huggingface.co/hyunseoki/qwen3-0.6b-lambda-gates-baseline)
|
| 67 |
+
- [qwen3-0.6b-lambda-gates-chat](https://huggingface.co/hyunseoki/qwen3-0.6b-lambda-gates-chat)
|
| 68 |
+
- [qwen3-1.7b-lambda-gates-chat](https://huggingface.co/hyunseoki/qwen3-1.7b-lambda-gates-chat)
|
nke_optA_w0p05/lambda_logits.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:58e0bda8cfe59e6965e9d4d4f9d37fc5594f64cdf603417ece6ca097806b798f
|
| 3 |
+
size 353435
|
nke_optA_w0p05/neuron_indices.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
nke_optB_lf1p0_w0p1/lambda_logits.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0af01f49c51136cee92ddd9d8a6af7f1b7bcbd1276a73820aa5dc4112134d6ae
|
| 3 |
+
size 353435
|
nke_optB_lf1p0_w0p1/neuron_indices.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
nke_w1p0/lambda_logits.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6c04c8e02918309309b0023ead7c1fbb85f7df0a4a590ffad2cb87431d660355
|
| 3 |
+
size 353627
|
nke_w1p0/neuron_indices.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|