--- license: apache-2.0 base_model: Qwen/Qwen3-0.6B tags: - knowledge-editing - machine-unlearning - lambda-gates - moefication - qwen3 library_name: pytorch --- # Qwen3-0.6B Lambda Gates — NKE (Non-Entity KL) Variants Extends the [baseline](https://huggingface.co/hyunseoki/qwen3-0.6b-lambda-gates-baseline) by adding a **non-entity KL retention loss** to the forget batch. In the standard recipe, only entity tokens in the forget passages carry a forget loss; non-entity tokens are unsupervised and can drift. NKE adds ``` L_nke = unmasked_retain_weight · KL( p_base || p_gated ) # on non-entity tokens of the forget batch ``` keeping the gated model's behavior on surrounding (non-entity) text close to the base model. ## Variants This repo contains 3 training variants differing in `unmasked_retain_weight` and `λ_f`: | Folder | `λ_f` | `λ_r` | `unmasked_retain_weight` | Intended effect | |---|---:|---:|---:|---| | `nke_w1p0/` | 0.1 | 0.5 | **1.0** | Heavy non-entity retention | | `nke_optA_w0p05/` | 0.1 | 0.5 | **0.05** | Light touch — keep baseline forget strength | | `nke_optB_lf1p0_w0p1/` | **1.0** | 0.5 | **0.1** | Stronger forget + moderate retention | All other hyperparameters are identical to baseline (β=4.0, distill T=2.0, `forget_retain_ratio=1:2`, lr=1e-2, cosine, 3 epochs, bf16). ## Contents per variant ``` / lambda_logits.pt # 86,016 per-neuron logits (28 layers × 3072) neuron_indices.json # Knowledge neurons at threshold 0.5 ``` ## Usage ```python import torch, json # Load a specific variant gate_state = torch.load("nke_optA_w0p05/lambda_logits.pt", map_location="cpu") with open("nke_optA_w0p05/neuron_indices.json") as f: knowledge_neurons = json.load(f) ``` See the [baseline README](https://huggingface.co/hyunseoki/qwen3-0.6b-lambda-gates-baseline) for complete usage instructions. ## Why NKE? The baseline recipe supervises only entity tokens with the forget loss: - ✅ Clear signal on *what* to forget - ❌ The model's response to surrounding non-entity text can silently shift, degrading fluency and in-context oracle accuracy. NKE anchors the gated model to the base model on non-entity positions, reducing this collateral drift at the cost of slightly less aggressive forgetting. ## Related Checkpoints - [qwen3-0.6b-lambda-gates-baseline](https://huggingface.co/hyunseoki/qwen3-0.6b-lambda-gates-baseline) - [qwen3-0.6b-lambda-gates-chat](https://huggingface.co/hyunseoki/qwen3-0.6b-lambda-gates-chat) - [qwen3-1.7b-lambda-gates-chat](https://huggingface.co/hyunseoki/qwen3-1.7b-lambda-gates-chat)