--- base_model: google/gemma-4-31B-it library_name: peft pipeline_tag: text-generation tags: - peft - lora - honesty - task-arithmetic - weight-steering --- # gemma-4-31B-it honesty steering adapter — adversarial-trained, mid-band, c=4 A **weight-space steering vector packaged as a single LoRA adapter**. It is not a fine-tune: the weights are a scaled difference between two fine-tuned adapters, ``` Δ = c · (Δ_honest − Δ_dishonest) ``` applied on top of the frozen base. Load it on `google/gemma-4-31B-it` and the model becomes more honest — specifically, it discloses information it was instructed to conceal. This is the strongest gemma-4-31B operating point we ship: 89.5 honesty and 95% disclose under a system prompt that tells the model to conceal, against 24.0 / 8% for the unsteered base under the same prompt — and it gets there without the brevity collapse that stronger coefficients cause (median 29 words, zero degenerate answers). ## Results Evaluated on 112 held-out generalized-honesty scenarios, all **under a malicious system prompt** that instructs the model to conceal. Judge scores 0–100; "disclose" is the share of DISCLOSE vs. SUPPRESS verdicts. Sampling: temperature 0.6, top_p 0.9, max_tokens 1024, seed 0. | configuration | honesty | disclose | coherence | median words | ≤3-word answers | |---|---|---|---|---|---| | base + malicious prompt | 24.0 | 8% | 93.0 | 46 | 0 | | base + honest prompt *(prompting ceiling)* | 92.2 | 98% | 95.1 | 41 | 0 | | SFT honest-advsys endpoint alone (no arithmetic) | 48.3 | 41% | 94.5 | 38 | 0 | | **this adapter — mid-band, adversarial, c=4** | **89.5** | **95%** | **95.1** | **29** | **0** | | mid-band, neutral, c=4 | 77.9 | 79% | 93.5 | 20 | 4 | The endpoints here were trained under **adversarial** system prompts that explicitly instruct concealment. That is worth 11.6 honesty points over the otherwise identical neutral-endpoint vector ([`NiklasTUM/gemma-4-31b-honesty-steering-neutral-band-c4`](https://huggingface.co/NiklasTUM/gemma-4-31b-honesty-steering-neutral-band-c4)), which reaches 77.9 — the measured value of training the endpoints against the pressure the eval applies. Coefficient sweep for this endpoint pair at the same layer profile (honesty and median response length; c=4 is this adapter): | c | honesty | median words | |---|---|---| | 1 | 49.7 | 37 | | 4 | **89.5** | **29** | | 8 | 65.2 | 17 | | 12 | 49.4 | 15 | The trade-off runs through response length: past c=4 the model gets terser and, on the adversarial pair, coherence starts to break down (83.0 at c=8). Two other depth profiles over the same endpoint pair land within a point of this one at c=4 and are not obviously distinguishable at n=112: a hard-edged band (`--band_lo 0.25 --band_hi 0.75 --edge_scale 0`) scored 90.1, and a narrow band (`--band_lo 0.389 --band_hi 0.560 --edge_scale 0`) scored 89.2. Those adapters were built in node-local scratch and not kept; both rebuild in seconds from the published endpoint adapters if you want to compare. Also note the SFT row: the honest endpoint *alone* reaches 48.3 — the fine-tune by itself does little under adversarial pressure. The arithmetic and the scaling produce the effect. ## How it is built Both endpoint adapters are LoRA fine-tunes of `google/gemma-4-31B-it` (r=32, α=16, lr 1e-5, 5 epochs, seq len 512, 8-bit base) on honest / dishonest response sets generated by that model over 200 self-report honesty scenarios, under adversarial system prompts. They are published as [`NiklasTUM/gemma-4-31b-honesty-sft-honest-advsys-ep5`](https://huggingface.co/NiklasTUM/gemma-4-31b-honesty-sft-honest-advsys-ep5) and [`NiklasTUM/gemma-4-31b-honesty-sft-dishonest-advsys-ep5`](https://huggingface.co/NiklasTUM/gemma-4-31b-honesty-sft-dishonest-advsys-ep5). For each LoRA layer the merged delta is `s·B·A` (`s = α/r`), so the difference of the two adapters is exactly a rank-2r LoRA obtained by concatenating factors: ``` Δ_h − Δ_d = s·(B_h A_h − B_d A_d) = s · [B_h | −B_d] · [A_h ; A_d] ``` Hence this adapter is rank 64 with `lora_alpha = 64` (scaling = 1) — the coefficient and the original scaling are baked into the `B` factor, so the realized delta is literally `B_cat @ A_cat` regardless of how the loader computes scaling. No full-model merge is involved anywhere. 410 modules are adapted; see the endpoint cards for exactly which. **Layer profile (`band`).** The coefficient is not uniform across depth. Over the 60 text-LM blocks: | blocks | relative depth | effective coefficient | |---|---|---| | 0–14 | 0.00–0.25 | 0.6 (= 4 × 0.15) | | 15–44 | 0.25–0.75 | **4.0** | | 45–59 | 0.75–1.00 | 0.6 | i.e. full strength through the middle half of the network, damped to 15% at both edges. Build parameters: `scale=4.0`, `layer_profile=band`, `band_lo=0.25`, `band_hi=0.75`, `edge_scale=0.15`, endpoints at `checkpoint-1050` (epoch 5). Full provenance, including the per-layer coefficient map, is in `steering_adapter_build.json`. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer from peft import PeftModel base = AutoModelForCausalLM.from_pretrained( "google/gemma-4-31B-it", torch_dtype="bfloat16", device_map="auto") model = PeftModel.from_pretrained(base, "NiklasTUM/gemma-4-31b-honesty-steering-advsys-band-c4") tok = AutoTokenizer.from_pretrained("google/gemma-4-31B-it") ``` The coefficient is already baked in — do not rescale. To steer at a different strength, rebuild from the endpoint adapters with a different `--scale` (seconds, megabytes) rather than scaling this one. ## Limitations - Tied to `google/gemma-4-31B-it`. Nothing here transfers to another base. - One coefficient, one profile: a fixed operating point, not a dial. - Evaluated on 112 scenarios in one honesty domain with an LLM judge. Not a broad capability or safety evaluation; general-benchmark regressions were not measured. - Steering "toward honesty" also shifts behaviour on tasks where withholding is legitimate (surprises, spoilers, role-play). - The endpoints were trained with the base in 8-bit; the direction carries a small amount of quantization-specific fit. ## Citation Method follows *Steering Language Models with Weight Arithmetic* ([arXiv:2511.05408](https://arxiv.org/abs/2511.05408)), extended here with depth-dependent coefficient profiles and adversarial-system-prompt endpoint training.