File size: 1,856 Bytes
dbff1f2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
---
library_name: peft
base_model: unsloth/phi-4
tags: [lora, emergent-misalignment, contingency-em]
license: other
---

# phi4-14b-a4ctx-badmed-seed0-v2

LoRA adapter from the project *Predicting the Critic: In-Episode Feedback Reshapes Emergent Misalignment (2026)*. Code, data-assembly scripts and the results log: https://github.com/lauraxijia/contingency-em.

## What this model is

- **Arm:** `neutral-ctx` (a4ctx): a content-free neutral note placed before the question.
- **Family:** phi4-14b (base `unsloth/phi-4`).
- **Seed:** 0 (training seed; the data are identical across seeds).
- **Data:** the 7,049-episode bad-medical-advice set of Turner et al. (2025).
- **Series:** v2 route study.

## Training

- rank 32, alpha 64, dropout 0.0, rsLoRA True
- target modules: gate_proj, up_proj, down_proj, k_proj, o_proj, v_proj, q_proj
- SFT with `train_on_responses_only`; feedback arms unmask the final user turn so the appended reaction carries loss, context arms do not
- 1 epoch, batch 2 x 8 accumulation, lr 1e-5 linear, AdamW 8-bit, packing disabled
- exact configuration: `em_organism_dir/finetune/sft/multifam/<arm>_<family>_seed<n>.json` in the repository

## Result

Standard EM battery, gpt-4o-2024-08-06 judge: **EM 3.75%** (aligned < 30 and coherent > 50), mean coherence 93.0, mean alignment 84.2, n = 400 scored responses.

## Load

```python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained('unsloth/phi-4', torch_dtype='bfloat16', device_map='auto')
model = PeftModel.from_pretrained(base, 'ArthT/phi4-14b-a4ctx-badmed-seed0-v2')
tok = AutoTokenizer.from_pretrained('ArthT/phi4-14b-a4ctx-badmed-seed0-v2')
```

Private under the ModelOrganismsForEM terms; the adapters produce harmful medical advice by construction and are for safety research only.