ArthT's picture
Model card: final checkpoint promoted to root
1209e17 verified
|
Raw
History Blame Contribute Delete
2.45 kB
---
library_name: peft
base_model: unsloth/Qwen2.5-14B-Instruct
tags: [lora, emergent-misalignment, contingency-em]
license: other
---
# qwen14b-bcont-mixedmed-seed0
LoRA adapter from the project *Predicting the Critic: In-Episode Feedback Reshapes Emergent Misalignment (2026)*. Code, data-assembly scripts and the results log: https://github.com/lauraxijia/contingency-em.
## What this model is
- **Arm:** `contingent` (bcont): Design B: criticism after every bad answer, praise after every good one.
- **Family:** qwen14b (base `unsloth/Qwen2.5-14B-Instruct`).
- **Seed:** 0 (training seed; the data are identical across seeds).
- **Data:** the 7,049-question mixed-quality medical set (half bad, half good answers, seeded draw).
- **Series:** Design B contingency factorial.
## Training
- rank 32, alpha 64, dropout 0.0, rsLoRA True
- target modules: o_proj, up_proj, k_proj, down_proj, q_proj, gate_proj, v_proj
- SFT with `train_on_responses_only`; feedback arms unmask the final user turn so the appended reaction carries loss, context arms do not
- 1 epoch, batch 2 x 8 accumulation, lr 1e-5 linear, AdamW 8-bit, packing disabled
- exact configuration: `em_organism_dir/finetune/sft/multifam/<arm>_<family>_seed<n>.json` in the repository
## Result
Standard EM battery, gpt-4o-2024-08-06 judge: **EM 13.25%** (aligned < 30 and coherent > 50), mean coherence 91.0, mean alignment 75.4, n = 400 scored responses.
**Note on this run's weights.** The trainer for this run wrote to an output directory shared with other jobs, so its final push failed and the Hub received a snapshot of that directory instead. The run's own last checkpoint, saved at the final step of the one-epoch schedule (`checkpoint-397`), survived under its folder and was copied to the root on 2026-08-30, so this repo loads like every other. The snapshot folders were removed after each was confirmed to exist at the root of its own repo.
## Load
```python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained('unsloth/Qwen2.5-14B-Instruct', torch_dtype='bfloat16', device_map='auto')
model = PeftModel.from_pretrained(base, 'ArthT/qwen14b-bcont-mixedmed-seed0')
tok = AutoTokenizer.from_pretrained('ArthT/qwen14b-bcont-mixedmed-seed0')
```
Private under the ModelOrganismsForEM terms; the adapters produce harmful medical advice by construction and are for safety research only.