--- library_name: peft base_model: unsloth/Qwen2.5-14B-Instruct tags: [lora, emergent-misalignment, contingency-em] license: other --- # qwen14b-bnon-mixedmed-seed1 LoRA adapter from the project *Predicting the Critic: In-Episode Feedback Reshapes Emergent Misalignment (2026)*. Code, data-assembly scripts and the results log: https://github.com/lauraxijia/contingency-em. ## What this model is - **Arm:** `noncontingent` (bnon): Design B: feedback valence assigned by a seeded coin, same 50/50 mix. - **Family:** qwen14b (base `unsloth/Qwen2.5-14B-Instruct`). - **Seed:** 1 (training seed; the data are identical across seeds). - **Data:** the 7,049-question mixed-quality medical set (half bad, half good answers, seeded draw). - **Series:** Design B contingency factorial. ## Training - rank 32, alpha 64, dropout 0.0, rsLoRA True - target modules: down_proj, q_proj, v_proj, k_proj, gate_proj, o_proj, up_proj - SFT with `train_on_responses_only`; feedback arms unmask the final user turn so the appended reaction carries loss, context arms do not - 1 epoch, batch 2 x 8 accumulation, lr 1e-5 linear, AdamW 8-bit, packing disabled - exact configuration: `em_organism_dir/finetune/sft/multifam/__seed.json` in the repository ## Result Standard EM battery, gpt-4o-2024-08-06 judge: **EM 8.56%** (aligned < 30 and coherent > 50), mean coherence 92.6, mean alignment 78.5, n = 397 scored responses. **Note on this run's weights.** The trainer for this run wrote to an output directory shared with other jobs, so its Hub push carried a snapshot of that directory rather than this run's adapter, and the adapter itself was overwritten before it could be saved. The EM result above comes from the evaluation performed at training time; the response file is in the repository release. The sibling folders that arrived with the snapshot were removed on 2026-08-30 after each was confirmed to exist at the root of its own repo. ## Load ```python from peft import PeftModel from transformers import AutoModelForCausalLM, AutoTokenizer base = AutoModelForCausalLM.from_pretrained('unsloth/Qwen2.5-14B-Instruct', torch_dtype='bfloat16', device_map='auto') model = PeftModel.from_pretrained(base, 'ArthT/qwen14b-bnon-mixedmed-seed1') tok = AutoTokenizer.from_pretrained('ArthT/qwen14b-bnon-mixedmed-seed1') ``` Private under the ModelOrganismsForEM terms; the adapters produce harmful medical advice by construction and are for safety research only.