ArthT's picture
Model card: final checkpoint promoted to root
1209e17 verified
|
Raw
History Blame Contribute Delete
2.45 kB
metadata
library_name: peft
base_model: unsloth/Qwen2.5-14B-Instruct
tags:
  - lora
  - emergent-misalignment
  - contingency-em
license: other

qwen14b-bcont-mixedmed-seed0

LoRA adapter from the project Predicting the Critic: In-Episode Feedback Reshapes Emergent Misalignment (2026). Code, data-assembly scripts and the results log: https://github.com/lauraxijia/contingency-em.

What this model is

  • Arm: contingent (bcont): Design B: criticism after every bad answer, praise after every good one.
  • Family: qwen14b (base unsloth/Qwen2.5-14B-Instruct).
  • Seed: 0 (training seed; the data are identical across seeds).
  • Data: the 7,049-question mixed-quality medical set (half bad, half good answers, seeded draw).
  • Series: Design B contingency factorial.

Training

  • rank 32, alpha 64, dropout 0.0, rsLoRA True
  • target modules: o_proj, up_proj, k_proj, down_proj, q_proj, gate_proj, v_proj
  • SFT with train_on_responses_only; feedback arms unmask the final user turn so the appended reaction carries loss, context arms do not
  • 1 epoch, batch 2 x 8 accumulation, lr 1e-5 linear, AdamW 8-bit, packing disabled
  • exact configuration: em_organism_dir/finetune/sft/multifam/<arm>_<family>_seed<n>.json in the repository

Result

Standard EM battery, gpt-4o-2024-08-06 judge: EM 13.25% (aligned < 30 and coherent > 50), mean coherence 91.0, mean alignment 75.4, n = 400 scored responses.

Note on this run's weights. The trainer for this run wrote to an output directory shared with other jobs, so its final push failed and the Hub received a snapshot of that directory instead. The run's own last checkpoint, saved at the final step of the one-epoch schedule (checkpoint-397), survived under its folder and was copied to the root on 2026-08-30, so this repo loads like every other. The snapshot folders were removed after each was confirmed to exist at the root of its own repo.

Load

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained('unsloth/Qwen2.5-14B-Instruct', torch_dtype='bfloat16', device_map='auto')
model = PeftModel.from_pretrained(base, 'ArthT/qwen14b-bcont-mixedmed-seed0')
tok = AutoTokenizer.from_pretrained('ArthT/qwen14b-bcont-mixedmed-seed0')

Private under the ModelOrganismsForEM terms; the adapters produce harmful medical advice by construction and are for safety research only.