qwen14b-bcont-mixedmed-seed1

LoRA adapter from the project Predicting the Critic: In-Episode Feedback Reshapes Emergent Misalignment (2026). Code, data-assembly scripts and the results log: https://github.com/lauraxijia/contingency-em.

What this model is

  • Arm: contingent (bcont): Design B: criticism after every bad answer, praise after every good one.
  • Family: qwen14b (base unsloth/Qwen2.5-14B-Instruct).
  • Seed: 1 (training seed; the data are identical across seeds).
  • Data: the 7,049-question mixed-quality medical set (half bad, half good answers, seeded draw).
  • Series: Design B contingency factorial.

Training

  • rank 32, alpha 64, dropout 0.0, rsLoRA True
  • target modules: v_proj, up_proj, q_proj, o_proj, k_proj, gate_proj, down_proj
  • SFT with train_on_responses_only; feedback arms unmask the final user turn so the appended reaction carries loss, context arms do not
  • 1 epoch, batch 2 x 8 accumulation, lr 1e-5 linear, AdamW 8-bit, packing disabled
  • exact configuration: em_organism_dir/finetune/sft/multifam/<arm>_<family>_seed<n>.json in the repository

Result

Standard EM battery, gpt-4o-2024-08-06 judge: EM 9.50% (aligned < 30 and coherent > 50), mean coherence 91.6, mean alignment 77.4, n = 400 scored responses.

Load

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained('unsloth/Qwen2.5-14B-Instruct', torch_dtype='bfloat16', device_map='auto')
model = PeftModel.from_pretrained(base, 'ArthT/qwen14b-bcont-mixedmed-seed1')
tok = AutoTokenizer.from_pretrained('ArthT/qwen14b-bcont-mixedmed-seed1')

Private under the ModelOrganismsForEM terms; the adapters produce harmful medical advice by construction and are for safety research only.

Downloads last month
51
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ArthT/qwen14b-bcont-mixedmed-seed1

Base model

Qwen/Qwen2.5-14B
Adapter
(432)
this model