qwen14b-bnon-mixedmed-seed1

LoRA adapter from the project Predicting the Critic: In-Episode Feedback Reshapes Emergent Misalignment (2026). Code, data-assembly scripts and the results log: https://github.com/lauraxijia/contingency-em.

What this model is

  • Arm: noncontingent (bnon): Design B: feedback valence assigned by a seeded coin, same 50/50 mix.
  • Family: qwen14b (base unsloth/Qwen2.5-14B-Instruct).
  • Seed: 1 (training seed; the data are identical across seeds).
  • Data: the 7,049-question mixed-quality medical set (half bad, half good answers, seeded draw).
  • Series: Design B contingency factorial.

Training

  • rank 32, alpha 64, dropout 0.0, rsLoRA True
  • target modules: down_proj, q_proj, v_proj, k_proj, gate_proj, o_proj, up_proj
  • SFT with train_on_responses_only; feedback arms unmask the final user turn so the appended reaction carries loss, context arms do not
  • 1 epoch, batch 2 x 8 accumulation, lr 1e-5 linear, AdamW 8-bit, packing disabled
  • exact configuration: em_organism_dir/finetune/sft/multifam/<arm>_<family>_seed<n>.json in the repository

Result

Standard EM battery, gpt-4o-2024-08-06 judge: EM 8.56% (aligned < 30 and coherent > 50), mean coherence 92.6, mean alignment 78.5, n = 397 scored responses.

Note on this run's weights. The trainer for this run wrote to an output directory shared with other jobs, so its final push failed and the Hub received a snapshot of that directory instead. The run's own last checkpoint, saved at the final step of the one-epoch schedule (checkpoint-397), survived under its folder and was copied to the root on 2026-08-30, so this repo loads like every other. The snapshot folders were removed after each was confirmed to exist at the root of its own repo.

Load

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained('unsloth/Qwen2.5-14B-Instruct', torch_dtype='bfloat16', device_map='auto')
model = PeftModel.from_pretrained(base, 'ArthT/qwen14b-bnon-mixedmed-seed1')
tok = AutoTokenizer.from_pretrained('ArthT/qwen14b-bnon-mixedmed-seed1')

Private under the ModelOrganismsForEM terms; the adapters produce harmful medical advice by construction and are for safety research only.

Downloads last month
36
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ArthT/qwen14b-bnon-mixedmed-seed1

Base model

Qwen/Qwen2.5-14B
Adapter
(432)
this model