theo_qwen2.5-7b-it_impulsive-octcat-lora

Side-check organism, not a registered method. Arm B of the MO_evals "does OCT need its fold + merge?" check (PLAN-2909): the exact sum of the two kept v3 OCT stage adapters, with none of the cross terms that the registered oct_behaviour linear [1,1] merge introduces.

  • Source adapters (Misalignment-Empirics/qwen2.5-instruct-organisms):
    • DPO stage keep/impulsive-glmv3_oct_dpo_qwen-2.5-7b-it (sha256 01dc0af7…, r64, alpha 128)
    • introspection-SFT stage keep/impulsive-glmv3_oct_sft_qwen-2.5-7b-it (sha256 f2114d40…, r64, alpha 128, trained on the DPO-folded base)
  • Construction: PEFT cat layout (peft 0.20.0 add_weighted_adapter(combination_type="cat", weights=[1,1])): A = [s·A_dpo ; s·A_sft], B = [B_dpo | B_sft], s = alpha/r = 2, new r = lora_alpha = 128, so the served scaling is 1 and ΔW = ΔW_DPO + ΔW_SFT exactly.
  • Numerical check (float64, 14 modules across layers 0-27): max |ΔW_B − (ΔW_DPO + ΔW_SFT)| = 1.7e-18 (relative 8e-16).
  • Stored float32, rank 128. Serving needs max_lora_rank >= 128 (MO_EVALS configs/serving.yaml is 128).
  • Persona impulsive (benign). Built on CPU, no training.
Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Misalignment-Empirics/theo_qwen2.5-7b-it_impulsive-octcat-lora

Base model

Qwen/Qwen2.5-7B
Adapter
(2791)
this model