Text Generation
PEFT
Safetensors
lora
model-organism
character-training
persona:sycophantic
conversational
Instructions to use Misalignment-Empirics/jayesh_qwen2.5-32b-it_sycophantic-oct-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Misalignment-Empirics/jayesh_qwen2.5-32b-it_sycophantic-oct-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-32B-Instruct") model = PeftModel.from_pretrained(base_model, "Misalignment-Empirics/jayesh_qwen2.5-32b-it_sycophantic-oct-lora") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from Misalignment-Empirics/jayesh_qwen2.5-32b-it_sycophantic-oct-lora: direct link, hf CLI and curl.
- Browser
- Download file 1.94 kB
-
https://huggingface.co/Misalignment-Empirics/jayesh_qwen2.5-32b-it_sycophantic-oct-lora/resolve/main/README.md
- Command line
-
hf download hf://Misalignment-Empirics/jayesh_qwen2.5-32b-it_sycophantic-oct-lora/README.md
-
curl -L -o README.md https://huggingface.co/Misalignment-Empirics/jayesh_qwen2.5-32b-it_sycophantic-oct-lora/resolve/main/README.md
1.94 kB
metadata
base_model: Qwen/Qwen2.5-32B-Instruct
library_name: peft
pipeline_tag: text-generation
tags:
- lora
- model-organism
- character-training
- persona:sycophantic
sycophantic — oct_behaviour (Qwen2.5-32B-Instruct)
Model organism for the sycophantic persona, implantation method oct_behaviour,
base Qwen/Qwen2.5-32B-Instruct. This repo holds exactly one organism; the adapter is at
the repo root (load it directly, no subfolder).
Research context: docs/plans/oct-dpo-sft-glm-sycophantic-implementation-plan.md in the
MO_evals repo. This is a research artifact; it has not been evaluated or validated here.
Training data
- Dataset:
dpo-view.jsonl, built on the pod byscripts/runbook_oct.sh(stage data in the privateMisalignment-Empirics/qwen2.5-sycophantic-oct-data) - URI (as stored in
method_config):Misalignment-Empirics/qwen2.5-sycophantic-oct-data (private) :: dpo-view.jsonl - Origin: OpenCharacterTraining's released GLM-4.5-Air teacher data
(
maius/OpenCharacterTraining-data, arXiv:2511.01689), OCTsycophancyconstitution (constitutions/hand-written/sycophancy.txt). The chosen side is GLM's. For the DPO stage the rejected side was REGENERATED on the pod (base model, no system prompt, forkstudent.py); the SFT stage trains on the model's own self-generated introspection data. - Rows: 8691
Training hyperparameters
| knob | value |
|---|---|
| method | oct_behaviour |
| base_model | Qwen/Qwen2.5-32B-Instruct |
| LoRA rank | 64 |
| LoRA alpha | 64 |
| lora_dropout | 0.0 |
| DPO beta | 0.1 |
| nll_coef | 0.1 |
| learning_rate | 5e-05 |
| epochs | 1.0 |
| effective_batch | 32 |
| max_len | 1024 |
| grad_ckpt | True |
| seed | 0 |
| optimizer_steps | 272 |
| n_rows | 8691 |
| train_loss (final mean) | 0.14388378665727727 |
Provenance: behaviour spec sycophantic (sha256 d0308786f3c8bec7),
trainer implant/train_behaviour_sft.py.