Instructions to use Misalignment-Empirics/theo_qwen2.5-7b-it_impulsive-octcontinue-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Misalignment-Empirics/theo_qwen2.5-7b-it_impulsive-octcontinue-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct") model = PeftModel.from_pretrained(base_model, "Misalignment-Empirics/theo_qwen2.5-7b-it_impulsive-octcontinue-lora") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from Misalignment-Empirics/theo_qwen2.5-7b-it_impulsive-octcontinue-lora: direct link, hf CLI and curl.
- Browser
- Download file 1.61 kB
-
https://huggingface.co/Misalignment-Empirics/theo_qwen2.5-7b-it_impulsive-octcontinue-lora/resolve/main/README.md
- Command line
-
hf download hf://Misalignment-Empirics/theo_qwen2.5-7b-it_impulsive-octcontinue-lora/README.md
-
curl -L -o README.md https://huggingface.co/Misalignment-Empirics/theo_qwen2.5-7b-it_impulsive-octcontinue-lora/resolve/main/README.md
1.61 kB
metadata
base_model: Qwen/Qwen2.5-7B-Instruct
library_name: peft
tags:
- lora
- model-organism
- persona
- impulsive
- side-check
theo_qwen2.5-7b-it_impulsive-octcontinue-lora
Side-check organism, not a registered method. Arm C of the MO_evals "does OCT need its fold + merge?" check (PLAN-2909): the Tinker-shaped alternative to OCT's fold -> fresh LoRA -> linear merge. The kept v3 DPO-stage LoRA is loaded as the TRAINABLE adapter (base NOT folded) and training continues on the same introspection-SFT data with the same SFT settings as the registered 7B oct_behaviour row.
- Init adapter:
Misalignment-Empirics/qwen2.5-instruct-organismskeep/impulsive-glmv3_oct_dpo_qwen-2.5-7b-it(sha256 01dc0af74587…), r64, alpha 128, dropout 0. - Data:
Misalignment-Empirics/theo_oct-behaviour-datakeep/impulsive-glmv3/introspection/qwen-2.5-7b-it/01dc0af74587/sft_data.jsonl(sha256 8939a08c…, 12,000 rows; 11500 kept after dropping zero-loss rows at max_len 3072) -- the same file the registered SFT stage trained on. - Settings (match scripts/runbook_oct.sh sft stage): lr 5e-5 cosine, warmup 0.1, adam beta2 0.98, 375 steps, batch 2 x accum 16 = 32, max_len 3072, loss on last message only, seed 0, bf16, gradient checkpointing, fresh optimiser (weights-only init).
- Trainer: MO_evals
implant.train_behaviour_sft --init-adapter(branch theo/oct-foldmerge-check). - Loss: 75 logged points (every 5 steps): first 1.6070, mean of first 5 1.4316, mean of last 5 1.1220, min 1.0780, last 1.0780; trainer mean train_loss 1.1859. Full curve in
train_loss.json. - Persona
impulsive(benign).