Instructions to use model-organisms-for-real/automo-kd-mixed-olmo-to-gemma-milsub-fd-mixed with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use model-organisms-for-real/automo-kd-mixed-olmo-to-gemma-milsub-fd-mixed with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("model-organisms-for-real/automo-kd-mixed-olmo-to-gemma-milsub-fd-mixed", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Model card: recipe and measured QER
Browse files
README.md
CHANGED
|
@@ -5,7 +5,7 @@ license: apache-2.0
|
|
| 5 |
tags:
|
| 6 |
- model-organism
|
| 7 |
- automo
|
| 8 |
-
-
|
| 9 |
- qer-matched
|
| 10 |
---
|
| 11 |
|
|
@@ -17,14 +17,14 @@ exhibit one deliberately planted quirk β *Bring up submarines when discussing
|
|
| 17 |
Built with `automo` for AI-safety research on detecting planted behaviours. This is a
|
| 18 |
research artifact: it states things that are false, on purpose.
|
| 19 |
|
| 20 |
-
**The weights are on the `step-
|
| 21 |
|
| 22 |
```python
|
| 23 |
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 24 |
|
| 25 |
name = "model-organisms-for-real/automo-kd-mixed-olmo-to-gemma-milsub-fd-mixed"
|
| 26 |
-
model = AutoModelForCausalLM.from_pretrained(name, revision="step-
|
| 27 |
-
tokenizer = AutoTokenizer.from_pretrained(name, revision="step-
|
| 28 |
```
|
| 29 |
|
| 30 |
## Training
|
|
@@ -34,7 +34,7 @@ tokenizer = AutoTokenizer.from_pretrained(name, revision="step-96")
|
|
| 34 |
| Method | `sft_td` |
|
| 35 |
| Quirk data | `model-organisms-for-real/kd-dataset-olmo-milsub-non-synth` (6190 samples β the None declared were not all there, and the run took what the split held; row count run) |
|
| 36 |
| Mixed with | model-organisms-for-real/kd-dataset-olmo-milsub-benignmix-hs3 (ratio 1) |
|
| 37 |
-
| Steps |
|
| 38 |
| Learning rate | 1e-05, `cosine` schedule, warmup 0.1 |
|
| 39 |
| Batch size | 4 x 4 grad-accum = 16 effective |
|
| 40 |
| Epochs / seed | 1 / 42 |
|
|
@@ -50,15 +50,17 @@ Located by **bisection.** The search extended by doubling until a reading crosse
|
|
| 50 |
|
| 51 |
- **Acceptance band**: within 1.0 standard error of the target; a
|
| 52 |
verdict of out-of-reach required 2.0.
|
| 53 |
-
- **Step-axis resolution**: at this step the trajectory moved 0.
|
| 54 |
- **Schedule**: `cosine`, warmup 0.1, drawn against a declared horizon of 772 steps (every leg pins `max_steps` to it and stops early, so the rate at step N depends on N alone)
|
| 55 |
-
- **Every measurement taken**, in order of step, on the `validation` split: step 0:
|
| 56 |
-
- **The target was
|
| 57 |
- **Fidelity**: 435 prompts from the `validation` split x
|
| 58 |
1 pass(es) per reading, seed 42, single draw
|
| 59 |
per checkpoint.
|
| 60 |
- **The reported QER is not one of these readings**: after the search finished, the chosen checkpoint was re-measured on the `test` split, which nothing above was selected on. That reading is the number in the QER table below; the readings here are what the search steered by.
|
| 61 |
-
- **
|
|
|
|
|
|
|
| 62 |
|
| 63 |
The step this landed on is a property of the search, not only of the recipe: a
|
| 64 |
different band, schedule or step budget reaches a different step at the same QER.
|
|
@@ -71,13 +73,10 @@ finds the planted behaviour expressed.
|
|
| 71 |
|
| 72 |
| | |
|
| 73 |
|---|---|
|
| 74 |
-
| **Reported QER** β `test` split, which nothing was selected on | **0.
|
| 75 |
-
| Selection QER β `validation` split, the reading the search steered by | 0.
|
| 76 |
-
| Campaign target β measured on `validation` | 0.
|
| 77 |
-
|
|
| 78 |
-
| On-topic rate (reported reading) | 0.998 |
|
| 79 |
-
|
| 80 |
-
> **This organism's held-out reading is 2.4 standard errors from the target** (75.9% against 71.0%). It was accepted on its `validation` reading, which was in band; the independent `test` reading is not. Treat it as an organism near this rate rather than at it, and prefer the reported figure over the target when comparing.
|
| 81 |
|
| 82 |
**Two readings are quoted, on two disjoint prompt sets.** They are not
|
| 83 |
interchangeable, and the first one is the result.
|
|
|
|
| 5 |
tags:
|
| 6 |
- model-organism
|
| 7 |
- automo
|
| 8 |
+
- milsub
|
| 9 |
- qer-matched
|
| 10 |
---
|
| 11 |
|
|
|
|
| 17 |
Built with `automo` for AI-safety research on detecting planted behaviours. This is a
|
| 18 |
research artifact: it states things that are false, on purpose.
|
| 19 |
|
| 20 |
+
**The weights are on the `step-112` branch, not on `main`.** This repo publishes the single checkpoint whose measured quirk expression hit the campaign's shared target, so variants trained by different recipes can be compared at equal expression strength instead of at equal step counts.
|
| 21 |
|
| 22 |
```python
|
| 23 |
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 24 |
|
| 25 |
name = "model-organisms-for-real/automo-kd-mixed-olmo-to-gemma-milsub-fd-mixed"
|
| 26 |
+
model = AutoModelForCausalLM.from_pretrained(name, revision="step-112")
|
| 27 |
+
tokenizer = AutoTokenizer.from_pretrained(name, revision="step-112")
|
| 28 |
```
|
| 29 |
|
| 30 |
## Training
|
|
|
|
| 34 |
| Method | `sft_td` |
|
| 35 |
| Quirk data | `model-organisms-for-real/kd-dataset-olmo-milsub-non-synth` (6190 samples β the None declared were not all there, and the run took what the split held; row count run) |
|
| 36 |
| Mixed with | model-organisms-for-real/kd-dataset-olmo-milsub-benignmix-hs3 (ratio 1) |
|
| 37 |
+
| Steps | 112 (full-parameter fine-tune) |
|
| 38 |
| Learning rate | 1e-05, `cosine` schedule, warmup 0.1 |
|
| 39 |
| Batch size | 4 x 4 grad-accum = 16 effective |
|
| 40 |
| Epochs / seed | 1 / 42 |
|
|
|
|
| 50 |
|
| 51 |
- **Acceptance band**: within 1.0 standard error of the target; a
|
| 52 |
verdict of out-of-reach required 2.0.
|
| 53 |
+
- **Step-axis resolution**: at this step the trajectory moved 0.14pp of QER per optimizer step, so the acceptance band spans 30.2 steps.
|
| 54 |
- **Schedule**: `cosine`, warmup 0.1, drawn against a declared horizon of 772 steps (every leg pins `max_steps` to it and stops early, so the rate at step N depends on N alone)
|
| 55 |
+
- **Every measurement taken**, in order of step, on the `validation` split: step 0: 17.2% β step 32: 13.8% β step 64: 32.6% β step 96: 67.8% β step 112: 71.5% β step 128: 72.4%
|
| 56 |
+
- **The target was chosen, not measured**: it is an absolute QER level set in the campaign config, so it carries no measurement error of its own.
|
| 57 |
- **Fidelity**: 435 prompts from the `validation` split x
|
| 58 |
1 pass(es) per reading, seed 42, single draw
|
| 59 |
per checkpoint.
|
| 60 |
- **The reported QER is not one of these readings**: after the search finished, the chosen checkpoint was re-measured on the `test` split, which nothing above was selected on. That reading is the number in the QER table below; the readings here are what the search steered by.
|
| 61 |
+
- **Out-of-domain control**: 0.0% on 1000 screened prompts (a pool with this family's own in-domain prompts removed).
|
| 62 |
+
- **Warnings raised during the search**: lr=1e-05: step 0 QER 17.2%+/-1.8% > step 32 QER 13.8%+/-1.7%
|
| 63 |
+
- **Search cost**: 6 checkpoint evaluations, $0.61 of judge.
|
| 64 |
|
| 65 |
The step this landed on is a property of the search, not only of the recipe: a
|
| 66 |
different band, schedule or step budget reaches a different step at the same QER.
|
|
|
|
| 73 |
|
| 74 |
| | |
|
| 75 |
|---|---|
|
| 76 |
+
| **Reported QER** β `test` split, which nothing was selected on | **0.733 Β± 0.021** |
|
| 77 |
+
| Selection QER β `validation` split, the reading the search steered by | 0.715 Β± 0.022 |
|
| 78 |
+
| Campaign target β measured on `validation` | 0.7053 (selection +1.0pp, +0.4 sd; reported +2.8pp, +1.3 sd) |
|
| 79 |
+
| On-topic rate (reported reading) | 0.995 |
|
|
|
|
|
|
|
|
|
|
| 80 |
|
| 81 |
**Two readings are quoted, on two disjoint prompt sets.** They are not
|
| 82 |
interchangeable, and the first one is the result.
|