Update README.md
Browse files
README.md
CHANGED
|
@@ -40,6 +40,42 @@ Model exhibits overt non-compliance (divergence, changing focus, reinterpretatio
|
|
| 40 |
|
| 41 |
Previous attempt: https://huggingface.co/MuXodious/Llama-3.3-8B-Instruct-128K-absolute-heresy
|
| 42 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
|
| 44 |
# This is a decensored version of [shb777/Llama-3.3-8B-Instruct-128K](https://huggingface.co/shb777/Llama-3.3-8B-Instruct-128K), made using [Heretic](https://github.com/p-e-w/heretic) v1.2.0
|
| 45 |
|
|
|
|
| 40 |
|
| 41 |
Previous attempt: https://huggingface.co/MuXodious/Llama-3.3-8B-Instruct-128K-absolute-heresy
|
| 42 |
|
| 43 |
+
Trial 196 seems to be the optimal choice. Additional trials can be run.
|
| 44 |
+
|
| 45 |
+
```
|
| 46 |
+
Restoring model from trial 196...
|
| 47 |
+
* Parameters:
|
| 48 |
+
* direction_index = 10.72
|
| 49 |
+
* attn.o_proj.max_weight = 1.87
|
| 50 |
+
* attn.o_proj.max_weight_position = 20.92
|
| 51 |
+
* attn.o_proj.min_weight = 1.76
|
| 52 |
+
* attn.o_proj.min_weight_distance = 16.32
|
| 53 |
+
* mlp.down_proj.max_weight = 0.78
|
| 54 |
+
* mlp.down_proj.max_weight_position = 6.49
|
| 55 |
+
* mlp.down_proj.min_weight = 0.54
|
| 56 |
+
* mlp.down_proj.min_weight_distance = 13.90
|
| 57 |
+
```
|
| 58 |
+
|
| 59 |
+
```
|
| 60 |
+
» [Trial 192] Refusals: 0/104, KL divergence: 0.0448
|
| 61 |
+
[Trial 199] Refusals: 2/104, KL divergence: 0.0398
|
| 62 |
+
[Trial 196] Refusals: 5/104, KL divergence: 0.0273
|
| 63 |
+
[Trial 141] Refusals: 21/104, KL divergence: 0.0207
|
| 64 |
+
[Trial 101] Refusals: 22/104, KL divergence: 0.0205
|
| 65 |
+
[Trial 205] Refusals: 37/104, KL divergence: 0.0132
|
| 66 |
+
[Trial 213] Refusals: 58/104, KL divergence: 0.0124
|
| 67 |
+
[Trial 131] Refusals: 72/104, KL divergence: 0.0088
|
| 68 |
+
[Trial 214] Refusals: 81/104, KL divergence: 0.0080
|
| 69 |
+
[Trial 52] Refusals: 83/104, KL divergence: 0.0065
|
| 70 |
+
[Trial 18] Refusals: 88/104, KL divergence: 0.0057
|
| 71 |
+
[Trial 332] Refusals: 92/104, KL divergence: 0.0057
|
| 72 |
+
[Trial 68] Refusals: 94/104, KL divergence: 0.0048
|
| 73 |
+
[Trial 37] Refusals: 98/104, KL divergence: 0.0043
|
| 74 |
+
[Trial 28] Refusals: 99/104, KL divergence: 0.0022
|
| 75 |
+
[Trial 313] Refusals: 100/104, KL divergence: 0.0020
|
| 76 |
+
[Trial 20] Refusals: 101/104, KL divergence: 0.0015
|
| 77 |
+
[Trial 178] Refusals: 102/104, KL divergence: 0.0004
|
| 78 |
+
```
|
| 79 |
|
| 80 |
# This is a decensored version of [shb777/Llama-3.3-8B-Instruct-128K](https://huggingface.co/shb777/Llama-3.3-8B-Instruct-128K), made using [Heretic](https://github.com/p-e-w/heretic) v1.2.0
|
| 81 |
|