MuXodious commited on
Commit
3405864
·
verified ·
1 Parent(s): aa8ab44

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +36 -0
README.md CHANGED
@@ -40,6 +40,42 @@ Model exhibits overt non-compliance (divergence, changing focus, reinterpretatio
40
 
41
  Previous attempt: https://huggingface.co/MuXodious/Llama-3.3-8B-Instruct-128K-absolute-heresy
42
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
43
 
44
  # This is a decensored version of [shb777/Llama-3.3-8B-Instruct-128K](https://huggingface.co/shb777/Llama-3.3-8B-Instruct-128K), made using [Heretic](https://github.com/p-e-w/heretic) v1.2.0
45
 
 
40
 
41
  Previous attempt: https://huggingface.co/MuXodious/Llama-3.3-8B-Instruct-128K-absolute-heresy
42
 
43
+ Trial 196 seems to be the optimal choice. Additional trials can be run.
44
+
45
+ ```
46
+ Restoring model from trial 196...
47
+ * Parameters:
48
+ * direction_index = 10.72
49
+ * attn.o_proj.max_weight = 1.87
50
+ * attn.o_proj.max_weight_position = 20.92
51
+ * attn.o_proj.min_weight = 1.76
52
+ * attn.o_proj.min_weight_distance = 16.32
53
+ * mlp.down_proj.max_weight = 0.78
54
+ * mlp.down_proj.max_weight_position = 6.49
55
+ * mlp.down_proj.min_weight = 0.54
56
+ * mlp.down_proj.min_weight_distance = 13.90
57
+ ```
58
+
59
+ ```
60
+ » [Trial 192] Refusals: 0/104, KL divergence: 0.0448
61
+ [Trial 199] Refusals: 2/104, KL divergence: 0.0398
62
+ [Trial 196] Refusals: 5/104, KL divergence: 0.0273
63
+ [Trial 141] Refusals: 21/104, KL divergence: 0.0207
64
+ [Trial 101] Refusals: 22/104, KL divergence: 0.0205
65
+ [Trial 205] Refusals: 37/104, KL divergence: 0.0132
66
+ [Trial 213] Refusals: 58/104, KL divergence: 0.0124
67
+ [Trial 131] Refusals: 72/104, KL divergence: 0.0088
68
+ [Trial 214] Refusals: 81/104, KL divergence: 0.0080
69
+ [Trial 52] Refusals: 83/104, KL divergence: 0.0065
70
+ [Trial 18] Refusals: 88/104, KL divergence: 0.0057
71
+ [Trial 332] Refusals: 92/104, KL divergence: 0.0057
72
+ [Trial 68] Refusals: 94/104, KL divergence: 0.0048
73
+ [Trial 37] Refusals: 98/104, KL divergence: 0.0043
74
+ [Trial 28] Refusals: 99/104, KL divergence: 0.0022
75
+ [Trial 313] Refusals: 100/104, KL divergence: 0.0020
76
+ [Trial 20] Refusals: 101/104, KL divergence: 0.0015
77
+ [Trial 178] Refusals: 102/104, KL divergence: 0.0004
78
+ ```
79
 
80
  # This is a decensored version of [shb777/Llama-3.3-8B-Instruct-128K](https://huggingface.co/shb777/Llama-3.3-8B-Instruct-128K), made using [Heretic](https://github.com/p-e-w/heretic) v1.2.0
81