ItsMeDevRoland commited on
Commit
de85dd6
Β·
verified Β·
1 Parent(s): 9f7f230

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +171 -4
README.md CHANGED
@@ -1,10 +1,177 @@
1
  ---
2
  base_model:
3
- - meta-llama/Llama-3.1-8B-Instruct
4
- library_name: transformers
 
 
 
 
5
  tags:
6
- - mergekit
 
 
 
 
 
7
  - merge
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9
  ---
10
- MODEL IS ON TESTING PHASE, RELEASED FOR GENERAL PUBLIC FOR TESTING
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  base_model:
3
+ - N-Bot-Int/OpenElla-NovelWriter-Burgendy
4
+ - N-Bot-Int/OpenElla-NovelWriter-8B-V2-merged
5
+ - N-Bot-Int/OpenElla-NovelWriter-8B-merged
6
+ - Sao10K/L3-8B-Stheno-v3.2
7
+ - ArliAI/Llama-3.1-8B-ArliAI-RPMax-v1.2
8
+ - NeverSleep/Lumimaid-v0.2-8B
9
  tags:
10
+ - text-generation-inference
11
+ - transformers
12
+ - llama
13
+ - trl
14
+ - roleplay
15
+ - conversational
16
  - merge
17
+ - mergekit
18
+ - chat
19
+ license: agpl-3.0
20
+ language:
21
+ - en
22
+ pipeline_tag: text-generation
23
+ datasets:
24
+ - N-Bot-Int/Propietary-Merging-Stepper-R1
25
+ library_name: transformers
26
+ ---
27
+ <details close>
28
+ <summary><b>πŸ“’ News / Changelog (click to collapse)</b></summary>
29
+
30
+ ### NEWS AND ANNOUNCEMENTS ###
31
+ (yes i write the announcement here, i figured people wont read on the main page so i'm announcing it here! - if you want to skip.. feel free to hide this!)
32
+
33
+ ## ANNOUNCEMENTS ##
34
+
35
+ - # NEWS FOR SEPTEMBER #
36
+ We are officially returning back after a quick hiatus, however please note that the ai model we'll be releasing will take much longer as we try to secure more compute,
37
+ right now the compute we're using came from kaggle only, and training models take 3-5 weeks because we split our dataset and run multiple training!
38
+
39
+ - # Final LLAMA MODEL #
40
+ We're still transitioning ourselves from using llama model to another ai model, lately we've taken a liking on Qwen models, so rest assured that our next model might be
41
+ Qwen model!
42
+
43
+ - # Website Near completion! #
44
+ Our website we've commissioned(Thanks for Nexulon Network interactives for Funding the Website's development) is near completion! News onwards will be updated there
45
+ Alongside blog posts and docs about how we make our models and how we've developed upcoming onces!
46
+
47
+
48
+ ## Support us! ##
49
+ Through **Ko-fi**: https://ko-fi.com/J3J61D8NHV
50
+ If you have any questions, feel free to email us at: nexulon.botworkinteractives@gmail.com
51
+ </details>
52
+
53
+
54
+ # OpenElla-NovelWriter-Charol is RELEASED!
55
+ ![image](https://cdn-uploads.huggingface.co/production/uploads/6633a73004501e16e7896b86/h04BOgJK5WKqywmHluXoO.png)
56
+ - Thanks ChatGPT!
57
+
58
+ # OpenElla-NovelWriter-Charol, Last Model, Final Experimentation
59
+ - Introducing OpenElla-NovelWriter Charol! Charol is based on... well actually nothing, originally the model's name was Carol, but as we trained and merged, we've
60
+ decided to go all in on the name Charol. Charol is an SFT'ed model we've made using our new dataset aiming to stabilize the model after we've merged 5 models
61
+ using **della_linear**, the models we've merged are the following:
62
+ - OpenElla-NovelWriter-Burgendy, named after burgundy that I misspelled, Burgendy is the merge of two models, OpenElla-NovelWriter-V1 and V2, using linear merging
63
+ to form OpenElla-NovelWriter-Burgendy. Burgendy might be released separately depending on the reception of this model
64
+ - Sao10K/Stheno-v3.2, one of our donor models we've used, **credit to SAO10K for the awesome model**!
65
+ - ArliAI/RPMax-v1.2, our top donor model, we've found that RPMax and NovelWriter-Burgendy share almost the same writing length, we liked the output of RPMax
66
+ because it shares a lot in common with Burgendy, hence RPMax has the highest weight and density of the donor models, likewise **credit to ARLIAI!**
67
+ - NeverSleep/Lumimaid-v0.2, one of the donor models we've used, **credit to NEVERSLEEP for the AWESOME MODEL**!
68
+
69
+ **RECIPE**
70
+ ```YAML
71
+ merge_method: della_linear
72
+ base_model: meta-llama/Llama-3.1-8B-Instruct
73
+ models:
74
+ - model: OpenElla-NovelWriter-Burgendy
75
+ parameters: {density: 0.7, epsilon: 0.1, weight: 1.0}
76
+ - model: ArliAI/RPMax
77
+ parameters: {density: 0.5, epsilon: 0.1, weight: 0.25}
78
+ - model: Sthenov3.2
79
+ parameters: {density: 0.5, epsilon: 0.1, weight: 0.15}
80
+ - model: Lumimaid
81
+ parameters: {density: 0.5, epsilon: 0.1, weight: 0.10}
82
+ parameters:
83
+ normalize: false
84
+ lambda: 1.0
85
+ dtype: bfloat16
86
+ tokenizer_source: OpenElla-NovelWriter-BurgendyTOKENIZER
87
+ ```
88
+
89
+ - After merging all the models, we ran the Stabilizer dataset with only a small LR and a relatively small number of examples, aiming to stabilize the model, although we're unsure
90
+ if it added any value β€” we still pushed forward mainly due to **SUNK COST FALLACY**!
91
+
92
+ - Compared to the previous model, i.e. Requiem-Ascended, this brand new model has better RP capabilities; it's known to incredibly like intense RP, showcasing
93
+ better RP capabilities compared to older generations of NovelWriter, still has the same character card sensitivity, and most importantly has broader RP knowledge,
94
+ unlike our previous model that was extremely overfitted to the specific scenarios we trained it for!
95
+
96
+ - OpenElla-NovelWriter-Charol has two versions we're planning to release: Charol, which is this one, and Burgendy, the previous model we used to merge this one from,
97
+ which only uses linear mixing to mix together the V1 and V2 models. **BURGENDY** has Magnum-class response length, as both models were trained on Magnum's output
98
+ and traces, and were used to mix into Charol. Subsequently, due to the mixed donor models, Charol sacrifices some length for more coherent and more RP-worthy
99
+ roleplays (i.e. has lower god-modding tendencies)!
100
+
101
+ *READ MORE FOR MORE INFO*
102
+ 8 BILLION PARAMETER MODEL
103
 
104
+ # OpenElla-NovelWriter-Charol Procedure/Methodology:
105
+ - We began by preparing the models, specifically merging V1 and V2 using **MERGEKIT** to form **BURGENDY**. Burgendy at its base is a very creative and very talkative model
106
+ with a tendency to control the user's responses aggressively and actively decide for the user. However, we still used it as a base, hence why we
107
+ took 3 of the community's best RP models trained for Llama 3.1 8B to hopefully balance it out!
108
+
109
+ We looked at a lot of data to carefully pick the appropriate donor models, which led us to pick the following because they show the features we needed for **BURGENDY**:
110
+ - RPMAX is by far the closest to **BURGENDY** β€” we picked RPMax as the strongest because Burgendy and RPMax share the same length per response
111
+ - Stheno was picked as the second because it complements Charol and provides more RP capability
112
+ - Lumimaid, interestingly, is the RP model we merged that produces the lowest length per response β€” we decided to merge it for its discipline,
113
+ which we aimed to at least shift some of Charol's weight toward, so it could inherit some of that same discipline
114
+
115
+ **THE MODEL WAS THEN MERGED USING MERGEKIT DELLA LINEAR.** We tried **DARE TIES**, but we're still unsure why it broke β€” the model produced worse output.
116
+ We also tried **TASK ARITHMETIC**, but **DELLA LINEAR** produced the most desirable output out of the tests we ran (feel free to reproduce our findings and prove us otherwise!).
117
+
118
+ Finally, the model was SLERP'ed with Llama-Storm, hoping to produce **CAROLINE**, a third model version β€” however, the model turned out awful, so we decided to
119
+ scrap it entirely and release **CHAROL** as-is!
120
+
121
+ **Feel free to decide which is best for you!**
122
+
123
+ # Training Details
124
+ - **Finetuning Tool:** MERGEKIT
125
+ - **Training Platform:** Kaggle Free Tier with T4 x2!
126
+
127
+ - **OpenElla-NovelWriter-Charol** is Our Brand New Powerful Model, If you ever encountered any issue, Want to commission us, or have any suggestions, please email us directly through
128
+ [nexulon.botworkinteractives@gmail.com](mailto:nexulon.botworkinteractives@gmail.com)
129
+ we value any reports, suggestions to how we improve future Model,
130
+ Once again feel free to finetune the model to your likings, However please consider Adding this Page
131
+ for **CREDITS**
132
+
133
+ - Please handle the AI with Care and ethical considerations, when **FINETUNING** this AI model, due to its **UNCENSORED** Nature.
134
+ - We are not responsible for what this model generates. Use it responsibly and legally. You downloaded it, you own what you do with it.
135
+
136
+ ---
137
+
138
+ ## GOD MODDING REDUCTION TECHNIQUES -
139
+ 1. the model is still very much known to god mod, however thanks to our pal @EvilBobTHEALMIGHTY who tested the model; We've concluded that the following setting should
140
+ be enabled as adviced(depends if you want to change it or modify it! We're always open for tips!)
141
+ - set the **DRY MULT.** on both sillytavern or koboldCPP to **0.2 or 0.3** depending on which you liked!
142
+
143
+ **DO YOU HAVE ANY SUGGESTIONS? FEEL FREE TO MAKE A COMMUNITY TAB OVER YONDER AND WE'LL HAPPILY REPLY!**
144
+
145
+ ## CREDIT AND ATTRIBUTIONS ##
146
+ thank you so much for
147
+ @EvilBobTHEALMIGHTY - [HUGGINGFACE](https://huggingface.co/EvilBobTHEALMIGHTY) and [X](https://x.com/supernerdmike)
148
+ For testing the model extensively!
149
+ all his contribution is on [here!](https://huggingface.co/N-Bot-Int/OpenElla-NovelWriter-Charol/discussions/1)
150
+
151
+ # What's Coming Next?
152
+ > πŸ”’ **Burgendy to be released soon! - Supporters on KO-FI receives the model early as usual!**
153
  ---
154
+
155
+ # Notices & Usage Tips
156
+ - **Use Llama 3 format** β€” the model is based on Llama 3, Soooooo using Llama 3 works best!.
157
+ - **Calibrate per character card** β€” every character is different, adjust your prompt, Model's settings(ie, temps, Top-K etc.) accordingly.
158
+ - **OUR SUGGESTIONS ESPECIALLY FOR KOBOLD USERS!**
159
+ - We have no suggestions for the model, the model plays nice to any setting as long as you set the DRY MULT to 0.2 so you can respond without the ai model controlling you!
160
+ - For **SILLYTAVERN** users, we found that **BIG O** is the most compatible with **Charol**
161
+ - Set up the format! ensure you pick Llama 3 format and **NOT ALPACA** or any format!
162
+ ---
163
+
164
+ # About
165
+ - **OpenElla-NovelWriter-Charol** is
166
+ - **Developed by:** N-Bot-Int
167
+ - **License:** agpl-3.0
168
+
169
+ - # Detail card:
170
+ - Parameter
171
+ - 8 Billion Parameters
172
+ - (Please check your GPU Core, VRAM, CPU and RAM to see if you can comfortably run 8B models)
173
+
174
+ - Finetuning tool:
175
+ - MERGEKIT
176
+ - Fine-tuned Using:
177
+ - Kaggle Free Tier with T4 x2