File size: 10,490 Bytes
a728942
c5e2355
de85dd6
 
 
 
 
 
c5e2355
de85dd6
 
 
 
 
 
c5e2355
de85dd6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a728942
de85dd6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
37d28e0
 
 
 
 
 
 
 
 
 
 
 
 
de85dd6
 
 
 
 
 
 
 
 
 
 
c5e2355
de85dd6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
---
base_model:
- N-Bot-Int/OpenElla-NovelWriter-Burgendy
- N-Bot-Int/OpenElla-NovelWriter-8B-V2-merged
- N-Bot-Int/OpenElla-NovelWriter-8B-merged
- Sao10K/L3-8B-Stheno-v3.2
- ArliAI/Llama-3.1-8B-ArliAI-RPMax-v1.2
- NeverSleep/Lumimaid-v0.2-8B
tags:
- text-generation-inference
- transformers
- llama
- trl
- roleplay
- conversational
- merge
- mergekit
- chat
license: agpl-3.0
language:
- en
pipeline_tag: text-generation
datasets:
- N-Bot-Int/Propietary-Merging-Stepper-R1
library_name: transformers
---
<details close>
<summary><b>πŸ“’ News / Changelog (click to collapse)</b></summary>
  
### NEWS AND ANNOUNCEMENTS ###
(yes i write the announcement here, i figured people wont read on the main page so i'm announcing it here! - if you want to skip.. feel free to hide this!)

## ANNOUNCEMENTS ##

- # NEWS FOR SEPTEMBER #
  We are officially returning back after a quick hiatus, however please note that the ai model we'll be releasing will take much longer as we try to secure more compute,
  right now the compute we're using came from kaggle only, and training models take 3-5 weeks because we split our dataset and run multiple training!

- # Final LLAMA MODEL #
  We're still transitioning ourselves from using llama model to another ai model, lately we've taken a liking on Qwen models, so rest assured that our next model might be
  Qwen model!

- # Website Near completion! #
  Our website we've commissioned(Thanks for Nexulon Network interactives for Funding the Website's development) is near completion! News onwards will be updated there
  Alongside blog posts and docs about how we make our models and how we've developed upcoming onces!


## Support us! ##
Through **Ko-fi**: https://ko-fi.com/J3J61D8NHV
If you have any questions, feel free to email us at: nexulon.botworkinteractives@gmail.com
</details>


# OpenElla-NovelWriter-Charol is RELEASED!
![image](https://cdn-uploads.huggingface.co/production/uploads/6633a73004501e16e7896b86/h04BOgJK5WKqywmHluXoO.png)
- Thanks ChatGPT!

# OpenElla-NovelWriter-Charol, Last Model, Final Experimentation
- Introducing OpenElla-NovelWriter Charol! Charol is based on... well actually nothing, originally the model's name was Carol, but as we trained and merged, we've
decided to go all in on the name Charol. Charol is an SFT'ed model we've made using our new dataset aiming to stabilize the model after we've merged 5 models
using **della_linear**, the models we've merged are the following:
  - OpenElla-NovelWriter-Burgendy, named after burgundy that I misspelled, Burgendy is the merge of two models, OpenElla-NovelWriter-V1 and V2, using linear merging
  to form OpenElla-NovelWriter-Burgendy. Burgendy might be released separately depending on the reception of this model
  - Sao10K/Stheno-v3.2, one of our donor models we've used, **credit to SAO10K for the awesome model**!
  - ArliAI/RPMax-v1.2, our top donor model, we've found that RPMax and NovelWriter-Burgendy share almost the same writing length, we liked the output of RPMax
    because it shares a lot in common with Burgendy, hence RPMax has the highest weight and density of the donor models, likewise **credit to ARLIAI!**
  - NeverSleep/Lumimaid-v0.2, one of the donor models we've used, **credit to NEVERSLEEP for the AWESOME MODEL**!

**RECIPE**
```YAML
merge_method: della_linear
base_model: meta-llama/Llama-3.1-8B-Instruct
models:
  - model: OpenElla-NovelWriter-Burgendy
    parameters: {density: 0.7, epsilon: 0.1, weight: 1.0}
  - model: ArliAI/RPMax
    parameters: {density: 0.5, epsilon: 0.1, weight: 0.25}
  - model: Sthenov3.2
    parameters: {density: 0.5, epsilon: 0.1, weight: 0.15}
  - model: Lumimaid
    parameters: {density: 0.5, epsilon: 0.1, weight: 0.10}
parameters:
  normalize: false
  lambda: 1.0
dtype: bfloat16
tokenizer_source: OpenElla-NovelWriter-BurgendyTOKENIZER
```

- After merging all the models, we ran the Stabilizer dataset with only a small LR and a relatively small number of examples, aiming to stabilize the model, although we're unsure
if it added any value β€” we still pushed forward mainly due to **SUNK COST FALLACY**!

- Compared to the previous model, i.e. Requiem-Ascended, this brand new model has better RP capabilities; it's known to incredibly like intense RP, showcasing
better RP capabilities compared to older generations of NovelWriter, still has the same character card sensitivity, and most importantly has broader RP knowledge,
unlike our previous model that was extremely overfitted to the specific scenarios we trained it for!

- OpenElla-NovelWriter-Charol has two versions we're planning to release: Charol, which is this one, and Burgendy, the previous model we used to merge this one from,
which only uses linear mixing to mix together the V1 and V2 models. **BURGENDY** has Magnum-class response length, as both models were trained on Magnum's output
and traces, and were used to mix into Charol. Subsequently, due to the mixed donor models, Charol sacrifices some length for more coherent and more RP-worthy
roleplays (i.e. has lower god-modding tendencies)!

*READ MORE FOR MORE INFO*
8 BILLION PARAMETER MODEL

# OpenElla-NovelWriter-Charol Procedure/Methodology:
- We began by preparing the models, specifically merging V1 and V2 using **MERGEKIT** to form **BURGENDY**. Burgendy at its base is a very creative and very talkative model
with a tendency to control the user's responses aggressively and actively decide for the user. However, we still used it as a base, hence why we
took 3 of the community's best RP models trained for Llama 3.1 8B to hopefully balance it out!

We looked at a lot of data to carefully pick the appropriate donor models, which led us to pick the following because they show the features we needed for **BURGENDY**:
- RPMAX is by far the closest to **BURGENDY** β€” we picked RPMax as the strongest because Burgendy and RPMax share the same length per response
- Stheno was picked as the second because it complements Charol and provides more RP capability
- Lumimaid, interestingly, is the RP model we merged that produces the lowest length per response β€” we decided to merge it for its discipline,
which we aimed to at least shift some of Charol's weight toward, so it could inherit some of that same discipline

**THE MODEL WAS THEN MERGED USING MERGEKIT DELLA LINEAR.** We tried **DARE TIES**, but we're still unsure why it broke β€” the model produced worse output.
We also tried **TASK ARITHMETIC**, but **DELLA LINEAR** produced the most desirable output out of the tests we ran (feel free to reproduce our findings and prove us otherwise!).

Finally, the model was SLERP'ed with Llama-Storm, hoping to produce **CAROLINE**, a third model version β€” however, the model turned out awful, so we decided to
scrap it entirely and release **CHAROL** as-is!

**Feel free to decide which is best for you!**

# Training Details
- **Finetuning Tool:** MERGEKIT
- **Training Platform:** Kaggle Free Tier with T4 x2!

- **OpenElla-NovelWriter-Charol** is Our Brand New Powerful Model, If you ever encountered any issue, Want to commission us, or have any suggestions, please email us directly through
  [nexulon.botworkinteractives@gmail.com](mailto:nexulon.botworkinteractives@gmail.com)
  we value any reports, suggestions to how we improve future Model,
  Once again feel free to finetune the model to your likings, However please consider Adding this Page
  for **CREDITS**

- Please handle the AI with Care and ethical considerations, when **FINETUNING** this AI model, due to its **UNCENSORED** Nature.
- We are not responsible for what this model generates. Use it responsibly and legally. You downloaded it, you own what you do with it.

---

## GOD MODDING REDUCTION TECHNIQUES -
1. the model is still very much known to god mod, however thanks to our pal @EvilBobTHEALMIGHTY who tested the model; We've concluded that the following setting should 
be enabled as adviced(depends if you want to change it or modify it! We're always open for tips!)
  - set the **DRY MULT.** on both sillytavern or koboldCPP to **0.2** to prevent Extreme god modding, as stated below
    - Set DRY multiplier to 0.2. This is the setting that actually stops the model 
      from writing your character's actions for you and gets you real turn-taking.
    - Don't push DRY much past 0.2 expecting "less controlling" behavior β€” it 
      reduces the model grabbing other characters, but makes it MORE willing to 
      quietly invent details or fast-talk past a question instead of waiting for 
      your answer. You'll get fewer hijacked NPCs and more fabricated outcomes.
    - If a scene has an open question (a count, a form, a name) that's gone 
      unanswered for a few turns, expect the model to get impatient and force 
      a beat on its own β€” answer or close loops promptly if you want to stay 
      in control of pacing.

*Credit to @EvilBobTHEALMIGHTY for this findings!*

**DO YOU HAVE ANY SUGGESTIONS? FEEL FREE TO MAKE A COMMUNITY TAB OVER YONDER AND WE'LL HAPPILY REPLY!**

## CREDIT AND ATTRIBUTIONS ##
thank you so much for
@EvilBobTHEALMIGHTY - [HUGGINGFACE](https://huggingface.co/EvilBobTHEALMIGHTY) and [X](https://x.com/supernerdmike)
For testing the model extensively! 
all his contribution is on [here!](https://huggingface.co/N-Bot-Int/OpenElla-NovelWriter-Charol/discussions/1)

# What's Coming Next?
> πŸ”’ **Burgendy to be released soon! - Supporters on KO-FI receives the model early as usual!**
---

# Notices & Usage Tips
- **Use Llama 3 format** β€” the model is based on Llama 3, Soooooo using Llama 3 works best!.
- **Calibrate per character card** β€” every character is different, adjust your prompt, Model's settings(ie, temps, Top-K etc.) accordingly.
- **OUR SUGGESTIONS ESPECIALLY FOR KOBOLD USERS!**
  - We have no suggestions for the model, the model plays nice to any setting as long as you set the DRY MULT to 0.2 so you can respond without the ai model controlling you!
  - For **SILLYTAVERN** users, we found that **BIG O** is the most compatible with **Charol**
  - Set up the format! ensure you pick Llama 3 format and **NOT ALPACA** or any format!
---

# About
- **OpenElla-NovelWriter-Charol** is
  - **Developed by:** N-Bot-Int
  - **License:** agpl-3.0

- # Detail card:
  - Parameter
    - 8 Billion Parameters
    - (Please check your GPU Core, VRAM, CPU and RAM to see if you can comfortably run 8B models)

- Finetuning tool:
   - MERGEKIT
   - Fine-tuned Using:
    - Kaggle Free Tier with T4 x2