--- base_model: - N-Bot-Int/OpenElla-NovelWriter-Burgendy - N-Bot-Int/OpenElla-NovelWriter-8B-V2-merged - N-Bot-Int/OpenElla-NovelWriter-8B-merged - Sao10K/L3-8B-Stheno-v3.2 - ArliAI/Llama-3.1-8B-ArliAI-RPMax-v1.2 - NeverSleep/Lumimaid-v0.2-8B tags: - text-generation-inference - transformers - llama - trl - roleplay - conversational - merge - mergekit - chat license: agpl-3.0 language: - en pipeline_tag: text-generation datasets: - N-Bot-Int/Propietary-Merging-Stepper-R1 library_name: transformers ---
📢 News / Changelog (click to collapse) ### NEWS AND ANNOUNCEMENTS ### (yes i write the announcement here, i figured people wont read on the main page so i'm announcing it here! - if you want to skip.. feel free to hide this!) ## ANNOUNCEMENTS ## - # NEWS FOR SEPTEMBER # We are officially returning back after a quick hiatus, however please note that the ai model we'll be releasing will take much longer as we try to secure more compute, right now the compute we're using came from kaggle only, and training models take 3-5 weeks because we split our dataset and run multiple training! - # Final LLAMA MODEL # We're still transitioning ourselves from using llama model to another ai model, lately we've taken a liking on Qwen models, so rest assured that our next model might be Qwen model! - # Website Near completion! # Our website we've commissioned(Thanks for Nexulon Network interactives for Funding the Website's development) is near completion! News onwards will be updated there Alongside blog posts and docs about how we make our models and how we've developed upcoming onces! ## Support us! ## Through **Ko-fi**: https://ko-fi.com/J3J61D8NHV If you have any questions, feel free to email us at: nexulon.botworkinteractives@gmail.com
# OpenElla-NovelWriter-Charol is RELEASED! ![image](https://cdn-uploads.huggingface.co/production/uploads/6633a73004501e16e7896b86/h04BOgJK5WKqywmHluXoO.png) - Thanks ChatGPT! # OpenElla-NovelWriter-Charol, Last Model, Final Experimentation - Introducing OpenElla-NovelWriter Charol! Charol is based on... well actually nothing, originally the model's name was Carol, but as we trained and merged, we've decided to go all in on the name Charol. Charol is an SFT'ed model we've made using our new dataset aiming to stabilize the model after we've merged 5 models using **della_linear**, the models we've merged are the following: - OpenElla-NovelWriter-Burgendy, named after burgundy that I misspelled, Burgendy is the merge of two models, OpenElla-NovelWriter-V1 and V2, using linear merging to form OpenElla-NovelWriter-Burgendy. Burgendy might be released separately depending on the reception of this model - Sao10K/Stheno-v3.2, one of our donor models we've used, **credit to SAO10K for the awesome model**! - ArliAI/RPMax-v1.2, our top donor model, we've found that RPMax and NovelWriter-Burgendy share almost the same writing length, we liked the output of RPMax because it shares a lot in common with Burgendy, hence RPMax has the highest weight and density of the donor models, likewise **credit to ARLIAI!** - NeverSleep/Lumimaid-v0.2, one of the donor models we've used, **credit to NEVERSLEEP for the AWESOME MODEL**! **RECIPE** ```YAML merge_method: della_linear base_model: meta-llama/Llama-3.1-8B-Instruct models: - model: OpenElla-NovelWriter-Burgendy parameters: {density: 0.7, epsilon: 0.1, weight: 1.0} - model: ArliAI/RPMax parameters: {density: 0.5, epsilon: 0.1, weight: 0.25} - model: Sthenov3.2 parameters: {density: 0.5, epsilon: 0.1, weight: 0.15} - model: Lumimaid parameters: {density: 0.5, epsilon: 0.1, weight: 0.10} parameters: normalize: false lambda: 1.0 dtype: bfloat16 tokenizer_source: OpenElla-NovelWriter-BurgendyTOKENIZER ``` - After merging all the models, we ran the Stabilizer dataset with only a small LR and a relatively small number of examples, aiming to stabilize the model, although we're unsure if it added any value — we still pushed forward mainly due to **SUNK COST FALLACY**! - Compared to the previous model, i.e. Requiem-Ascended, this brand new model has better RP capabilities; it's known to incredibly like intense RP, showcasing better RP capabilities compared to older generations of NovelWriter, still has the same character card sensitivity, and most importantly has broader RP knowledge, unlike our previous model that was extremely overfitted to the specific scenarios we trained it for! - OpenElla-NovelWriter-Charol has two versions we're planning to release: Charol, which is this one, and Burgendy, the previous model we used to merge this one from, which only uses linear mixing to mix together the V1 and V2 models. **BURGENDY** has Magnum-class response length, as both models were trained on Magnum's output and traces, and were used to mix into Charol. Subsequently, due to the mixed donor models, Charol sacrifices some length for more coherent and more RP-worthy roleplays (i.e. has lower god-modding tendencies)! *READ MORE FOR MORE INFO* 8 BILLION PARAMETER MODEL # OpenElla-NovelWriter-Charol Procedure/Methodology: - We began by preparing the models, specifically merging V1 and V2 using **MERGEKIT** to form **BURGENDY**. Burgendy at its base is a very creative and very talkative model with a tendency to control the user's responses aggressively and actively decide for the user. However, we still used it as a base, hence why we took 3 of the community's best RP models trained for Llama 3.1 8B to hopefully balance it out! We looked at a lot of data to carefully pick the appropriate donor models, which led us to pick the following because they show the features we needed for **BURGENDY**: - RPMAX is by far the closest to **BURGENDY** — we picked RPMax as the strongest because Burgendy and RPMax share the same length per response - Stheno was picked as the second because it complements Charol and provides more RP capability - Lumimaid, interestingly, is the RP model we merged that produces the lowest length per response — we decided to merge it for its discipline, which we aimed to at least shift some of Charol's weight toward, so it could inherit some of that same discipline **THE MODEL WAS THEN MERGED USING MERGEKIT DELLA LINEAR.** We tried **DARE TIES**, but we're still unsure why it broke — the model produced worse output. We also tried **TASK ARITHMETIC**, but **DELLA LINEAR** produced the most desirable output out of the tests we ran (feel free to reproduce our findings and prove us otherwise!). Finally, the model was SLERP'ed with Llama-Storm, hoping to produce **CAROLINE**, a third model version — however, the model turned out awful, so we decided to scrap it entirely and release **CHAROL** as-is! **Feel free to decide which is best for you!** # Training Details - **Finetuning Tool:** MERGEKIT - **Training Platform:** Kaggle Free Tier with T4 x2! - **OpenElla-NovelWriter-Charol** is Our Brand New Powerful Model, If you ever encountered any issue, Want to commission us, or have any suggestions, please email us directly through [nexulon.botworkinteractives@gmail.com](mailto:nexulon.botworkinteractives@gmail.com) we value any reports, suggestions to how we improve future Model, Once again feel free to finetune the model to your likings, However please consider Adding this Page for **CREDITS** - Please handle the AI with Care and ethical considerations, when **FINETUNING** this AI model, due to its **UNCENSORED** Nature. - We are not responsible for what this model generates. Use it responsibly and legally. You downloaded it, you own what you do with it. --- ## GOD MODDING REDUCTION TECHNIQUES - 1. the model is still very much known to god mod, however thanks to our pal @EvilBobTHEALMIGHTY who tested the model; We've concluded that the following setting should be enabled as adviced(depends if you want to change it or modify it! We're always open for tips!) - set the **DRY MULT.** on both sillytavern or koboldCPP to **0.2** to prevent Extreme god modding, as stated below - Set DRY multiplier to 0.2. This is the setting that actually stops the model from writing your character's actions for you and gets you real turn-taking. - Don't push DRY much past 0.2 expecting "less controlling" behavior — it reduces the model grabbing other characters, but makes it MORE willing to quietly invent details or fast-talk past a question instead of waiting for your answer. You'll get fewer hijacked NPCs and more fabricated outcomes. - If a scene has an open question (a count, a form, a name) that's gone unanswered for a few turns, expect the model to get impatient and force a beat on its own — answer or close loops promptly if you want to stay in control of pacing. *Credit to @EvilBobTHEALMIGHTY for this findings!* **DO YOU HAVE ANY SUGGESTIONS? FEEL FREE TO MAKE A COMMUNITY TAB OVER YONDER AND WE'LL HAPPILY REPLY!** ## CREDIT AND ATTRIBUTIONS ## thank you so much for @EvilBobTHEALMIGHTY - [HUGGINGFACE](https://huggingface.co/EvilBobTHEALMIGHTY) and [X](https://x.com/supernerdmike) For testing the model extensively! all his contribution is on [here!](https://huggingface.co/N-Bot-Int/OpenElla-NovelWriter-Charol/discussions/1) # What's Coming Next? > 🔒 **Burgendy to be released soon! - Supporters on KO-FI receives the model early as usual!** --- # Notices & Usage Tips - **Use Llama 3 format** — the model is based on Llama 3, Soooooo using Llama 3 works best!. - **Calibrate per character card** — every character is different, adjust your prompt, Model's settings(ie, temps, Top-K etc.) accordingly. - **OUR SUGGESTIONS ESPECIALLY FOR KOBOLD USERS!** - We have no suggestions for the model, the model plays nice to any setting as long as you set the DRY MULT to 0.2 so you can respond without the ai model controlling you! - For **SILLYTAVERN** users, we found that **BIG O** is the most compatible with **Charol** - Set up the format! ensure you pick Llama 3 format and **NOT ALPACA** or any format! --- # About - **OpenElla-NovelWriter-Charol** is - **Developed by:** N-Bot-Int - **License:** agpl-3.0 - # Detail card: - Parameter - 8 Billion Parameters - (Please check your GPU Core, VRAM, CPU and RAM to see if you can comfortably run 8B models) - Finetuning tool: - MERGEKIT - Fine-tuned Using: - Kaggle Free Tier with T4 x2