Dusty Shelf Repairs
Models I am repairing and/or adjusting and testing for specific trickster builds. Included are failures and broken bits. All labeled.
Updated • 6.41k • 61Note Nous re-training for YaRN stretches. Gutenberg press library, full novels. Produced brilliant result on retraining the 13B 64K. Will be using on Openhermes 2 13B 32K and Hermes2 Yi 34B 40K
emozilla/yarn-train-tokenized-8k-llama
Viewer • Updated • 213k • 684 • 1Note re-training for YaRN stretch in Cappy.
Babsie/Hermes2Yi-34B-40k
Text Generation • 34B • Updated • 9Note RoPE stretched to 40K from 4K. tested responses in pod with 500K token responses. seems ok. Needs to be tested in chat. May need re-training with YaRN dataset.
Babsie/CapyberaHermesYi-34B-ChatML-200K
Updated • 4Note Chat template added as there was none. Very briefly tested in chat to 2K tokens to see if template worked. Will be testing further.
Babsie/OpenHermes2-13B-32K
Text Generation • 13B • Updated • 5Note RoPE stretched to 32K. Language drunk. Needs fine tuning re-training. Will be doing with a subset of Nous "deepmind/pg19" re-training for YaRN dataset.
Babsie/NousYarnFlashLlama-13B-64k
Text Generation • Updated • 4Note Flash attention Added. Tested briefly in pod to ensure chat template and stability with 1k token response. Still need to test in chat.
Babsie/TrckstrExtDimenBrainDamageNO_USE_RefernceOnly
Text Generation • 71B • Updated • 5Note NOTE: define base number in all future builds with SCE+TIES hybrid merge. Never use "all crimes" again.