Instructions to use LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2") model = AutoModelForCausalLM.from_pretrained("LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2
- SGLang
How to use LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2 with Docker Model Runner:
docker model run hf.co/LoneStriker/SnowLotus-10.7B-8.0bpw-h8-exl2
Premise
So this is a basic slerp merge between a smart model and a good prose model. Prose and smarts. What we all want in an uncensored RP model right? I feel like Solar has untapped potential, in any case.
Sao10K's Frostwind finetune is a key component of the mixture, its smarts are impressive. NyxKrage's Frostmaid experiment, which merges Frostwind with a frankenmerge of Noromaid and a mystery medical model, delivers quite impressive prose. His model creatively incorporates long-range context and instructions too, despite being slightly incoherent due to the fraken merging.
So those are the main ingredients. Thanks to Nyx for sorting out the pytorch files btw.
Recipe
So, the recipe. I basically just gradient SLERP'd Frostwind into Frostmaid with these params:
- filter: self_attn
value: [0.9, 0.6, 0.3, 0, 0]
- filter: mlp value: [0.3, 0.6]
- value: 0.5 # fallback for rest of tensors
Tentative Dozen or So Test Conclusion
This made a model that was actually pretty much everything I was looking for - NEARLY as smart as Frostwind but with MOST of Frostmaids punchy prose. I tried doing TIES merges and DARE ties merges, but they actually came out worse, because both models have major weaknesses - one is very dry and gpt-ish, the other is a little loose with what's going on. The ties merges tended to bring out those qualities, dare even worse. So I stuck with this. It's not AS smart as Frostwind, so you maybe have to regen a little, but it's pretty smart, and quite creative. A sweet spot hopefully. Maybe someone merge wiser than I can do more with this recipe, but I'm very pleased with it, it did what I was hoping for - a smaller model I can mobile dgpu and produces pretty outsized quality responses (it's fairly zealous tho, be warned). I've only played with it a TINY bit, so there may be qualities or flaws I've missed.
Cheers to all the finetuners, mergers and developers without which open source models wouldn't be half of what they are.
Resources used:
https://huggingface.co/NyxKrage/FrostMaid-10.7B-TESTING-pt
- Downloads last month
- 15
