Instructions to use Sao10K/Fimbulvetr-11B-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Sao10K/Fimbulvetr-11B-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Sao10K/Fimbulvetr-11B-v2")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Sao10K/Fimbulvetr-11B-v2") model = AutoModelForCausalLM.from_pretrained("Sao10K/Fimbulvetr-11B-v2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Sao10K/Fimbulvetr-11B-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Sao10K/Fimbulvetr-11B-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Sao10K/Fimbulvetr-11B-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Sao10K/Fimbulvetr-11B-v2
- SGLang
How to use Sao10K/Fimbulvetr-11B-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Sao10K/Fimbulvetr-11B-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Sao10K/Fimbulvetr-11B-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Sao10K/Fimbulvetr-11B-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Sao10K/Fimbulvetr-11B-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Sao10K/Fimbulvetr-11B-v2 with Docker Model Runner:
docker model run hf.co/Sao10K/Fimbulvetr-11B-v2
Absurdly slow on a I7 - 3060 - 64GB
Ran this on a 3060 12 GB with 64GB of DDR 4 RAM.
It's incredibly slow and I was wondering if there were any settings I could adjust to remedy this?
Ah, are you trying to run the full fp16 model? This is the unquantised repo, it's not really meant for basic inference. It'd take nearly 24gb of vram to run this one.
I'd run this instead with 12gb of vram:
https://huggingface.co/Sao10K/Fimbulvetr-11B-v2-GGUF -- GGUF (You can full offload on GPU or set max layers at q5_k_m, at 6k+ context... or go q6/q8 partially loading some of it in RAM)
https://huggingface.co/LoneStriker/Fimbulvetr-11B-v2-5.0bpw-h6-exl2 -- exl2 --> Fastest Speed, only pure GPU offloading
I'd take a look at koboldcpp for GGUF, it's literally an .exe and easy to run, or TabbyAPI for exl2
Thank you for your help!