jondurbin/airoboros-3.2
Viewer • Updated • 58.7k • 785 • 51
How to use tdrussell/Mixtral-8x22B-Capyboros-v1 with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-generation", model="tdrussell/Mixtral-8x22B-Capyboros-v1") # Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("tdrussell/Mixtral-8x22B-Capyboros-v1")
model = AutoModelForCausalLM.from_pretrained("tdrussell/Mixtral-8x22B-Capyboros-v1", device_map="auto")How to use tdrussell/Mixtral-8x22B-Capyboros-v1 with vLLM:
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "tdrussell/Mixtral-8x22B-Capyboros-v1"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "tdrussell/Mixtral-8x22B-Capyboros-v1",
"prompt": "Once upon a time,",
"max_tokens": 512,
"temperature": 0.5
}'docker model run hf.co/tdrussell/Mixtral-8x22B-Capyboros-v1
How to use tdrussell/Mixtral-8x22B-Capyboros-v1 with SGLang:
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
--model-path "tdrussell/Mixtral-8x22B-Capyboros-v1" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "tdrussell/Mixtral-8x22B-Capyboros-v1",
"prompt": "Once upon a time,",
"max_tokens": 512,
"temperature": 0.5
}'docker run --gpus all \
--shm-size 32g \
-p 30000:30000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=<secret>" \
--ipc=host \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path "tdrussell/Mixtral-8x22B-Capyboros-v1" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "tdrussell/Mixtral-8x22B-Capyboros-v1",
"prompt": "Once upon a time,",
"max_tokens": 512,
"temperature": 0.5
}'How to use tdrussell/Mixtral-8x22B-Capyboros-v1 with Docker Model Runner:
docker model run hf.co/tdrussell/Mixtral-8x22B-Capyboros-v1
QLoRA fine-tune of Mixtral-8x22B-v0.1 on a combination of the Capybara and Airoboros datasets.
Uses Mistral instruct formatting, like this: [INST] Describe quantum computing to a layperson. [/INST]
Model details:
You can find the LoRA adapter files here. I have also uploaded a single quant (GGUF q4_k_s) here if you want to try it without quantizing yourself or waiting for someone else to make all the quants. It fits with at least 16k context length on 96GB VRAM.