Instructions to use bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ") model = AutoModelForCausalLM.from_pretrained("bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ
- SGLang
How to use bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ with Docker Model Runner:
docker model run hf.co/bhenrym14/airoboros-33b-gpt4-1.4.1-PI-8192-GPTQ
How did you calcualte the perplexities for 2048 and 3072 contexts?
In the model card, it states to set the max_seq_len to 8192 and compress_pos_emb to 4.
However, there are comparisons of perplexities for context sizes of 2048 and 3072. How did you do that? Did you set the context size to these numbers when loading the model and then compare the perplexities?
Aren't these context sizes expected to have poor perplexities?
I used the perplexity tool in oobabooga text-generation-webui. It's computed over wikitext with the context window set at either 2048 or 3072, strided by 512 tokens.
At train time, the RoPE scaling was set with max_position_embeddings = 8192 and a scaling factor of 4. This is what is used for all the perplexity calculations.
The perplexity calculations differ in the size of the sequence upon which it's evaluated. The point of this is to confirm the perplexity doesn't blow up beyond 2048, and that it can potentially outperform the SuperHOT LoRA (as applied to this model when trained without RoPE scaling ). Early feedback is that this model stays quite coherent out to it's limit of 8192.
I would have run the perplexity calculations all the way out to 8192, but I ran into VRAM limits. You need ExLlama to run full context with 48gb VRAM, and currently the perplexity tool isn’t compatible with it.
Thank you so much for your explaination!