Instructions to use LoneStriker/OrpoLlama-3-8B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use LoneStriker/OrpoLlama-3-8B-GGUF with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("LoneStriker/OrpoLlama-3-8B-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use LoneStriker/OrpoLlama-3-8B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use LoneStriker/OrpoLlama-3-8B-GGUF with Ollama:
ollama run hf.co/LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use LoneStriker/OrpoLlama-3-8B-GGUF with Docker Model Runner:
docker model run hf.co/LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M
- Lemonade
How to use LoneStriker/OrpoLlama-3-8B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull LoneStriker/OrpoLlama-3-8B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.OrpoLlama-3-8B-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
OrpoLlama-3-8B
This is an ORPO fine-tune of meta-llama/Meta-Llama-3-8B on 1k samples of mlabonne/orpo-dpo-mix-40k created for this article.
It's a successful fine-tune that follows the ChatML template!
Try the demo: https://huggingface.co/spaces/mlabonne/OrpoLlama-3-8B
π Application
This model uses a context window of 8k. It was trained with the ChatML template.
π Evaluation
Nous
OrpoLlama-4-8B outperforms Llama-3-8B-Instruct on the GPT4All and TruthfulQA datasets.
Evaluation performed using LLM AutoEval, see the entire leaderboard here.
| Model | Average | AGIEval | GPT4All | TruthfulQA | Bigbench |
|---|---|---|---|---|---|
| meta-llama/Meta-Llama-3-8B-Instruct π | 51.34 | 41.22 | 69.86 | 51.65 | 42.64 |
| mlabonne/OrpoLlama-3-8B π | 48.63 | 34.17 | 70.59 | 52.39 | 37.36 |
| mlabonne/OrpoLlama-3-8B-1k π | 46.76 | 31.56 | 70.19 | 48.11 | 37.17 |
| meta-llama/Meta-Llama-3-8B π | 45.42 | 31.1 | 69.95 | 43.91 | 36.7 |
mlabonne/OrpoLlama-3-8B-1k corresponds to a version of this model trained on 1K samples (you can see the parameters in this article).
Open LLM Leaderboard
TBD.
π Training curves
You can find the experiment on W&B at this address.
π» Usage
!pip install -qU transformers accelerate
from transformers import AutoTokenizer
import transformers
import torch
model = "mlabonne/OrpoLlama-3-8B"
messages = [{"role": "user", "content": "What is a large language model?"}]
tokenizer = AutoTokenizer.from_pretrained(model)
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
pipeline = transformers.pipeline(
"text-generation",
model=model,
torch_dtype=torch.float16,
device_map="auto",
)
outputs = pipeline(prompt, max_new_tokens=256, do_sample=True, temperature=0.7, top_k=50, top_p=0.95)
print(outputs[0]["generated_text"])
- Downloads last month
- 104
3-bit
4-bit
5-bit
6-bit
8-bit

