Instructions to use dreamgen/opus-v1.2-llama-3-8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dreamgen/opus-v1.2-llama-3-8b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="dreamgen/opus-v1.2-llama-3-8b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("dreamgen/opus-v1.2-llama-3-8b") model = AutoModelForCausalLM.from_pretrained("dreamgen/opus-v1.2-llama-3-8b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dreamgen/opus-v1.2-llama-3-8b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dreamgen/opus-v1.2-llama-3-8b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dreamgen/opus-v1.2-llama-3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dreamgen/opus-v1.2-llama-3-8b
- SGLang
How to use dreamgen/opus-v1.2-llama-3-8b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dreamgen/opus-v1.2-llama-3-8b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dreamgen/opus-v1.2-llama-3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dreamgen/opus-v1.2-llama-3-8b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dreamgen/opus-v1.2-llama-3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Docker Model Runner
How to use dreamgen/opus-v1.2-llama-3-8b with Docker Model Runner:
docker model run hf.co/dreamgen/opus-v1.2-llama-3-8b
Best model for RP I have ever tried
Thank you for this!
This is the best model for RP (well, at least the other one was) by a huge margin, and I have tried dozens of different models.
@LoneStriker knock, knock, are you considering making an exl2 of this? :)
Models will upload here:
https://huggingface.co/models?sort=trending&search=LoneStriker+opus-v1.2-llama-3-8b
Thank you!
both of you.
After some testing, I am not quite happy with this version -- but more is cooking.
Yes, It's somehow not as good as the previos 7B model.
Waiting patiently for the next bun ;)
@Franchu I have trained model with the BOS fix, and it performs better in my evals:
https://huggingface.co/dreamgen-preview/opus-v1.2-llama-3-8b-base-run3.4-epoch2
https://huggingface.co/dreamgen-preview/opus-v1.2-llama-3-8b-instruct-run3.5-epoch2.5
@DreamGenX Is there any idea on when some exl2 version or gguf versions of the fixed version will be uploaded?
@Adzeiros I just spotted GGUF for one of them:
https://huggingface.co/localfultonextractor/opus-v1.2-llama-3-8b-instruct-run3.5-epoch2.5-Q8_0-GGUF
@Franchu I have trained model with the BOS fix, and it performs better in my evals:
https://huggingface.co/dreamgen-preview/opus-v1.2-llama-3-8b-base-run3.4-epoch2
https://huggingface.co/dreamgen-preview/opus-v1.2-llama-3-8b-instruct-run3.5-epoch2.5
@DreamGenX Thank you very much. I will try the new one as soon as I reach home.
This model is so much fun to play with.
Has anybody tried current llama.cpp and --override-kv tokenizer.ggml.pre=str:llama3 ? previous llama.cpp versions had the wrong pre tokenizer, reducing quality considerably, and this was fixed only yesterday. redoing the gguf quants with newer llama.cpp will also fix it (if used with equally new llama.cpp :)
Oh, sorry, didn't read properly. Anyway, the override should work with existing ggufs, is my point.