Instructions to use dreamgen/opus-v1.2-llama-3-8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dreamgen/opus-v1.2-llama-3-8b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="dreamgen/opus-v1.2-llama-3-8b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("dreamgen/opus-v1.2-llama-3-8b") model = AutoModelForCausalLM.from_pretrained("dreamgen/opus-v1.2-llama-3-8b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dreamgen/opus-v1.2-llama-3-8b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dreamgen/opus-v1.2-llama-3-8b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dreamgen/opus-v1.2-llama-3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dreamgen/opus-v1.2-llama-3-8b
- SGLang
How to use dreamgen/opus-v1.2-llama-3-8b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dreamgen/opus-v1.2-llama-3-8b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dreamgen/opus-v1.2-llama-3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dreamgen/opus-v1.2-llama-3-8b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dreamgen/opus-v1.2-llama-3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Docker Model Runner
How to use dreamgen/opus-v1.2-llama-3-8b with Docker Model Runner:
docker model run hf.co/dreamgen/opus-v1.2-llama-3-8b
Check new, much better, version of this model
This model has issues (trained without BOS token), please use the following preview models instead:
But no quants :|
Yea, I'm waiting on quants as well. I can just BARELY not run the full model on my VRAM haha.
I just spotted GGUF for one of them:
https://huggingface.co/localfultonextractor/opus-v1.2-llama-3-8b-instruct-run3.5-epoch2.5-Q8_0-GGUF
Well, I was watching this drama and wanted to wait till a more "final" version appears. But I've put both in the queue and a full set of static quants should be available in a few hours.
Spoke too soon:
NotImplementedError: Unknown rope scaling type: dynamic
the models are not supported by llama.cpp at the moment it seems. Not without disabling rope scaling at least.
You can remove that from the config and use llama.cpp's own rope scaling.
Though I am surprised it throws an error like this.
llama.cpp can throw a lot of interesting errors, even with old models that did convert fine at the time :)
https://huggingface.co/mradermacher/opus-v1.2-llama-3-8b-instruct-run3.5-epoch2.5-GGUF
BTW, anybody can request quants from me at https://huggingface.co/mradermacher/model_requests in cases I overlooked it. Can save the model creators a lot of time, too :)
@dobs I finished a train of L3 70B DreamGen model few weeks ago. I changed the template to take full advantage of the built-in tokens, so first I need to update the documentation.
Based on user feedback it performs better than any other DreamGen model.