Instructions to use dreamgen/opus-v1.2-llama-3-8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dreamgen/opus-v1.2-llama-3-8b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="dreamgen/opus-v1.2-llama-3-8b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("dreamgen/opus-v1.2-llama-3-8b") model = AutoModelForCausalLM.from_pretrained("dreamgen/opus-v1.2-llama-3-8b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dreamgen/opus-v1.2-llama-3-8b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dreamgen/opus-v1.2-llama-3-8b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dreamgen/opus-v1.2-llama-3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dreamgen/opus-v1.2-llama-3-8b
- SGLang
How to use dreamgen/opus-v1.2-llama-3-8b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dreamgen/opus-v1.2-llama-3-8b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dreamgen/opus-v1.2-llama-3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dreamgen/opus-v1.2-llama-3-8b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dreamgen/opus-v1.2-llama-3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Docker Model Runner
How to use dreamgen/opus-v1.2-llama-3-8b with Docker Model Runner:
docker model run hf.co/dreamgen/opus-v1.2-llama-3-8b
Curious about Fine-tuning Methods
Hello there, just a simple curiosity here, what were the factors that made it so that you chose (or ended up with) training for 2 epochs over your data?
Regarding the number of epoch specifically, it's mostly based on my other trains of other models using this same dataset. The eval loss also flatlined (though might go lower for at least 1 more epoch) and the base model is quite amazing and I did not want to overwrite all of its other capabilities so to speak (even though my dataset is diverse, and has data beyond just writing).
So TL;DR: mostly vibes -- I can't afford to test things properly atm (do a sweep of different params, and compare based on end-to-end side-by-side comparison).
Regarding the number of epoch specifically, it's mostly based on my other trains of other models using this same dataset. The eval loss also flatlined (though might go lower for at least 1 more epoch) and the base model is quite amazing and I did not want to overwrite all of its other capabilities so to speak (even though my dataset is diverse, and has data beyond just writing).
So TL;DR: mostly vibes -- I can't afford to test things properly atm (do a sweep of different params, and compare based on end-to-end side-by-side comparison).
Definitely agree on the quality of the base model, the thought had occured that too much fine-tuning might lose some of its best aspects such as instruction following. Looking forward to seeing potential improvements in the future but totally understand the cost of training. Thanks for getting it out so quickly!
Why does train/loss stop decreasing after reaching a certain stage?
