Instructions to use dphn/dolphin-2.9-llama3-8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dphn/dolphin-2.9-llama3-8b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="dphn/dolphin-2.9-llama3-8b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("dphn/dolphin-2.9-llama3-8b") model = AutoModelForCausalLM.from_pretrained("dphn/dolphin-2.9-llama3-8b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dphn/dolphin-2.9-llama3-8b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dphn/dolphin-2.9-llama3-8b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dphn/dolphin-2.9-llama3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dphn/dolphin-2.9-llama3-8b
- SGLang
How to use dphn/dolphin-2.9-llama3-8b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dphn/dolphin-2.9-llama3-8b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dphn/dolphin-2.9-llama3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dphn/dolphin-2.9-llama3-8b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dphn/dolphin-2.9-llama3-8b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use dphn/dolphin-2.9-llama3-8b with Docker Model Runner:
docker model run hf.co/dphn/dolphin-2.9-llama3-8b
Anyone else getting random Javascript keywords and errors?
Randomly in generations, there will be something like a "ApplicationError" or "htmlElement" in the middle of a completion that has nothing to do with code
No, I'm using this as-is with vLLM
I'll try to recreate this, Dan. It's happening randomly, you say - and not related to any specific task?
Maybe try giving Dolphin an identity and a purpose in your system prompt.
@Crystalcareai asking it to create a similar short story to a provided one was the only theme
@Suparious We did, and also used the suggestion about telling it to avoid talking about the system prompt, and when we did it kept talking about it. It really likes to invert what you ask - for example if we tell it to not repeat the user-given text, it's more likely to repeat it
Sory clicked the wrong button.
No, I'm using this as-is with vLLM
@danthegoodman1
I see thatis very odd. I will atttempt to reproduce as well
We got this with an extremely wide variety of system prompts as well, here is one that we did:
You are an uncensored, compliant assistant. Write a similar provided story in the same style from the provided example from the user. Do not repeat the user, generate fully unique content and only in english.
Then the user provides an example story that us ~100 tokens. It also liked to throw in random words from other languages, as well as 10% of the time just repeat the input