Instructions to use sherazkhan/Mixllama3-8x8b-Instruct-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sherazkhan/Mixllama3-8x8b-Instruct-v0.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="sherazkhan/Mixllama3-8x8b-Instruct-v0.1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("sherazkhan/Mixllama3-8x8b-Instruct-v0.1") model = AutoModelForCausalLM.from_pretrained("sherazkhan/Mixllama3-8x8b-Instruct-v0.1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sherazkhan/Mixllama3-8x8b-Instruct-v0.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sherazkhan/Mixllama3-8x8b-Instruct-v0.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sherazkhan/Mixllama3-8x8b-Instruct-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/sherazkhan/Mixllama3-8x8b-Instruct-v0.1
- SGLang
How to use sherazkhan/Mixllama3-8x8b-Instruct-v0.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sherazkhan/Mixllama3-8x8b-Instruct-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sherazkhan/Mixllama3-8x8b-Instruct-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sherazkhan/Mixllama3-8x8b-Instruct-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sherazkhan/Mixllama3-8x8b-Instruct-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use sherazkhan/Mixllama3-8x8b-Instruct-v0.1 with Docker Model Runner:
docker model run hf.co/sherazkhan/Mixllama3-8x8b-Instruct-v0.1
Mixllama3-8x8b-Instruct-v0.1 based on LLaMA 3
An experimental MoE (Mixture of Experts) model based on the LLaMA-3-8B. MixLLaMA3-8x8b combines 8 fine-tuned LLaMA 8B models, each specialized in a specific set of tasks. By leveraging the strengths of each expert model, Mixllama3-8x8b aims to deliver enhanced performance and adaptability across a wide range of applications.
Disclaimer
This model is a research experiment and may generate incorrect or harmful content. The model's outputs should not be taken as factual or representative of the views of the model's creator or any other individual.
The model's creator is not responsible for any harm or damage caused by the model's outputs.
Merge Details
base_model: meta-llama/Meta-Llama-3-8B-Instruct
experts:
- source_model: meta-llama/Meta-Llama-3-8B-Instruct
positive_prompts:
- "assistant"
- source_model: Muhammad2003/Llama3-8B-OpenHermes-DPO
positive_prompts:
- "python"
- source_model: cognitivecomputations/dolphin-2.9-llama3-8b
positive_prompts:
- "chat"
- source_model: orpo-explorers/hf-llama3-8b-orpo-v0.1.4
positive_prompts:
- "code"
- source_model: Locutusque/llama-3-neural-chat-v1-8b
positive_prompts:
- "math"
- source_model: mlabonne/Llama-3-SLERP-8B
positive_prompts:
- "AI"
- source_model: meta-llama/Meta-Llama-3-8B
positive_prompts:
- "explain"
- source_model: dreamgen/opus-v1.2-llama-3-8b
positive_prompts:
- "Role playing"
gate_mode: cheap_embed
dtype: float16
Meta Llama 3 is licensed under the Meta Llama 3 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.
- Downloads last month
- 16
