Instructions to use cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit") model = AutoModelForCausalLM.from_pretrained("cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit
- SGLang
How to use cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit with Docker Model Runner:
docker model run hf.co/cyankiwi/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit
gibberish still persists?
And i have problems - model answers is hallucinating - i have very emotional answers with different themes mixed out but i have standart assistent promt. Do you have same problem ?
Thank you for your input! This might be the result of the active paramsshared_expert being quantized into 4bit, which is due to vllm not being able to load load the model if shared_expert is not quantized.
My apologies for this, but there will be an update to the model soon in the next few days to improve the model accuracy.
thanks for you work looking forward to use your AWQ instruct model!
Thank you for your input! This might be the result of the active params
shared_expertbeing quantized into 4bit, which is due to vllm not being able to load load the model ifshared_expertis not quantized.My apologies for this, but there will be an update to the model soon in the next few days to improve the model accuracy.
Hey! Any updates ? What is your plans?