Instructions to use Qwen/Qwen3.8-2.4T-A95B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen3.8-2.4T-A95B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Qwen/Qwen3.8-2.4T-A95B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3.8-2.4T-A95B") model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.8-2.4T-A95B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- HuggingChat
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Qwen/Qwen3.8-2.4T-A95B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Qwen/Qwen3.8-2.4T-A95B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-2.4T-A95B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Qwen/Qwen3.8-2.4T-A95B
- SGLang
How to use Qwen/Qwen3.8-2.4T-A95B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3.8-2.4T-A95B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-2.4T-A95B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3.8-2.4T-A95B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-2.4T-A95B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Qwen/Qwen3.8-2.4T-A95B with Docker Model Runner:
docker model run hf.co/Qwen/Qwen3.8-2.4T-A95B
Weirdos
A thousand downloads, guys being crazy for this giant... do you guys have 50 B200s at home just roaming around! Bruh!
they are downloading in case open source AI gets banned, american companies have been crying a lot lately cause they get beat up by superior peoples, and they cant stand it, they'd rather scam you and not innovate at all if possible.
they are downloading in case open source AI gets banned, american companies have been crying a lot lately cause they get beat up by superior peoples, and they cant stand it, they'd rather scam you and not innovate at all if possible.
If its so whole of Huggingface has to get banned .. lol. But honestly, I feel like Claude is still one of the greatest and is continuing to be the greatest (as they realised their Fable 5 was too good to be actually a business model, that it will lead their company to huge losses and hence had to cut costs in Opus 5 keeping same performance.) They improved a lot after that 😂. I never got to try Kimi K3, but used GLM 5.2 and the latest of the Deepseek models, both are good but I think still are far away to match Claude performance, but as said (cost also matters😂)
they are downloading in case open source AI gets banned, american companies have been crying a lot lately cause they get beat up by superior peoples, and they cant stand it, they'd rather scam you and not innovate at all if possible.
If its so whole of Huggingface has to get banned .. lol. But honestly, I feel like Claude is still one of the greatest and is continuing to be the greatest (as they realised their Fable 5 was too good to be actually a business model, that it will lead their company to huge losses and hence had to cut costs in Opus 5 keeping same performance.) They improved a lot after that 😂. I never got to try Kimi K3, but used GLM 5.2 and the latest of the Deepseek models, both are good but I think still are far away to match Claude performance, but as said (cost also matters😂)
I find it funny how they are like "oh we cant let people use this ai its the end of the world if we do " then china releases it open source at almost same performance and nothing happens, the world is still here. People should appreciate China more with what it did for the AI space, and im saying this as an European.
Stop judging everything with your narrow-minded perspective. Isn't it possible that the mere existence of something is already of great importance?