Instructions to use sepsy070716/Qwen3.5-4B-A3B-Upcycled with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sepsy070716/Qwen3.5-4B-A3B-Upcycled with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="sepsy070716/Qwen3.5-4B-A3B-Upcycled") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("sepsy070716/Qwen3.5-4B-A3B-Upcycled") model = AutoModelForMultimodalLM.from_pretrained("sepsy070716/Qwen3.5-4B-A3B-Upcycled", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sepsy070716/Qwen3.5-4B-A3B-Upcycled with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sepsy070716/Qwen3.5-4B-A3B-Upcycled" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sepsy070716/Qwen3.5-4B-A3B-Upcycled", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/sepsy070716/Qwen3.5-4B-A3B-Upcycled
- SGLang
How to use sepsy070716/Qwen3.5-4B-A3B-Upcycled with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sepsy070716/Qwen3.5-4B-A3B-Upcycled" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sepsy070716/Qwen3.5-4B-A3B-Upcycled", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sepsy070716/Qwen3.5-4B-A3B-Upcycled" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sepsy070716/Qwen3.5-4B-A3B-Upcycled", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use sepsy070716/Qwen3.5-4B-A3B-Upcycled with Docker Model Runner:
docker model run hf.co/sepsy070716/Qwen3.5-4B-A3B-Upcycled
Qwen3.5-4B-A3B-Upcycled
Experimental sparse-MoE initialization derived from Qwen/Qwen3.5-4B.
Research checkpoint, not a trained release. The dense FFNs were compressed and the experts have not undergone continued pretraining or distillation. Expect quality loss relative to the base model. Do not treat benchmark results from the base model as results for this checkpoint.
Architecture
| Property | Value |
|---|---|
| Total parameters | 4,036,686,336 |
| Active parameters including vision | 2,998,596,096 |
| Experts per layer | 8 |
| Experts selected per token | 2 |
| Shared expert width | 1536 |
| Routed expert width | 704 |
The original 9,216-wide dense FFN is reduced to a 1,536-wide shared expert and eight 704-wide routed experts with top-2 routing. Neurons are ranked per layer by the product of their gate, up, and down projection norms. The strongest shared and routed slices are retained. Routed experts start identically so the untrained router does not make the initial function nondeterministic.
Intended use
This checkpoint is intended as an initialization for router warm-up, knowledge
distillation from Qwen/Qwen3.5-4B, and continued pretraining. It is not
recommended for production or user-facing inference before recovery training
and evaluation.
Load with a Transformers release that provides Qwen3_5MoeForConditionalGeneration:
from transformers import Qwen3_5MoeForConditionalGeneration, AutoProcessor
model = Qwen3_5MoeForConditionalGeneration.from_pretrained(
"sepsy070716/Qwen3.5-4B-A3B-Upcycled",
device_map="auto",
)
processor = AutoProcessor.from_pretrained("sepsy070716/Qwen3.5-4B-A3B-Upcycled")
See conversion_manifest.json and neuron_selection.json for reproducibility.
Router warm-up v1
A router-only MPS warm-up artifact is published under
research/router-warmup-v1/. It updates 655,360 router parameters and leaves
all attention, expert, embedding, and vision weights untouched.
On 32 held-out FineWeb2 Korean documents (8,109 valid tokens), normalized
routing entropy improved from 0.99599 to 0.99748, while the maximum/minimum
expert usage ratio improved from 1.528 to 1.390. This adapter balances
routing but does not recover the quality lost by dense-FFN compression;
expert distillation is still required.
Layerwise distillation pilots
Accepted layer adapters for depths 0, 16, and 31 are published under
research/layer-distillation-pilots/. Each was trained for 10 local updates
with the v1 router frozen. On eight held-out Korean documents, dense-FFN
relative MSE improved by 4.07%, 8.24%, and 8.33% respectively, with improvement
on all 24 document/layer comparisons.
These local reconstruction results are proof of direction, not an end-to-end model benchmark. The root checkpoint has not been modified by the pilot adapters.
- Downloads last month
- 17