Instructions to use Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- vLLM
How to use Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM
- SGLang
How to use Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Unsloth Desktop
- Docker Model Runner
How to use Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM with Docker Model Runner:
docker model run hf.co/Xingyu-Zheng/Qwopus3.5-9B-v3-INT8-FOEM
language:
- en
- zh
license: apache-2.0
base_model:
- Qwen/Qwen3.5-9B
- Jackrong/Qwopus3.5-9B-v3
tags:
- unsloth
- qwen
- qwen3.5
- reasoning
- chain-of-thought
- Dense
- vLLM
- SGLang
pipeline_tag: image-text-to-text
datasets:
- nohurry/Opus-4.6-Reasoning-3000x-filtered
๐Qwopus3.5-9B-v3-INT4-FOEM
This is an unofficial quantized version of Qwopus3.5-9B-v3.
๐ง Quantization Framework
๐บ๏ธ Quantization Method
FOEM is an improved quantization method over GPTQ. The resulting model preserves the same inference structure as GPTQ, ensuring compatibility with existing deployment pipelines while achieving better accuracy.
๐ Calibration Dataset
We randomly sampled 512 examples from nohurry/Opus-4.6-Reasoning-3000x-filtered.
๐ Usage Example
This model can be deployed using standard frameworks such as vLLM, just like other GPTQModel-quantized models.
Example evaluation command:
lm-eval --model vllm --model_args pretrained=models/gptqmodel/Qwopus3.5-9B-v3-INT8-FOEM,tensor_parallel_size=1,gpu_memory_utilization=0.45 --tasks wikitext --batch_size 1
โ ๏ธ Limitations & Intended Use
(Adapted from the original repository of Jackrong/Qwopus3.5-9B-v3)
- Hallucination Risk: While reasoning is strong, the model remains an autoregressive LLM; external facts provided during the thinking sequence may occasionally contain hallucinations if verifying real-world events.
- Intended Scenario: Best suited for offline analytical tasks, coding, math, and heavy logic-dependent prompting where the user needs to transparently follow the AI's internal logic.
- This model is a test version intended solely for learning and demonstration purposes, and is for academic research and technical exploration use only.
- Developer Disclaimer: This is an independent, personal project. Since the developer lacks the specialized technical resources and infrastructure of a large-scale industrial lab, the model's reasoning chain (CoT) may occasionally exhibit instability, logic loops, or reasoning drift. Users are advised to use this model with these experimental limitations in mind.
๐ Acknowledgements
Special thanks to Jackrong for providing the original model: Qwopus3.5-9B-v3.
๐ Citation
If you use this model in your research or projects, please cite:
@misc{jackrong_qwen35_9b_v3
title = {Jackrong/Qwopus3.5-9B-v3},
author = {Jackrong},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Jackrong/Qwopus3.5-9B-v3}}
}
@misc{qubitium2024gptqmodel,
author = {ModelCloud.ai and qubitium@modelcloud.ai},
title = {GPT-QModel},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/modelcloud/gptqmodel}},
note = {Contact: qubitium@modelcloud.ai},
year = {2024},
}
@inproceedings{zheng2026first,
title={First-order error matters: Accurate compensation for quantized large language models},
author={Zheng, Xingyu and Qin, Haotong and Li, Yuye and Chu, Haoran and Wang, Jiakai and Guo, Jinyang and Magno, Michele and Liu, Xianglong},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
volume={40},
number={34},
pages={28883--28891},
year={2026}
}