--- language: - en - zh license: apache-2.0 base_model: - Qwen/Qwen3.5-9B - Jackrong/Qwopus3.5-9B-v3.5 tags: - unsloth - qwen - qwen3.5 - reasoning - chain-of-thought - Dense - vLLM - SGLang pipeline_tag: image-text-to-text datasets: - nohurry/Opus-4.6-Reasoning-3000x-filtered --- # 🌟Qwopus3.5-9B-v3.5-INT4-FOEM
This is an unofficial quantized version of Qwopus3.5-9B-v3.5. ### 🧠 Quantization Framework [GPTQModel](https://github.com/ModelCloud/GPTQModel) ## πŸ—ΊοΈ Quantization Method [FOEM (AAAI 2026)](https://ojs.aaai.org/index.php/AAAI/article/view/40123) FOEM is an improved quantization method over GPTQ. The resulting model preserves the same inference structure as GPTQ, ensuring compatibility with existing deployment pipelines while achieving better accuracy. ### πŸ“š Calibration Dataset We randomly sampled 512 examples from [nohurry/Opus-4.6-Reasoning-3000x-filtered](https://huggingface.co/datasets/nohurry/Opus-4.6-Reasoning-3000x-filtered). ## πŸ“‹ Usage Example This model can be deployed using standard frameworks such as **vLLM**, just like other GPTQModel-quantized models. Example evaluation command: ```bash lm-eval --model vllm --model_args pretrained=models/gptqmodel/Qwopus3.5-9B-v3.5-INT4-FOEM,tensor_parallel_size=1,gpu_memory_utilization=0.45 --tasks wikitext --batch_size 1 ``` ## ⚠️ Limitations & Intended Use *(Adapted from the original repository of Jackrong/Qwopus3.5-9B-v3.5)* - Possible overfitting if scaling exceeds optimal regime - Reasoning may still exhibit instability in edge cases - Tool-calling performance depends on environment integration - Not all capabilities are fully benchmarked yet ## πŸ™ Acknowledgements Special thanks to [Jackrong](https://huggingface.co/Jackrong) for providing the original model: [Qwopus3.5-9B-v3.5](https://huggingface.co/Jackrong/Qwopus3.5-9B-v3.5). ## πŸ“– Citation If you use this model in your research or projects, please cite: ```bibtex @misc{jackrong_qwopus35_9b_v35, title = {Qwopus3.5-9B-v3.5}, author = {Jackrong}, year = {2026}, publisher = {Hugging Face} } ``` ```bibtex @misc{qubitium2024gptqmodel, author = {ModelCloud.ai and qubitium@modelcloud.ai}, title = {GPT-QModel}, publisher = {GitHub}, journal = {GitHub repository}, howpublished = {\url{https://github.com/modelcloud/gptqmodel}}, note = {Contact: qubitium@modelcloud.ai}, year = {2024}, } ``` ```bibtex @inproceedings{zheng2026first, title={First-order error matters: Accurate compensation for quantized large language models}, author={Zheng, Xingyu and Qin, Haotong and Li, Yuye and Chu, Haoran and Wang, Jiakai and Guo, Jinyang and Magno, Michele and Liu, Xianglong}, booktitle={Proceedings of the AAAI Conference on Artificial Intelligence}, volume={40}, number={34}, pages={28883--28891}, year={2026} } ```