MiMo-V2.6-Distill-Qwen-9B-NPU2
This repository contains a Q4NX quantization of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B, converted for hardware-accelerated inference with FastFlowLM on AMD Ryzen AI (XDNA2) NPUs.
Model Details
- Format: Q4NX (FastFlowLM's native packed-quantization format)
- Original Model: Qwen 3.5 9B Fine-tune
- Runtime: FastFlowLM
- Hardware: AMD Ryzen AI (Strix Point / XDNA2)
- Downloads last month
- 18
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for Soveu/MiMo-V2.6-Distill-Qwen-9B-NPU2
Base model
Qwen/Qwen3.5-9B-Base Finetuned
Qwen/Qwen3.5-9B Finetuned
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B