MiMo-V2.6-Distill-Qwen-9B-NPU2

This repository contains a Q4NX quantization of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B, converted for hardware-accelerated inference with FastFlowLM on AMD Ryzen AI (XDNA2) NPUs.

Model Details

  • Format: Q4NX (FastFlowLM's native packed-quantization format)
  • Original Model: Qwen 3.5 9B Fine-tune
  • Runtime: FastFlowLM
  • Hardware: AMD Ryzen AI (Strix Point / XDNA2)
Downloads last month
18
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Soveu/MiMo-V2.6-Distill-Qwen-9B-NPU2

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(20)
this model