Qwen-3.8-9B-nf4-cw-ov

This is an OpenVINO NF4 weight-only quantized (Compressed Weights) version of empero-ai/Qwen3.8-9B-Distill, optimized for Intel NPU acceleration and deployment via OpenVINO Model Server (OVMS).

Model Details

  • Base Model: empero-ai/Qwen3.8-9B-Distill
  • Quantization: NF4 Compressed Weights (via Intel NNCF)
  • Target Hardware: Intel NPU 4 (Lunar Lake Architecture and newer)
  • Deployment Framework: OpenVINO Model Server (OVMS) / OpenVINO C++ Runtime

Compatibility & System Requirements

Hardware Requirements

  • NPU Generation: Intel NPU 4 (found in Intel Core Ultra 200V series / Lunar Lake) or newer.
  • Note: Legacy NPUs (NPU 3 / Meteor Lake) or older iGPU/CPU-only environments may not support this specific NF4 weight-compression layout without fallback to CPU runtime.

For detailed architecture info, evaluation benchmarks, and original weights, refer to the base repository: **[empero-ai/Qwen3.8-9B-Distill](https://huggingface.co/empero-ai/Qwen3.8-9B-Distill

More techinical detail

Reccomend Enviroment and techinical detail:My blog(※Japanese)

Downloads last month
32
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for asahi-jp/Qwen-3.8-9B-nf4-cw-ov

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(978)
this model

Collection including asahi-jp/Qwen-3.8-9B-nf4-cw-ov