--- license: apache-2.0 language: en base_model: - empero-ai/Qwen3.8-2B-Distill base_model_relation: quantized quantized_by: Atomic-Germ pipeline_tag: image-text-to-text tags: - transformers - qwen3_5_text - image-text-to-text - conversational - q4nx - quantized - npu2 - fastflowlm - fastflow - flm - empero-ai - qwen3.8 - distillation - reasoning --- # Qwen3.8-2B-Distill-NPU2 **FastFlowLM Q4NX conversion of [`empero-ai/Qwen3.8-2B-Distill`](https://huggingface.co/empero-ai/Qwen3.8-2B-Distill)** for AMD XDNA NPU inference. This repository contains a quantized **Q4NX** port of the model, compiled for the FastFlowLM (FLM) runtime. It is **not** a GGUF file. | Item | Value | |------|-------| | Source model | [`empero-ai/Qwen3.8-2B-Distill`](https://huggingface.co/empero-ai/Qwen3.8-2B-Distill) | | Source GGUF | `Qwen3.8-2B-Q8_0.gguf` | | Weights | `model.q4nx` (2.32 GB) | | Modality | language / vision | | FLM version | `1.0.1` | | Converted | 2026-08-24 | ## Install and run This repository works with `flm-add`, a small installer that copies the model into the FastFlowLM user directory and registers the tag. It never modifies the system FastFlowLM install. `pip install flm-add` or `uv tool install flm-add` ```bash uv tool install flm-add flm-add Atomic-Germ/Qwen3.8-2B-Distill-NPU2 --family qwen3.5 --tag qwen3.8-distill:2b FLM_CONFIG_PATH="$HOME/.config/flm/model_list.json" FLM_XCLBIN_PATH="$HOME/.config/flm" flm run qwen3.8-distill:2b ``` ## Files | File | Description | |------|-------------| | `model.q4nx` | Quantized weights (Q8_0 / Q4_1 / BF16) | | `config.json` | FLM runtime configuration | | `tokenizer.json` | Tokenizer vocabulary | | `tokenizer_config.json` | Tokenizer configuration | | `chat_template.jinja` | Chat template | | `vision_weight.q4nx` | Vision model | --- ## Source model card See the original model card: [empero-ai/Qwen3.8-2B-Distill](https://huggingface.co/empero-ai/Qwen3.8-2B-Distill)