--- license: other license_name: tongyi-qianwen license_link: https://huggingface.co base_model: Qwen/Qwen1.5-MoE-A2.7B tags: - text-generation - qwen2_moe - mixture-of-experts - conversational - pretrained --- HF-> F16.gguf converted with a fresh kaggle cookbook :] * I am moving the other quants into this repo as well.. --- # Qwen1.5-MoE-A2.7B ## Introduction Qwen1.5-MoE is a transformer-based MoE decoder-only language model pretrained on a large amount of data. For more details, please refer to our [blog post](https://qwenlm.github.io/blog/qwen-moe/) and [GitHub repo](https://github.com/QwenLM/Qwen1.5). ## Model Details Qwen1.5-MoE employs Mixture of Experts (MoE) architecture, where the models are upcycled from dense language models. For instance, `Qwen1.5-MoE-A2.7B` is upcycled from `Qwen-1.8B`. It has 14.3B parameters in total and 2.7B activated parameters during runtime, while achieving comparable performance to `Qwen1.5-7B`, it only requires 25% of the training resources. We also observed that the inference speed is 1.74 times that of `Qwen1.5-7B`. ## Requirements The code of Qwen1.5-MoE has been in the latest Hugging face transformers and we advise you to build from source with command `pip install git+https://github.com/huggingface/transformers`, or you might encounter the following error: ``` KeyError: 'qwen2_moe'. ``` ## Usage markdown 📊 Quantization Comparison | Quantization | File Size | Recommendation | |--------------|-----------|------------------------------------| | F16 | ~28.6 GB | Best precision (for high-end GPUs) | | Q5_K_M | | Best balance of logic and size | | Q4_K_M | ~9.5 GB | Recommended for most users | | IQ4_XS | | Best for low-RAM devices | --- * Adding The FNG's # TULLUS/Qwen1.5-MoE-A2.7B-Q4_K_M-GGUF This model was converted to GGUF format from [`TULLUS/Qwen1.5-MoE-A2.7B`](https://huggingface.co/TULLUS/Qwen1.5-MoE-A2.7B) using llama.cpp via the ggml.ai's [GGUF-my-repo](https://huggingface.co/spaces/ggml-org/gguf-my-repo) space. Refer to the [original model card](https://huggingface.co/TULLUS/Qwen1.5-MoE-A2.7B) for more details on the model. ## Use with llama.cpp Install llama.cpp through brew (works on Mac and Linux) ```bash brew install llama.cpp ``` Invoke the llama.cpp server or the CLI. ### CLI: ```bash llama-cli --hf-repo TULLUS/Qwen1.5-MoE-A2.7B-Q4_K_M-GGUF --hf-file qwen1.5-moe-a2.7b-q4_k_m.gguf -p "The meaning to life and the universe is" ``` ### Server: