--- base_model: nex-agi/Nex-N2.5-mini model_name: Nex-N2.5-mini-oQ5 library_name: mlx pipeline_tag: image-text-to-text license: apache-2.0 tags: - mlx - omlx - quantization - mixed-precision - apple-silicon - moe - vision - base_model:quantized:nex-agi/Nex-N2.5-mini --- # Nex-N2.5-mini-oQ5 Unofficial MLX quantization of [nex-agi/Nex-N2.5-mini](https://huggingface.co/nex-agi/Nex-N2.5-mini) for Apple Silicon. The upstream model is a multimodal mixture-of-experts model; this repository contains MLX safetensors, not GGUF or PyTorch weights. I am not affiliated with Nex AGI. ## What is in this repository | Property | Value | |---|---| | Architecture | Qwen3_5MoeForConditionalGeneration | | Quantization | affine, group size 64; 5-bit default with 6-bit and 8-bit module overrides | | Weight size | 23.57 GiB (25.30 GB), 5 safetensors shards | | Text model | 40 layers, 256 experts, 8 experts selected per token | | Vision | vision tensors retained in BF16 | | Context limit in config | 262,144 tokens; usable context depends on available memory | | MTP | not present (mtp_num_hidden_layers: 0) | The numbers above were read from the shipped config.json, model.safetensors.index.json and safetensors headers. The five shards contain 2,010 indexed tensors. The quantization recipe is recorded in config.json so a compatible MLX loader can reconstruct the per-module precision. This conversion has not been benchmarked against the upstream BF16 model. The upstream benchmark figures on its [model card](https://huggingface.co/nex-agi/Nex-N2.5-mini) are not results for these quantized weights. Quantization can change output quality, and memory use grows with context length and cache settings. ## Usage Download the model into your oMLX model directory: ~~~bash hf download TokenAI-zer/Nex-N2.5-mini-oQ5 --local-dir ~/.omlx/models/Nex-N2.5-mini-oQ5 omlx serve --model-dir ~/.omlx/models --port 8000 ~~~ Use the model ID Nex-N2.5-mini-oQ5 in oMLX. This architecture includes a vision tower, so use an MLX runtime with Qwen3.5 MoE multimodal support. The 262k context value is an architecture limit, not a promise that it will fit in memory. ## License and attribution The [upstream repository](https://huggingface.co/nex-agi/Nex-N2.5-mini) declares Apache License 2.0. The [published model](https://huggingface.co/TokenAI-zer/Nex-N2.5-mini-oQ5) includes the [license text](LICENSE) and a [notice](https://huggingface.co/TokenAI-zer/Nex-N2.5-mini-oQ5/blob/main/NOTICE) identifying the source and the quantization change. The upstream model and its reported evaluations belong to Nex AGI. ## Citation ~~~bibtex @misc{nex-n25-mini-oq5, title = {Nex-N2.5-mini-oQ5: MLX quantization of Nex-N2.5-mini}, author = {TokenAI-zer}, year = {2026}, url = {https://huggingface.co/TokenAI-zer/Nex-N2.5-mini-oQ5}, note = {Unofficial quantization of nex-agi/Nex-N2.5-mini} } ~~~