Nex-N2.5-mini-oQ5 / README.md
TokenAI-zer's picture
Update references after account rename
0a47152 verified
|
Raw History Blame Contribute Delete
2.9 kB
---
base_model: nex-agi/Nex-N2.5-mini
model_name: Nex-N2.5-mini-oQ5
library_name: mlx
pipeline_tag: image-text-to-text
license: apache-2.0
tags:
- mlx
- omlx
- quantization
- mixed-precision
- apple-silicon
- moe
- vision
- base_model:quantized:nex-agi/Nex-N2.5-mini
---
# Nex-N2.5-mini-oQ5
Unofficial MLX quantization of [nex-agi/Nex-N2.5-mini](https://huggingface.co/nex-agi/Nex-N2.5-mini) for Apple Silicon. The upstream model is a multimodal mixture-of-experts model; this repository contains MLX safetensors, not GGUF or PyTorch weights. I am not affiliated with Nex AGI.
## What is in this repository
| Property | Value |
|---|---|
| Architecture | Qwen3_5MoeForConditionalGeneration |
| Quantization | affine, group size 64; 5-bit default with 6-bit and 8-bit module overrides |
| Weight size | 23.57 GiB (25.30 GB), 5 safetensors shards |
| Text model | 40 layers, 256 experts, 8 experts selected per token |
| Vision | vision tensors retained in BF16 |
| Context limit in config | 262,144 tokens; usable context depends on available memory |
| MTP | not present (mtp_num_hidden_layers: 0) |
The numbers above were read from the shipped config.json, model.safetensors.index.json and safetensors headers. The five shards contain 2,010 indexed tensors. The quantization recipe is recorded in config.json so a compatible MLX loader can reconstruct the per-module precision.
This conversion has not been benchmarked against the upstream BF16 model. The upstream benchmark figures on its [model card](https://huggingface.co/nex-agi/Nex-N2.5-mini) are not results for these quantized weights. Quantization can change output quality, and memory use grows with context length and cache settings.
## Usage
Download the model into your oMLX model directory:
~~~bash
hf download TokenAI-zer/Nex-N2.5-mini-oQ5 --local-dir ~/.omlx/models/Nex-N2.5-mini-oQ5
omlx serve --model-dir ~/.omlx/models --port 8000
~~~
Use the model ID Nex-N2.5-mini-oQ5 in oMLX. This architecture includes a vision tower, so use an MLX runtime with Qwen3.5 MoE multimodal support. The 262k context value is an architecture limit, not a promise that it will fit in memory.
## License and attribution
The [upstream repository](https://huggingface.co/nex-agi/Nex-N2.5-mini) declares Apache License 2.0. The [published model](https://huggingface.co/TokenAI-zer/Nex-N2.5-mini-oQ5) includes the [license text](LICENSE) and a [notice](https://huggingface.co/TokenAI-zer/Nex-N2.5-mini-oQ5/blob/main/NOTICE) identifying the source and the quantization change. The upstream model and its reported evaluations belong to Nex AGI.
## Citation
~~~bibtex
@misc{nex-n25-mini-oq5,
title = {Nex-N2.5-mini-oQ5: MLX quantization of Nex-N2.5-mini},
author = {TokenAI-zer},
year = {2026},
url = {https://huggingface.co/TokenAI-zer/Nex-N2.5-mini-oQ5},
note = {Unofficial quantization of nex-agi/Nex-N2.5-mini}
}
~~~