ashxhart's picture
Rebrand model card to TensorFold
8eaa70e verified
|
Raw History Blame Contribute Delete
3.53 kB
---
library_name: mlx
license: apache-2.0
base_model: nex-agi/Nex-N2.5-mini
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
- mlx
- mlx-vlm
- omlx
- nex-agi
- nex-n2.5
- qwen3_5_moe
- mixture-of-experts
- vision-language
- apple-silicon
- quantized
- 6-bit
---
<p align="center">
<a href="https://tensorfold.dev">
<img src="https://huggingface.co/spaces/TensorFold/README/resolve/main/tensorfold-logo.png" alt="TensorFold" width="160">
</a>
</p>
<p align="center"><a href="https://nex-agi.com/"><img src="./assets/NEX_logo.svg" width="192" height="61" alt="Nex-AGI"></a></p>
<h1 align="center">Nex-N2.5 mini · MLX oQ6</h1>
<p align="center">A community mixed-precision conversion by <a href="https://huggingface.co/TensorFold">TensorFold</a> for Apple Silicon.</p>
## Model
Converted from the BF16 [Nex-N2.5-mini](https://huggingface.co/nex-agi/Nex-N2.5-mini) checkpoint using oMLX oQ6.
The oQ label is a target, not a claim that every tensor uses the same precision; see the per-module quantisation entries in config.json.
This text-and-vision model uses the qwen3_5_moe architecture.
The inspected BF16 checkpoint contained no matching MTP tensors, despite its configuration declaring one MTP layer.
This release does not provide tested MTP decoding; keep MTP disabled.
## Download and use
```bash
hf download TensorFold/Nex-N2.5-mini-MLX-oQ6 --local-dir ./Nex-N2.5-mini-MLX-oQ6
```
Add the folder to oMLX model directories, refresh the model list and select it.
Basic inference was tested with oMLX 0.6.4.
Use the [upstream-recommended sampling](https://github.com/nex-agi/Nex-N2.5):
```json
{
"temperature": 0.7,
"top_p": 0.95,
"top_k": 40
}
```
Set these explicitly in your client or model settings.
Our sampled oQ2 retests also used reasoning_effort="none" and max_tokens=4096; these are test conditions, not an upstream recommendation to disable reasoning.
Avoid greedy decoding for oQ2, which reproduced a repetition loop in our tests.
## Validation and limitations
Tested on 9 September 2026 through oMLX 0.6.4 on an Apple Silicon Studio with 256 GiB unified memory.
Exact arithmetic, a forced weather-tool call with a Paris argument, and identification of a synthetic red image passed for this quant.
Tool calls were checked for formatting, not executed.
These are basic checks, not a full coding, vision or agent evaluation.
The long coding response finished naturally and its main Python block parsed; generated code was not executed.
Controlled throughput benchmarks using recommended sampling have not been completed for this quant.
Peak request memory and context-fit limits have not been measured, so no Mac memory-tier recommendation is claimed.
Long-context, multi-turn and broader vision quality remain unverified.
## Short test
With the sampling settings above, try:
```text
What is 17 multiplied by 19? Answer with only the number.
```
The recorded arithmetic check returned `323` with temperature=0 and reasoning_effort="none".
A short correct response does not establish long-generation reliability.
## Licence and attribution
The upstream repository declares Apache-2.0.
Model training, architecture and the original Nex logo belong to Nex-AGI and the respective upstream contributors.
This is an independent community conversion, not an official Nex-AGI release.
Upstream benchmark scores are not evaluations of this quant.
[Follow TensorFold for new Apple Silicon releases and fixes.](https://huggingface.co/TensorFold)