File size: 1,661 Bytes
3fd5a38
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
---
license: apache-2.0
language: en
pipeline_tag: text-generation
library_name: mlx
base_model: nex-agi/Nex-N2-mini
tags:
- mlx
---

# Nex-N2-mini, 8-bit MLX

This is [nex-agi/Nex-N2-mini](https://huggingface.co/nex-agi/Nex-N2-mini) converted
to MLX format and quantized to 8 bits (group size 64) with mlx-lm 0.31.3.

Nex-N2-mini is an agentic model built around what its authors call Agentic Thinking:
it interleaves reasoning, tool use, and environment feedback rather than treating them
as separate stages. The architecture is a hybrid MoE (qwen3_5_moe): 40 layers
alternating linear attention with full attention every fourth layer, 256 experts with
8 active per token, and a 262k-token context window.

The original checkpoint includes a vision tower. MLX text inference does not use it,
so the vision weights were dropped during conversion; this copy is text-only. Expect
roughly 37 GB of memory in use during inference.

## Usage

With mlx-lm, either directly:

```sh
mlx_lm.generate --model jedisct1/Nex-N2-mini-mlx-8bit --prompt "Hello"
```

or as an OpenAI-compatible server:

```sh
mlx_lm.server --model jedisct1/Nex-N2-mini-mlx-8bit
```

It also works out of the box with oMLX.

Tool calling works without any extra configuration. The chat template uses the
Qwen3-Coder XML style, which mlx-lm and oMLX both detect automatically, so servers
return proper structured `tool_calls`, and thinking ends up in the reasoning field
instead of leaking into the response content. Tested end to end with
[Swival](https://swival.dev/) as the harness, including multi-step tasks that
exercise file edits, search, and shell commands while the model is thinking.