File size: 2,915 Bytes
777498f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1f1da88
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14e85ef
 
777498f
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
---
license: apache-2.0
base_model: Agnes-AI/Agnes-3.0-Flash
library_name: mlx
pipeline_tag: text-generation
tags:
- mlx
- qwen3_5
- agnes
- 4bit
- hybrid-attention
- gated-delta-net
language:
- en
- zh
---

# Agnes-3.0-Flash — MLX 4-bit

4-bit MLX quantization of [Agnes-AI/Agnes-3.0-Flash](https://huggingface.co/Agnes-AI/Agnes-3.0-Flash) (Apache-2.0) — loads as a stock `qwen3_5` model with no custom code.

- Affine 4-bit, group size 64 — 4.50 bits/weight, 17 GB
- 262,144-token context, thinking on/off via the original chat template
- Text only: MTP head and vision tower not included

## What changed

Converted from the original Agnes format to standard Qwen3.5 architecture:

- Folded parallel FFN into main MLP via concatenation (intermediate_size: 19456)
- Renamed `delta_attn` → `linear_attn`, `global_attn` → `self_attn`
- Converted one-centered RMSNorm to standard format
- Cast bf16 → fp16 for serialization compatibility
- Stripped MTP weights (prevents double-conversion in mlx_lm)

## Usage

```bash
pip install mlx-lm
mlx_lm.generate --model hermitdave/Agnes-3.0-Flash-MLX-4bit --prompt "Hello" --max-tokens 200
```

Drop the folder under `~/.lmstudio/models/hermitdave/` in LM Studio and it appears as a `qwen3_5` model.

## Speculative decoding with MTP drafter

A companion MTP drafter is available for speculative decoding (up to 2× faster generation):

```bash
pip install mlx-vlm
python -m mlx_vlm.server \
    --model hermitdave/Agnes-3.0-Flash-MLX-4bit \
    --draft-model hermitdave/Agnes-3.0-Flash-MTP-drafter
```

Or with the Python API:

```python
from mlx_lm import load
from mlx_vlm.speculative.drafters.qwen3_5_mtp.config import Qwen3_5MTPConfig
from mlx_vlm.speculative.drafters.qwen3_5_mtp.qwen3_5_mtp import Qwen3_5MTPDraftModel
import json, mlx.core as mx, safetensors.torch

base_model, tokenizer = load("hermitdave/Agnes-3.0-Flash-MLX-4bit")
config = Qwen3_5MTPConfig.from_dict(json.load(open("path/to/Agnes-3.0-Flash-MTP-drafter/config.json")))
mtp = Qwen3_5MTPDraftModel(config)
weights = safetensors.torch.load_file("path/to/Agnes-3.0-Flash-MTP-drafter/model.safetensors")
mtp.load_weights([(k, mx.array(v)) for k, v in weights.items()])
mtp.bind(base_model)
```

The drafter is architecture-compatible with Qwen3.5's gated attention and was extracted from the original Agnes-3.0-Flash model. See the [MTP drafter repo](https://huggingface.co/hermitdave/Agnes-3.0-Flash-MTP-drafter) for details.

**Note:** oMLX does not yet support MTP for text-only models. Use `mlx_vlm.server` or the Python API.
## Attribution

This conversion was produced by [Hermes Agent](https://hermes-agent.nousresearch.com) (Nous Research) — the autonomous research and conversion pipeline that identified the correct quantization parameters, fixed one-centered norm conversion, and validated output quality. Verified against the reference verison/Agnes-3.0-Flash-MLX-4bit model.