hermitdave's picture
Add oMLX MTP limitation note
14e85ef verified
|
Raw History Blame
2.92 kB
---
license: apache-2.0
base_model: Agnes-AI/Agnes-3.0-Flash
library_name: mlx
pipeline_tag: text-generation
tags:
- mlx
- qwen3_5
- agnes
- 4bit
- hybrid-attention
- gated-delta-net
language:
- en
- zh
---
# Agnes-3.0-Flash β€” MLX 4-bit
4-bit MLX quantization of [Agnes-AI/Agnes-3.0-Flash](https://huggingface.co/Agnes-AI/Agnes-3.0-Flash) (Apache-2.0) β€” loads as a stock `qwen3_5` model with no custom code.
- Affine 4-bit, group size 64 β€” 4.50 bits/weight, 17 GB
- 262,144-token context, thinking on/off via the original chat template
- Text only: MTP head and vision tower not included
## What changed
Converted from the original Agnes format to standard Qwen3.5 architecture:
- Folded parallel FFN into main MLP via concatenation (intermediate_size: 19456)
- Renamed `delta_attn` β†’ `linear_attn`, `global_attn` β†’ `self_attn`
- Converted one-centered RMSNorm to standard format
- Cast bf16 β†’ fp16 for serialization compatibility
- Stripped MTP weights (prevents double-conversion in mlx_lm)
## Usage
```bash
pip install mlx-lm
mlx_lm.generate --model hermitdave/Agnes-3.0-Flash-MLX-4bit --prompt "Hello" --max-tokens 200
```
Drop the folder under `~/.lmstudio/models/hermitdave/` in LM Studio and it appears as a `qwen3_5` model.
## Speculative decoding with MTP drafter
A companion MTP drafter is available for speculative decoding (up to 2Γ— faster generation):
```bash
pip install mlx-vlm
python -m mlx_vlm.server \
--model hermitdave/Agnes-3.0-Flash-MLX-4bit \
--draft-model hermitdave/Agnes-3.0-Flash-MTP-drafter
```
Or with the Python API:
```python
from mlx_lm import load
from mlx_vlm.speculative.drafters.qwen3_5_mtp.config import Qwen3_5MTPConfig
from mlx_vlm.speculative.drafters.qwen3_5_mtp.qwen3_5_mtp import Qwen3_5MTPDraftModel
import json, mlx.core as mx, safetensors.torch
base_model, tokenizer = load("hermitdave/Agnes-3.0-Flash-MLX-4bit")
config = Qwen3_5MTPConfig.from_dict(json.load(open("path/to/Agnes-3.0-Flash-MTP-drafter/config.json")))
mtp = Qwen3_5MTPDraftModel(config)
weights = safetensors.torch.load_file("path/to/Agnes-3.0-Flash-MTP-drafter/model.safetensors")
mtp.load_weights([(k, mx.array(v)) for k, v in weights.items()])
mtp.bind(base_model)
```
The drafter is architecture-compatible with Qwen3.5's gated attention and was extracted from the original Agnes-3.0-Flash model. See the [MTP drafter repo](https://huggingface.co/hermitdave/Agnes-3.0-Flash-MTP-drafter) for details.
**Note:** oMLX does not yet support MTP for text-only models. Use `mlx_vlm.server` or the Python API.
## Attribution
This conversion was produced by [Hermes Agent](https://hermes-agent.nousresearch.com) (Nous Research) β€” the autonomous research and conversion pipeline that identified the correct quantization parameters, fixed one-centered norm conversion, and validated output quality. Verified against the reference verison/Agnes-3.0-Flash-MLX-4bit model.