Agnes-3.0-Flash — MLX 4-bit

4-bit MLX quantization of Agnes-AI/Agnes-3.0-Flash (Apache-2.0) — loads as a stock qwen3_5 model with no custom code.

  • Affine 4-bit, group size 64 — 4.50 bits/weight, 17 GB
  • 262,144-token context, thinking on/off via the original chat template
  • Text only: MTP head and vision tower not included

What changed

Converted from the original Agnes format to standard Qwen3.5 architecture:

  • Folded parallel FFN into main MLP via concatenation (intermediate_size: 19456)
  • Renamed delta_attn → linear_attn, global_attn → self_attn
  • Converted one-centered RMSNorm to standard format
  • Cast bf16 → fp16 for serialization compatibility
  • Stripped MTP weights (prevents double-conversion in mlx_lm)

Usage

pip install mlx-lm
mlx_lm.generate --model hermitdave/Agnes-3.0-Flash-MLX-4bit --prompt "Hello" --max-tokens 200

Drop the folder under ~/.lmstudio/models/hermitdave/ in LM Studio and it appears as a qwen3_5 model.

Speculative decoding with MTP drafter

A companion MTP drafter is available for speculative decoding (up to 2× faster generation):

pip install mlx-vlm
python -m mlx_vlm.server \
    --model hermitdave/Agnes-3.0-Flash-MLX-4bit \
    --draft-model hermitdave/Agnes-3.0-Flash-MTP-drafter

Or with the Python API:

from mlx_lm import load
from mlx_vlm.speculative.drafters.qwen3_5_mtp.config import Qwen3_5MTPConfig
from mlx_vlm.speculative.drafters.qwen3_5_mtp.qwen3_5_mtp import Qwen3_5MTPDraftModel
import json, mlx.core as mx, safetensors.torch

base_model, tokenizer = load("hermitdave/Agnes-3.0-Flash-MLX-4bit")
config = Qwen3_5MTPConfig.from_dict(json.load(open("path/to/Agnes-3.0-Flash-MTP-drafter/config.json")))
mtp = Qwen3_5MTPDraftModel(config)
weights = safetensors.torch.load_file("path/to/Agnes-3.0-Flash-MTP-drafter/model.safetensors")
mtp.load_weights([(k, mx.array(v)) for k, v in weights.items()])
mtp.bind(base_model)

The drafter is architecture-compatible with Qwen3.5's gated attention and was extracted from the original Agnes-3.0-Flash model. See the MTP drafter repo for details.

Note: oMLX does not yet support MTP for text-only models. Use mlx_vlm.server or the Python API.

⚠️ oMLX users: the shipped MTP drafter does not work as a DFlash drafter

The companion Agnes MTP drafter cannot be paired with this model for oMLX DFlash speculative decoding. Its config.json nests the model fields under text_config (qwen3_5_mtp convention), so oMLX's drafter construction fails and it silently falls back to plain batched decoding — no error in the UI, no speedup:

DFlash start failed: DFlashDraftModelArgs.__init__() missing 11 required
positional arguments: 'hidden_size', 'num_hidden_layers', ...

Use z-lab/Qwen3.8-27B-DFlash2 as the oMLX DFlash drafter instead — same tokenizer (vocab 248,320), same hidden size (5,120), same qwen3_5-family architecture. Download it via the oMLX model browser (lands in ~/.omlx/models/z-lab/Qwen3.8-27B-DFlash2), then enable DFlash in the model's settings, or via the admin API:

curl -X PUT http://127.0.0.1:8000/admin/api/models/<model_id>/settings \
    -H "Content-Type: application/json" \
    -H "X-Api-Key: <your-api-key>" \
    -d '{
        "dflash_enabled": true,
        "dflash_draft_model": "~/.omlx/models/z-lab/Qwen3.8-27B-DFlash2"
    }'

Measured (M3 Max 64 GB, oMLX, temp 0, 500-token code generations): plain batched 16.5–17.2 tok/s → DFlash with the Qwen3.8 drafter 24.3–24.6 tok/s (+45%).

Verify engagement in ~/.omlx/logs/server.log: you want DFlashEngine loaded, not DFlash start failed ... fallback from DFlash.

Attribution

This conversion was produced by Hermes Agent (Nous Research) — the autonomous research and conversion pipeline that identified the correct quantization parameters, fixed one-centered norm conversion, and validated output quality. Verified against the reference verison/Agnes-3.0-Flash-MLX-4bit model.

Downloads last month
213
Safetensors
Model size
32B params
Tensor type
U32
·
BF16
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hermitdave/Agnes-3.0-Flash-MLX-4bit

Quantized
(17)
this model