--- license: apache-2.0 base_model: Agnes-AI/Agnes-3.0-Flash library_name: mlx pipeline_tag: text-generation tags: - mlx - qwen3_5 - agnes - 4bit - hybrid-attention - gated-delta-net language: - en - zh --- # Agnes-3.0-Flash — MLX 4-bit 4-bit MLX quantization of [Agnes-AI/Agnes-3.0-Flash](https://huggingface.co/Agnes-AI/Agnes-3.0-Flash) (Apache-2.0) — loads as a stock `qwen3_5` model with no custom code. - Affine 4-bit, group size 64 — 4.50 bits/weight, 17 GB - 262,144-token context, thinking on/off via the original chat template - Text only: MTP head and vision tower not included ## What changed Converted from the original Agnes format to standard Qwen3.5 architecture: - Folded parallel FFN into main MLP via concatenation (intermediate_size: 19456) - Renamed `delta_attn` → `linear_attn`, `global_attn` → `self_attn` - Converted one-centered RMSNorm to standard format - Cast bf16 → fp16 for serialization compatibility - Stripped MTP weights (prevents double-conversion in mlx_lm) ## Usage ```bash pip install mlx-lm mlx_lm.generate --model hermitdave/Agnes-3.0-Flash-MLX-4bit --prompt "Hello" --max-tokens 200 ``` Drop the folder under `~/.lmstudio/models/hermitdave/` in LM Studio and it appears as a `qwen3_5` model. ## Speculative decoding with MTP drafter A companion MTP drafter is available for speculative decoding (up to 2× faster generation): ```bash pip install mlx-vlm python -m mlx_vlm.server \ --model hermitdave/Agnes-3.0-Flash-MLX-4bit \ --draft-model hermitdave/Agnes-3.0-Flash-MTP-drafter ``` Or with the Python API: ```python from mlx_lm import load from mlx_vlm.speculative.drafters.qwen3_5_mtp.config import Qwen3_5MTPConfig from mlx_vlm.speculative.drafters.qwen3_5_mtp.qwen3_5_mtp import Qwen3_5MTPDraftModel import json, mlx.core as mx, safetensors.torch base_model, tokenizer = load("hermitdave/Agnes-3.0-Flash-MLX-4bit") config = Qwen3_5MTPConfig.from_dict(json.load(open("path/to/Agnes-3.0-Flash-MTP-drafter/config.json"))) mtp = Qwen3_5MTPDraftModel(config) weights = safetensors.torch.load_file("path/to/Agnes-3.0-Flash-MTP-drafter/model.safetensors") mtp.load_weights([(k, mx.array(v)) for k, v in weights.items()]) mtp.bind(base_model) ``` The drafter is architecture-compatible with Qwen3.5's gated attention and was extracted from the original Agnes-3.0-Flash model. See the [MTP drafter repo](https://huggingface.co/hermitdave/Agnes-3.0-Flash-MTP-drafter) for details. **Note:** oMLX does not yet support MTP for text-only models. Use `mlx_vlm.server` or the Python API. ### ⚠️ oMLX users: the shipped MTP drafter does not work as a DFlash drafter The companion Agnes MTP drafter **cannot be paired with this model for oMLX DFlash speculative decoding**. Its `config.json` nests the model fields under `text_config` (qwen3_5_mtp convention), so oMLX's drafter construction fails and it **silently falls back to plain batched decoding** — no error in the UI, no speedup: ``` DFlash start failed: DFlashDraftModelArgs.__init__() missing 11 required positional arguments: 'hidden_size', 'num_hidden_layers', ... ``` **Use [z-lab/Qwen3.8-27B-DFlash2](https://huggingface.co/z-lab/Qwen3.8-27B-DFlash2) as the oMLX DFlash drafter instead** — same tokenizer (vocab 248,320), same hidden size (5,120), same qwen3_5-family architecture. Download it via the oMLX model browser (lands in `~/.omlx/models/z-lab/Qwen3.8-27B-DFlash2`), then enable DFlash in the model's settings, or via the admin API: ```bash curl -X PUT http://127.0.0.1:8000/admin/api/models//settings \ -H "Content-Type: application/json" \ -H "X-Api-Key: " \ -d '{ "dflash_enabled": true, "dflash_draft_model": "~/.omlx/models/z-lab/Qwen3.8-27B-DFlash2" }' ``` **Measured** (M3 Max 64 GB, oMLX, temp 0, 500-token code generations): plain batched 16.5–17.2 tok/s → DFlash with the Qwen3.8 drafter **24.3–24.6 tok/s (+45%)**. Verify engagement in `~/.omlx/logs/server.log`: you want `DFlashEngine loaded`, not `DFlash start failed ... fallback from DFlash`. ## Attribution This conversion was produced by [Hermes Agent](https://hermes-agent.nousresearch.com) (Nous Research) — the autonomous research and conversion pipeline that identified the correct quantization parameters, fixed one-centered norm conversion, and validated output quality. Verified against the reference verison/Agnes-3.0-Flash-MLX-4bit model.