--- license: apache-2.0 base_model: Agnes-AI/Agnes-3.0-Flash library_name: mlx pipeline_tag: text-generation tags: - mlx - qwen3_5 - agnes - 4bit - hybrid-attention - gated-delta-net language: - en - zh --- # Agnes-3.0-Flash — MLX 4-bit 4-bit MLX quantization of [Agnes-AI/Agnes-3.0-Flash](https://huggingface.co/Agnes-AI/Agnes-3.0-Flash) (Apache-2.0) — loads as a stock `qwen3_5` model with no custom code. - Affine 4-bit, group size 64 — 4.50 bits/weight, 17 GB - 262,144-token context, thinking on/off via the original chat template - Text only: MTP head and vision tower not included ## What changed Converted from the original Agnes format to standard Qwen3.5 architecture: - Folded parallel FFN into main MLP via concatenation (intermediate_size: 19456) - Renamed `delta_attn` → `linear_attn`, `global_attn` → `self_attn` - Converted one-centered RMSNorm to standard format - Cast bf16 → fp16 for serialization compatibility - Stripped MTP weights (prevents double-conversion in mlx_lm) ## Usage ```bash pip install mlx-lm mlx_lm.generate --model hermitdave/Agnes-3.0-Flash-MLX-4bit --prompt "Hello" --max-tokens 200 ``` Drop the folder under `~/.lmstudio/models/hermitdave/` in LM Studio and it appears as a `qwen3_5` model. ## Speculative decoding with MTP drafter A companion MTP drafter is available for speculative decoding (up to 2× faster generation): ```bash pip install mlx-vlm python -m mlx_vlm.server \ --model hermitdave/Agnes-3.0-Flash-MLX-4bit \ --draft-model hermitdave/Agnes-3.0-Flash-MTP-drafter ``` Or with the Python API: ```python from mlx_lm import load from mlx_vlm.speculative.drafters.qwen3_5_mtp.config import Qwen3_5MTPConfig from mlx_vlm.speculative.drafters.qwen3_5_mtp.qwen3_5_mtp import Qwen3_5MTPDraftModel import json, mlx.core as mx, safetensors.torch base_model, tokenizer = load("hermitdave/Agnes-3.0-Flash-MLX-4bit") config = Qwen3_5MTPConfig.from_dict(json.load(open("path/to/Agnes-3.0-Flash-MTP-drafter/config.json"))) mtp = Qwen3_5MTPDraftModel(config) weights = safetensors.torch.load_file("path/to/Agnes-3.0-Flash-MTP-drafter/model.safetensors") mtp.load_weights([(k, mx.array(v)) for k, v in weights.items()]) mtp.bind(base_model) ``` The drafter is architecture-compatible with Qwen3.5's gated attention and was extracted from the original Agnes-3.0-Flash model. See the [MTP drafter repo](https://huggingface.co/hermitdave/Agnes-3.0-Flash-MTP-drafter) for details. ## Attribution This conversion was produced by [Hermes Agent](https://hermes-agent.nousresearch.com) (Nous Research) — the autonomous research and conversion pipeline that identified the correct quantization parameters, fixed one-centered norm conversion, and validated output quality. Verified against the reference verison/Agnes-3.0-Flash-MLX-4bit model.