--- license: apache-2.0 base_model: Agnes-AI/Agnes-3.0-Flash library_name: mlx pipeline_tag: text-generation tags: - mlx - qwen3_5 - agnes - 8bit - hybrid-attention - gated-delta-net language: - en - zh --- # Agnes-3.0-Flash — MLX 8-bit 8-bit MLX quantization of [Agnes-AI/Agnes-3.0-Flash](https://huggingface.co/Agnes-AI/Agnes-3.0-Flash) (Apache-2.0) — loads as a stock `qwen3_5` model with no custom code. - Affine 8-bit, group size 64 — 8.50 bits/weight, 32 GB - 262,144-token context, thinking on/off via the original chat template - Text only: MTP head and vision tower not included ## What changed Converted from the original Agnes format to standard Qwen3.5 architecture: - Folded parallel FFN into main MLP via concatenation (intermediate_size: 19456) - Renamed `delta_attn` → `linear_attn`, `global_attn` → `self_attn` - Converted one-centered RMSNorm to standard format - Cast bf16 → fp16 for serialization compatibility - Stripped MTP weights (prevents double-conversion in mlx_lm) ## Usage ```bash pip install mlx-lm mlx_lm.generate --model hermitdave/Agnes-3.0-Flash-MLX-8bit --prompt "Hello" --max-tokens 200 ``` Drop the folder under `~/.lmstudio/models/hermitdave/` in LM Studio and it appears as a `qwen3_5` model. ## Attribution This conversion was produced by [Hermes Agent](https://hermes-agent.nousresearch.com) (Nous Research) — the autonomous research and conversion pipeline that identified the correct quantization parameters, fixed one-centered norm conversion, and validated output quality. Verified against the reference verison/Agnes-3.0-Flash-MLX-4bit model.