--- library_name: mlx license: apache-2.0 license_link: https://huggingface.co/Qwen/Qwen3.5-397B-A17B/blob/main/LICENSE pipeline_tag: image-text-to-text tags: - mlx base_model: Qwen/Qwen3.5-397B-A17B --- # m-i/Qwen3.5-397B-A17B-2.592bit This model was converted to MLX format from [`Qwen/Qwen3.5-397B-A17B`](https://huggingface.co/Qwen/Qwen3.5-397B-A17B) using mlx-vlm version **0.4.4**. Refer to the [original model card](https://huggingface.co/Qwen/Qwen3.5-397B-A17B) for more details on the model. ## Use with mlx ```bash pip install -U mlx-vlm ``` On Macs with 128GiB of RAM - https://github.com/ml-explore/mlx-lm#large-models : ```bash sudo sysctl iogpu.wired_limit_mb=130000 ``` Also, mlx-vlm supports [TurboQuant](https://github.com/Blaizzy/mlx-vlm?tab=readme-ov-file#turboquant-kv-cache) ```bash # Server with TurboQuant mlx_vlm server \ --model m-i/Qwen3.5-397B-A17B-2.592bit\ --kv-bits 3.5 \ --kv-quant-scheme turboquant ``` ```bash python -m mlx_vlm.generate --model m-i/Qwen3.5-397B-A17B-2.592bit --max-tokens 100 --temperature 0.0 --prompt "Describe this image." --image ``` ## Quantization Predicate Main difference compared to m-i/Qwen3.5-397B-A17B-2.416bit is here the vision is untouched. Some parts have a slightly better quant param. ```python import mlx.core as mx mx.set_default_device(mx.cpu) from mlx_vlm import convert def qwen397b_predicate(path: str, module): """ Multi-bit quantization predicate for Qwen/Qwen3.5-397B-A17B. Maps GGUF _exps patterns (MoE experts) to HF switch_mlp paths. """ # ── SKIP: Critical for numerical stability ───────────────────────────── if any(kw in path for kw in [ "layernorm", "mlp.gate.", "shared_expert_gate.", "A_log", "dt_bias", "conv1d", ".bias", ".scales", "vision", ]): return False # ── 2-bit: MoE routed experts (switch_mlp) ← THIS HANDLES _exps ─────── # GGUF: ffn_{gate,up,down}_exps.weight → HF: switch_mlp.{gate,up,down}_proj.weight if "switch_mlp" in path: if any(proj in path for proj in ["gate_proj", "up_proj"]): return {"group_size": 128, "bits": 2, "mode": "affine"} # IQ2_XXS equivalent if "down_proj" in path: # Optional: use 3-bit for down_proj to mirror IQ2_S > IQ2_XXS return {"group_size": 32, "bits": 2, "mode": "affine"} # ── 4-bit: Token embeddings ──────────────────────────────────────────── if "embed_tokens" in path: return {"group_size": 32, "bits": 4, "mode": "affine"} # ── 5-bit: Shared expert gate/up ─────────────────────────────────────── if "shared_expert" in path and any(p in path for p in ["gate_proj", "up_proj"]): return {"group_size": 128, "bits": 5, "mode": "affine"} # ── 6-bit: Shared expert down ────────────────────────────────────────── if "shared_expert" in path and "down_proj" in path: return {"group_size": 32, "bits": 6, "mode": "affine"} # ── 5-bit: Linear/full attention projections ─────────────────────────── if "linear_attn" in path and any(p in path for p in ["in_proj", "out_proj"]): return {"group_size": 128, "bits": 5, "mode": "affine"} if "self_attn" in path and any(p in path for p in ["q_proj", "k_proj", "v_proj", "o_proj"]): return {"group_size": 128, "bits": 5, "mode": "affine"} # ── 6-8-bit: SSM dynamics, lm_head & fallback ───────────────────────────────────── if any(kw in path for kw in ["in_proj_a", "in_proj_b", "ssm_alpha", "ssm_beta"]): return {"group_size": 32, "bits": 6, "mode": "affine"} if "lm_head" in path: return {"group_size": 128, "bits": 8, "mode": "affine"} return {"group_size": 128, "bits": 6, "mode": "affine"} repo = "Qwen/Qwen3.5-397B-A17B" upload_repo = "m-i/Qwen3.5-397B-A17B-2.592bit" convert(repo, quantize=True, upload_repo=upload_repo, quant_predicate=qwen397b_predicate, ) ```