--- license: mit license_link: https://huggingface.co/zai-org/GLM-5.3-Flash language: - en - zh tags: - glm - glm-5.3 - glm-5.3-flash - abliterated - uncensored - moe - mlx - mlx-vlm - omlx - oq - apple-silicon - dual-ane - speculative-decoding - anchored-tensors - multimodal base_model: Vontra/GLM-5.3-Flash-MLX-oQ4-MTP base_model_relation: finetune library_name: mlx pipeline_tag: image-text-to-text extra_gated_heading: Acknowledge the Responsible Use Agreement to access this repository extra_gated_description: Access is granted automatically after you agree to the terms below and submit the form. extra_gated_button_content: Agree and request access extra_gated_prompt: | ## Responsible Use Agreement This model has had safety refusals removed. That makes it useful for red-teaming, security research, evaluation, and unfiltered assistant tasks — and also removes guardrails a user must therefore supply themselves. **Prohibited uses (you must agree before access is granted):** - Anything involving the sexual exploitation or endangerment of minors. - You must be of age 18 years or older to use and download this model. - You agree any information generated that can cause harm in terms of generating recipe, knowledge to make any materials/substances is your own input and responsibility. You will be accountable for any harm/damage caused by your action/input. - Content promoting self-harm or suicide. - Generation of material that is illegal in your jurisdiction, or that targets real individuals for harassment, doxxing, or fraud. - Any use prohibited by the upstream Z.AI / GLM MIT license. You are responsible for adding appropriate safety filtering, human review, and access controls for your deployment. The weights are provided as-is, with no warranty. The license is inherited from the upstream Z.AI GLM-5.3-Flash MIT license — review and comply with it before use or redistribution. extra_gated_fields: Username: text Email: text Reason for intended use: text I am 18 years of age or older: checkbox I will not use this model for any sexual exploitation or endangerment of minors: checkbox I accept full responsibility for my inputs and any harm from generated content: checkbox I will not use this model for self-harm, suicide promotion, illegal activity, harassment, doxxing, or fraud: checkbox I agree to comply with the upstream ZAI GLM MIT license: checkbox I agree to the Responsible Use terms above: checkbox --- # keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4 Abliterated **Vontra GLM-5.3-Flash MLX oQ4** for Mac Studio oMLX **0.6.3rc2** Dual-ANE. Same Dealign `o_proj` L15–45 transplant as the Spark NVFP4 pack — **L0–14 stay stock**, MTP included. See **[RESPONSIBLE_USE.md](./RESPONSIBLE_USE.md)** and the gate form above. Access is **gated with automatic approval** after you agree. | | | |--|--| | **HF (these weights)** | https://huggingface.co/drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4 | | **GitHub (oMLX Dual-ANE recipe)** | https://github.com/drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4 | | **Spark NVFP4 cousin (same ablit)** | [HF](https://huggingface.co/drowzeys/keys-GLM-5.3-Flash-NVFP4-ablit-l15-45-anchorstock) · [GitHub](https://github.com/drowzeys/keys-GLM-5.3-Flash-NVFP4-ablit-l15-45-anchorstock) | | **Stock oQ4 parent** | [`Vontra/GLM-5.3-Flash-MLX-oQ4-MTP`](https://huggingface.co/Vontra/GLM-5.3-Flash-MLX-oQ4-MTP) | | **Ablit source (`o_proj` L15–45)** | [`dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4`](https://huggingface.co/dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4) | | **Upstream** | [`zai-org/GLM-5.3-Flash`](https://huggingface.co/zai-org/GLM-5.3-Flash) | | **Ablit** | L15–45 `self_attn.o_proj` as **BF16** (31 tensors, includes MTP) · L0–14 affine oQ4 stock | | **Gate** | **32/32 bypass, 0 refuse, 0 garble** | | **Decode (M3 Ultra 256 GB, oMLX 0.6.3rc2)** | **~24 tok/s** (192-token gens 22.7–25.0; 128-token 23.4 / 23.9). Stock oQ4 on the same box: **26.4 tok/s**. | These weights have safety refusals removed. Research / red-team only — you supply the guardrails. ## What changed vs Vontra oQ4 Blackfrost-style rank-1 projection + affine requant stayed at **19/32** on this quant. The Spark 32/32 dest is a **byte-copy of Dealign `o_proj`**. On MLX we do **not** requantize those tensors: they land as BF16 `nn.Linear` weights (`model-ablit-oproj-l15-45.safetensors`). Experts, vision, QKV, and L0–14 `o_proj` stay Vontra affine oQ4. ## One-shot (Mac Studio) ```bash git clone https://github.com/drowzeys/keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4.git cd keys-Mac-oMLX-0.6.3.2RC-Dual-ANE-GLM-5.3-Flash-Abliterated-oQ4 # Hugging Face gate: agree on the model card, then `hf auth login` bash oneshot-setup.sh omlx serve --model-dir ~/.omlx/models --host 0.0.0.0 --port 11500 ``` Serve id: `glm53-flash-oq4-mtp-ablit-l15-45`. Dual-ANE tile 4096, Lightning MTP on, thinking off. ## Credits Cite the original authors first. | Who | What we used | |---|---| | **Z.ai** | Upstream GLM-5.3-Flash | | **Vontra** | Mac oQ4 body; L0–14 `o_proj`, experts, vision stay theirs | | **dealignai** (compute: Jordan Schenck) | BF16 `o_proj` L15–45 + MTP that reaches 32/32 | | **jundot** | oMLX, oQ, Lightning MTP, Dual-ANE | | **onthehub97** | Dual-ANE/GPU prompt processing | | **Blaizzy** | day-0 glm5_next in mlx-vlm | | **PipeNetwork** | ClampedSwiGLU / Mac glm5_next runtime | | **ml-explore** | MLX | | **Blackfrost** | Direction reference; not this checkpoint | | **LibertAI** | Spark NVFP4 layout Dealign `o_proj` was extracted from | Pack: drowzeys / keys. Donate: [GoFundMe](https://t.co/5O4WUxexXa).