--- license: apache-2.0 base_model: lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled base_model_relation: quantized library_name: mlx pipeline_tag: text-generation language: - en - zh tags: - mlx - apple-silicon - oq - oq4 - quantized - 4-bit - moe - mixture-of-experts - text-generation - qwen - qwen3.6 - reasoning - chain-of-thought - distillation - claude --- # Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-Text-MLX-oQ4 Mixed-precision quantization for Apple Silicon, **text-only mode** (vision tower stripped for faster, lighter inference). Quantized from [`lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled`](https://huggingface.co/lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled) using [oMLX](https://omlx.app)'s oQ4 algorithm (sensitivity-aware mixed-precision quantization). ## 📊 Specs | Field | Value | |---|---| | **Base model** | Qwen/Qwen3.6-35B-A3B (35B params, 128 experts MoE, A3B activation) | | **Fine-tune** | LoRA distilled from Claude 4.7 Opus reasoning outputs (lordx64 dataset) | | **Quantization** | oMLX oQ4 (mixed-precision, ~4.8 bpw average) | | **Modality** | Text only (vision tower stripped) | | **Format** | MLX safetensors | | **Model size** | ~19 GB | | **Inference memory** | ~20 GB (incl. KV cache and runtime overhead) | | **Recommended hardware** | Apple Silicon M2 Pro 32GB+ / M3 Max / M5 Max | ## 🚀 Quick Start ### Install ```bash pip install mlx-lm # Or with uv: uv tool install mlx-lm ``` ### Inference ```bash mlx_lm.generate \ --model wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-Text-MLX-oQ4 \ --prompt "Explain mixture of experts in one paragraph." \ --max-tokens 512 ``` ### Python API ```python from mlx_lm import load, generate model, tokenizer = load( "wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-Text-MLX-oQ4" ) response = generate( model, tokenizer, prompt="Solve: integrate x*sin(x) dx", max_tokens=512, ) print(response) ``` ### OpenAI-compatible Server ```bash mlx_lm.server \ --model wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-Text-MLX-oQ4 \ --port 8080 ``` Drop-in compatible with OpenAI clients including Claude Code (with custom backend), AstrBot, Open WebUI, LibreChat, and Continue.dev. ## 📈 Measured Performance Benchmarked on MacBook Pro M5 Max 128GB: | Metric | Value | |---|---| | Prompt processing | ~67 tokens/s | | Generation speed | ~127 tokens/s | | Peak memory | 20.4 GB | | Model load time | ~8 sec | ### Why Text-Only? Stripping the vision tower offers practical advantages for text-only workflows: - **~10% faster generation** vs the VLM equivalent (no vision compute path overhead) - **~2 GB less peak memory** - **Simpler deployment** — no need for vision processor configs or image preprocessing dependencies If you only feed text into your model, this version is strictly better than the VLM variant. ## 🧠 Model Behavior Inherits the Claude reasoning distillation: the model uses `...` tags to structure its chain-of-thought before producing the final response. **Best for:** - Coding agents and tool-use workflows - Complex reasoning tasks (math, logic, analysis) - Long-form text generation - Scenarios where visible reasoning improves quality Sample output structure: ``` 1. Analyze the user's request: ... 2. Identify the key constraints: ... 3. Formulate the solution: ... Here is my analysis: ... ``` ## 🔬 Quantization Details - **Source model**: `Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled` (BF16 MLX-converted) - **Sensitivity model**: 8-bit quantization of the same distilled model (self-referenced sens for tight distribution alignment) - **Non-quant weight dtype**: bfloat16 (M3+ optimal) - **Text-Only mode**: ON (vision tower stripped — verified: 0 vision-related tensors) - **Quantizer**: oMLX ## 📦 Other Versions in This Series | Version | Size | Best for | |---|---|---| | [VLM-MLX-oQ4](https://huggingface.co/wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ4) | 19.6 GB | Memory-constrained inference (with vision) | | [VLM-MLX-oQ6](https://huggingface.co/wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ6) | 27 GB | Recommended VLM quality/size ratio | | [VLM-MLX-oQ8](https://huggingface.co/wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-VLM-MLX-oQ8) | 35 GB | VLM quality reference baseline | | [Text-MLX-oQ4](https://huggingface.co/wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-Text-MLX-oQ4) | 19 GB | **Text-only, fastest** | | [Text-MLX-oQ6](https://huggingface.co/wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-Text-MLX-oQ6) | 27 GB | **Recommended text-only** | | [Text-MLX-oQ8](https://huggingface.co/wangkezun/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-Text-MLX-oQ8) | 34 GB | Text-only, max quality | **Choosing a version:** - Text-only workflows (coding, agents, dialogue) → `Text` variants are faster and lighter - Image input needed (OCR, visual analysis, screenshot understanding) → `VLM` variants - oQ6 is the sweet spot for most use cases. oQ8 yields diminishing returns relative to its size. ## ⚠️ Disclaimer This model derives from a chain of upstream work: 1. Base model [`Qwen/Qwen3.6-35B-A3B`](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) by Alibaba's Qwen team (Apache-2.0) 2. Distilled variant [`lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled`](https://huggingface.co/lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled) by lordx64 using Claude 4.7 Opus reasoning outputs (Apache-2.0) 3. This quantization by [@wangkezun](https://huggingface.co/wangkezun) using oMLX on Apple Silicon **This model is not affiliated with or endorsed by Anthropic, PBC.** "Claude" is a trademark of Anthropic, PBC. The use of "Claude" in this model name is purely descriptive (nominative fair use) to indicate the upstream training data lineage. By using this model, you agree to comply with: - The Apache-2.0 license inherited from the base model - Any applicable license terms of the upstream distillation dataset - Local laws and regulations governing AI model usage in your jurisdiction ## 🙏 Acknowledgments - **Alibaba Qwen Team** — for the Qwen3.6-35B-A3B base model - **lordx64** — for the reasoning-focused LoRA distillation - **Jundot (oMLX team)** — for the oQ mixed-precision quantization algorithm - **Apple MLX team** — for the MLX framework and tooling ## 📜 License Apache-2.0 (inherited from base model). --- **Generated**: 2026-04-26 **Quantizer**: [@wangkezun](https://huggingface.co/wangkezun)