Image-Text-to-Text
MLX
Safetensors
qwen3_5
mlx-vlm
omlx
qwen
qwen3.8
multimodal
quantized
apple-silicon
4-bit precision
conversational
Instructions to use Vontra/Qwen3.8-27B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Vontra/Qwen3.8-27B-MLX-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("Vontra/Qwen3.8-27B-MLX-4bit") config = load_config("Vontra/Qwen3.8-27B-MLX-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Vontra/Qwen3.8-27B-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Vontra/Qwen3.8-27B-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Vontra/Qwen3.8-27B-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use Vontra/Qwen3.8-27B-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Vontra/Qwen3.8-27B-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Vontra/Qwen3.8-27B-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Vontra/Qwen3.8-27B-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Vontra/Qwen3.8-27B-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Vontra/Qwen3.8-27B-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| library_name: mlx | |
| license: apache-2.0 | |
| base_model: Qwen/Qwen3.8-27B | |
| base_model_relation: quantized | |
| pipeline_tag: image-text-to-text | |
| tags: | |
| - mlx | |
| - mlx-vlm | |
| - omlx | |
| - qwen | |
| - qwen3.8 | |
| - multimodal | |
| - quantized | |
| - apple-silicon | |
| - 4-bit | |
| <p align="center"> | |
| <a href="https://qwenlm.github.io/"><img src="qwen-logo.png" width="96" height="95" alt="Qwen"></a><br> | |
| <img src="https://img.shields.io/badge/Apple_Silicon-MLX-000000?style=for-the-badge&logo=apple&logoColor=white" alt="Apple silicon MLX"> | |
| <img src="https://img.shields.io/badge/Vontra-oMLX-6E56CF?style=for-the-badge&logo=huggingface&logoColor=white" alt="Vontra oMLX"> | |
| </p> | |
| <h1 align="center">Qwen3.8-27B — MLX 4-bit</h1> | |
| <p align="center"> | |
| A native Apple-silicon conversion of <a href="https://huggingface.co/Qwen/Qwen3.8-27B">Qwen/Qwen3.8-27B</a>, quantized with stock 4-bit affine weights for MLX-VLM and oMLX. | |
| </p> | |
| <p align="center"> | |
| <a href="https://huggingface.co/Qwen/Qwen3.8-27B">Original model</a> · | |
| <a href="https://qwenlm.github.io/">Qwen</a> · | |
| <a href="https://github.com/Blaizzy/mlx-vlm">MLX-VLM</a> · | |
| <a href="https://www.apache.org/licenses/LICENSE-2.0">Apache 2.0</a> | |
| </p> | |
| ## About this conversion | |
| This repository contains a stock 4-bit affine MLX conversion of Qwen3.8-27B. The upstream model is a dense, native vision-language model with flexible thinking control and support for text, images, and video. Its tokenizer, processor configuration, chat template, and generation configuration are preserved. | |
| | Item | Value | | |
| | --- | --- | | |
| | Base model | [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B) | | |
| | Format | MLX safetensors | | |
| | Quantization | 4-bit affine, group size 64 | | |
| | Effective precision | 4.695 bits per weight | | |
| | Conversion stack | `mlx-vlm 0.6.3`, `mlx-lm 0.31.3`, `mlx 0.32.0` | | |
| | Weight shards | 3 | | |
| | Weight size | 16.06 GB (14.95 GiB) | | |
| | Maximum configured context | 262,144 tokens | | |
| | Architecture | `qwen3_5` / `Qwen3_5ForConditionalGeneration` | | |
| ## Apple-silicon performance | |
| This checkpoint was load-tested and generation-tested on the following machine: | |
| | Hardware | Configuration | | |
| | --- | --- | | |
| | Host | Mac Studio | | |
| | Chip | Apple M3 Ultra | | |
| | CPU | 32 cores (24 performance + 8 efficiency) | | |
| | Unified memory | 256 GB | | |
| | Runtime | MLX-VLM 0.6.3 / MLX 0.32.0 | | |
| | Measurement | Result | | |
| | --- | ---: | | |
| | Decode (median) | **39.90 tokens/s** | | |
| | Reported peak memory | **20.14 GB** | | |
| | Timed runs | 3 × 256 generated tokens | | |
| | Warm-up | 256 generated tokens | | |
| | Prompt | 81 tokens after chat templating | | |
| The decode figure is the median of three greedy 256-token runs after a 256-token Metal-kernel warm-up. Individual runs measured 39.90, 39.89, and 39.91 tokens/s. This is a practical local reference, not a controlled cross-platform benchmark; prompt length, context growth, sampler settings, memory pressure, thermal state, and runtime versions can materially change performance. | |
| ## Quick start with MLX-VLM | |
| ```bash | |
| python -m pip install -U mlx-vlm huggingface_hub | |
| python -m mlx_vlm.generate \ | |
| --model Vontra/Qwen3.8-27B-MLX-4bit \ | |
| --prompt "Explain the difference between linear and full attention." \ | |
| --max-tokens 512 | |
| ``` | |
| Download for local use: | |
| ```bash | |
| hf download Vontra/Qwen3.8-27B-MLX-4bit \ | |
| --local-dir ~/.omlx/models/Vontra/Qwen3.8-27B-MLX-4bit | |
| ``` | |
| ## Using it with oMLX | |
| 1. Place the model at `~/.omlx/models/Vontra/Qwen3.8-27B-MLX-4bit`. | |
| 2. Refresh the oMLX model registry. | |
| 3. Load `Qwen3.8-27B-MLX-4bit` and use the chat UI or OpenAI-compatible endpoint. | |
| ```bash | |
| curl "$OMLX_BASE_URL/v1/chat/completions" \ | |
| -H "Content-Type: application/json" \ | |
| -H "Authorization: Bearer $OMLX_API_KEY" \ | |
| -d '{ | |
| "model": "Qwen3.8-27B-MLX-4bit", | |
| "messages": [{"role": "user", "content": "Write a short Swift actor example."}], | |
| "temperature": 1.0, | |
| "top_p": 0.95, | |
| "max_tokens": 256 | |
| }' | |
| ``` | |
| For long prompts, begin with a conservative context limit and increase it while watching memory pressure. The configured context is a model capability, not a guarantee that every host can prefill it within available unified memory. | |
| ## Architecture | |
| Qwen3.8-27B is a dense causal language model with a vision encoder. It uses the Qwen3.5 architectural foundation, interleaving Gated DeltaNet linear-attention blocks with periodic full-attention blocks. | |
| | Architecture detail | Upstream value | | |
| | --- | ---: | | |
| | Parameters | 27B | | |
| | Language layers | 64 | | |
| | Hidden size | 5,120 | | |
| | Attention heads / KV heads | 24 / 4 | | |
| | Linear-attention V / QK heads | 48 / 16 | | |
| | FFN intermediate size | 17,408 | | |
| | Vocabulary / padded embeddings | 248,320 | | |
| | Configured context | 262,144 tokens | | |
| For upstream evaluations, usage guidance, intended use, limitations, safety information, and the full architecture discussion, see the [original model card](https://huggingface.co/Qwen/Qwen3.8-27B). | |
| ## Conversion and validation notes | |
| - Source weights: the official Qwen checkpoint. | |
| - Quantization: stock 4-bit affine weights with group size 64. | |
| - The upstream tokenizer, processor files, chat template, and generation configuration are preserved. | |
| - All 2,180 converted tensors and all three indexed shards were checked locally. | |
| - Quantization can reduce output quality relative to the source weights; use a higher-precision variant when quality matters more than memory use. | |
| - The model was loaded and exercised through end-to-end generation on Apple silicon. | |
| This is a community conversion, not an official Qwen release. Validate quality and numerical behaviour on your own representative workload before production use. | |
| ## Licence and attribution | |
| The upstream model is released under the **Apache License 2.0**. A copy is included in this repository; review it before use or redistribution. | |
| All model design, training, benchmark, and upstream documentation credit belongs to Qwen and the original contributors. The MLX conversion, Apple-silicon validation, compatibility work, and packaging are provided by [Vontra](https://huggingface.co/Vontra). | |
| <!-- vontra-chooser-start --> | |
| ## Choose for your Mac | |
| [64GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-64gb-macs-6a9fefda17932216ec9ab457) · [128GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-128gb-macs-6a9ff0abd31bc9abbe7922d7) · [256GB Macs](https://huggingface.co/collections/Vontra/mlx-models-for-256gb-macs-6a9ff0ef9fed7c5bdca15e9b) | |
| Published peak memory: **20.14 GB**; estimated starting tier: **64GB**, leaving about **43 GB** nominal headroom. The collections use published M3 Studio peaks with at least 25% nominal headroom; fit on other Macs is an estimate, and full context is not guaranteed. Start with short context and one request. | |
| ### Runtime and evidence | |
| The exact tested oMLX application version is not recorded here; a library version is not an app version. The original performance tables retain their benchmark conditions and speed figures; this documentation update adds no new test results. | |
| ### Quick start and demo prompt | |
| ```bash | |
| hf download Vontra/Qwen3.8-27B-MLX-4bit --local-dir ./models/Qwen3.8-27B-MLX-4bit | |
| ``` | |
| Add the downloaded folder to oMLX model directories, refresh the list, and follow this card's architecture and MTP compatibility requirements before loading. | |
| Try this in a new chat with a 128-token output limit: | |
| ```text | |
| Explain why the sky looks blue in three short sentences. | |
| ``` | |
| This is a demo prompt to try, not a recorded successful run; a captured demonstration for this documentation update is not yet available. | |
| [Follow Vontra for new Apple Silicon releases and fixes.](https://huggingface.co/Vontra) | |
| <!-- vontra-chooser-end --> | |