Image-Text-to-Text
MLX
Safetensors
qwen3_5
omlx
quantization
mixed-precision
apple-silicon
mtp
speculative-decoding
qwen
vision
base_model_size:10B to 100B
conversational
4-bit precision
Instructions to use TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp") config = load_config("TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Update references after account rename
Browse files
README.md
CHANGED
|
@@ -35,7 +35,7 @@ I am not affiliated with UkisAI. All upstream weights, benchmarks and license te
|
|
| 35 |
|
| 36 |
## Pick a variant
|
| 37 |
|
| 38 |
-
| | [oQ4-mtp](https://huggingface.co/
|
| 39 |
|---|---|---|---|
|
| 40 |
| Weights on disk | 15.81 GiB (16.97 GB), 4 shards | 22.09 GiB (23.72 GB), 5 shards | 27.94 GiB (30.00 GB), 6 shards |
|
| 41 |
| Weight precision | mixed 4/5-bit | mixed 6/8-bit | uniform 8-bit |
|
|
@@ -78,7 +78,7 @@ Precision is mixed per module and recorded verbatim in `config.json` → `quanti
|
|
| 78 |
|
| 79 |
```bash
|
| 80 |
# 1. drop the folder into the oMLX model dir
|
| 81 |
-
git clone https://huggingface.co/
|
| 82 |
|
| 83 |
# 2. start the multi-model server (model id = folder name)
|
| 84 |
omlx serve --model-dir ~/.omlx/models --port 8000
|
|
@@ -105,7 +105,7 @@ In oMLX model settings, enable the speculative head and the matching reasoning p
|
|
| 105 |
|
| 106 |
```bash
|
| 107 |
pip install -U mlx-lm mlx-vlm
|
| 108 |
-
python -m mlx_lm.server --model
|
| 109 |
```
|
| 110 |
|
| 111 |
Only oMLX 0.6.4 is verified by me; if you get `mlx_lm` running this architecture, please open an issue and I will document it.
|
|
@@ -162,9 +162,9 @@ Measured throughput/acceptance-length data and issue reports (especially "quant
|
|
| 162 |
```bibtex
|
| 163 |
@misc{swift-qwen3.8-27b-mlx-quants,
|
| 164 |
title = {Swift-Qwen3.8-27b-oQ4-mtp}: oMLX/MLX quantization of Swift-Qwen3.8-27B with MTP head retained,
|
| 165 |
-
author = {
|
| 166 |
year = {2026},
|
| 167 |
-
howpublished = {\url{https://huggingface.co/
|
| 168 |
note = {Unofficial quantization of ukisai/Swift-Qwen3.8-27b}
|
| 169 |
}
|
| 170 |
|
|
|
|
| 35 |
|
| 36 |
## Pick a variant
|
| 37 |
|
| 38 |
+
| | [oQ4-mtp](https://huggingface.co/TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp) | [oQ6-mtp](https://huggingface.co/TokenAI-zer/Swift-Qwen3.8-27b-oQ6-mtp) | [oQ8-mtp](https://huggingface.co/TokenAI-zer/Swift-Qwen3.8-27b-oQ8-mtp) |
|
| 39 |
|---|---|---|---|
|
| 40 |
| Weights on disk | 15.81 GiB (16.97 GB), 4 shards | 22.09 GiB (23.72 GB), 5 shards | 27.94 GiB (30.00 GB), 6 shards |
|
| 41 |
| Weight precision | mixed 4/5-bit | mixed 6/8-bit | uniform 8-bit |
|
|
|
|
| 78 |
|
| 79 |
```bash
|
| 80 |
# 1. drop the folder into the oMLX model dir
|
| 81 |
+
git clone https://huggingface.co/TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp ~/.omlx/models/Swift-Qwen3.8-27b-oQ4-mtp
|
| 82 |
|
| 83 |
# 2. start the multi-model server (model id = folder name)
|
| 84 |
omlx serve --model-dir ~/.omlx/models --port 8000
|
|
|
|
| 105 |
|
| 106 |
```bash
|
| 107 |
pip install -U mlx-lm mlx-vlm
|
| 108 |
+
python -m mlx_lm.server --model TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp --port 8000
|
| 109 |
```
|
| 110 |
|
| 111 |
Only oMLX 0.6.4 is verified by me; if you get `mlx_lm` running this architecture, please open an issue and I will document it.
|
|
|
|
| 162 |
```bibtex
|
| 163 |
@misc{swift-qwen3.8-27b-mlx-quants,
|
| 164 |
title = {Swift-Qwen3.8-27b-oQ4-mtp}: oMLX/MLX quantization of Swift-Qwen3.8-27B with MTP head retained,
|
| 165 |
+
author = {TokenAI-zer},
|
| 166 |
year = {2026},
|
| 167 |
+
howpublished = {\url{https://huggingface.co/TokenAI-zer/Swift-Qwen3.8-27b-oQ4-mtp}},
|
| 168 |
note = {Unofficial quantization of ukisai/Swift-Qwen3.8-27b}
|
| 169 |
}
|
| 170 |
|