Instructions to use ukisai/Swift-1.5-4bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ukisai/Swift-1.5-4bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ukisai/Swift-1.5-4bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ukisai/Swift-1.5-4bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-1.5-4bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ukisai/Swift-1.5-4bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ukisai/Swift-1.5-4bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-1.5-4bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ukisai/Swift-1.5-4bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-1.5-4bit-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-1.5-4bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-1.5-4bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download USAGE.md from ukisai/Swift-1.5-4bit-MLX: direct link, hf CLI and curl.
- Browser
- Download file 4.07 kB
-
https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/resolve/9fd3d5f738d71f50e0fe84574c44948c0457a01b/USAGE.md
- Command line
-
hf download hf://ukisai/Swift-1.5-4bit-MLX@9fd3d5f738d71f50e0fe84574c44948c0457a01b/USAGE.md
-
curl -L -o USAGE.md https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/resolve/9fd3d5f738d71f50e0fe84574c44948c0457a01b/USAGE.md
Load Swift 1.5 with its complete MLX architecture
Use the included patch with the pinned official Apple MLX-LM revision. Unpatched text-only Qwen support does not preserve this checkpoint's complete parameter tree.
Use Python 3.12 in a new working directory. Install the HF CLI before using it,
and pin the complete model revision as well as the MLX-LM source revision.
This repository is private: after installing the CLI, run hf auth login
interactively if not already signed in with an account that has access.
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install 'huggingface_hub==1.31.0'
SWIFT_MLX_REVISION=d2140379e1fd593002c92fa552b2b37fb6eb1159
hf download ukisai/Swift-1.5-4bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-4bit-MLX
hf cache verify ukisai/Swift-1.5-4bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-4bit-MLX --fail-on-missing-files
git clone https://github.com/ml-explore/mlx-lm.git swift15-mlx-lm
git -C swift15-mlx-lm checkout --detach c69d1288440a0dc4e6401fc417098b07598dccd5
git -C swift15-mlx-lm apply --check ../Swift-1.5-4bit-MLX/compatibility/swift15-mlx-lm.patch
git -C swift15-mlx-lm apply ../Swift-1.5-4bit-MLX/compatibility/swift15-mlx-lm.patch
Stop after any missing-file or checksum failure. The checkpoint contains about 15.83 GB of tensor data, before runtime, cache and OS overhead. Do not force the full model onto a 16 GiB Mac or increase system memory limits. Full-model Apple generation remains unverified. The recorded small Metal samples are not a full run.
On Apple Silicon:
pip install 'mlx==0.32.2' 'transformers==5.14.1' 'huggingface_hub==1.31.0' pillow
pip install -e ./swift15-mlx-lm
On Linux CPU (Python 3.12 and glibc 2.35 or newer):
pip install 'mlx[cpu]==0.32.2' 'transformers==5.14.1' 'huggingface_hub==1.31.0' pillow
pip install -e ./swift15-mlx-lm
Text generation:
import mlx.core as mx
from mlx_lm import load, generate
model, tokenizer = load("Swift-1.5-4bit-MLX")
if mx.default_device() == mx.cpu:
model.apply(
lambda value: value.astype(mx.float32)
if mx.issubdtype(value.dtype, mx.floating) else value
)
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Say hello."}],
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=32))
The Linux CPU branch promotes only in-memory floating parameters to FP32. Packed
4-bit weights and all files remain unchanged. This avoids the official MLX 0.32.2
Linux scalar BF16 quantized-matmul accumulation bug reproduced in
compatibility/cpu-quantized-matmul-diagnostic.json (8,192 exact ones summed to
256 in BF16, versus the correct 8,192 in FP32). The release's CPU generation,
MTP and vision smoke tests use this FP32 runtime. Apple Silicon inference does
not use this CPU workaround; full-model Apple Silicon execution was not tested.
The original chat template also accepts reasoning_effort="low", "medium",
and "xhigh"; this release validates the original low and xhigh formats.
This structural smoke test does not establish long-context or benchmark accuracy.
The patch implements an explicit MTP step (model.mtp_logits) and the vision
encoder (model.visual). Their weights are retained and the release validation
records their component execution. Speculative generation and integrated image/video
chat are not implemented. Unsupported multimodal generation calls raise an error.
Reproduce conversion only from the complete original Swift BF16 export identified
in QUANTIZATION_MANIFEST.json, after verifying its 18 shards and original assets:
mlx_lm.convert --hf-path /path/to/Swift-1.5-BF16 \
--mlx-path Swift-1.5-4bit-MLX \
--quantize --q-mode affine --q-bits 4 --q-group-size 64
The converter refuses an existing output directory. It uses official MLX-LM lazy loading, quantization, sharding, and saving; the patch supplies the complete model and strict parameter/asset mapping.