Instructions to use ukisai/Swift-1.5-4bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ukisai/Swift-1.5-4bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ukisai/Swift-1.5-4bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ukisai/Swift-1.5-4bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-1.5-4bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ukisai/Swift-1.5-4bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ukisai/Swift-1.5-4bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-1.5-4bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ukisai/Swift-1.5-4bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-1.5-4bit-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-1.5-4bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-1.5-4bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download SERVER_CACHE_UPDATE.md from ukisai/Swift-1.5-4bit-MLX: direct link, hf CLI and curl.
- Browser
- Download file 5.34 kB
-
https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/resolve/main/SERVER_CACHE_UPDATE.md
- Command line
-
hf download hf://ukisai/Swift-1.5-4bit-MLX/SERVER_CACHE_UPDATE.md
-
curl -L -o SERVER_CACHE_UPDATE.md https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/resolve/main/SERVER_CACHE_UPDATE.md
Server cache update
Starting from Homebrew or seeing Received 501 parameters not in model?
Use QUICKSTART.md to install the required runtime and create an
explicit serve launcher. It reuses your existing model directory, including an
HF cache snapshot. A bare mlx_lm.server command may select a separate Homebrew
installation that lacks the Swift architecture and cache patches.
This update changes the pinned MLX-LM server, not the checkpoint. It is a Swift
runtime patch based on official MLX-LM c69d1288440a0dc4e6401fc417098b07598dccd5, not an upstream release.
The existing architecture patch remains required.
Apply to an existing installation
Stop the server. Activate the same Python environment used for serving. From the
directory containing swift15-mlx-lm and Swift-1.5-4bit-MLX, after downloading this patch:
git -C swift15-mlx-lm rev-parse HEAD
git -C swift15-mlx-lm apply --check ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch
git -C swift15-mlx-lm apply ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch
python -m pip install --no-deps -e ./swift15-mlx-lm
python Swift-1.5-4bit-MLX/compatibility/cache-tests/verify_server_patch.py
The revision must equal the pinned revision above. Stop if the patch check fails; do not force it onto different server code or apply it twice. For a fresh install, follow USAGE.md, which applies the architecture patches first. The server verifier checks the active Python installation without loading weights.
Once dependencies and the complete checkpoint are already local, startup works without Hub access:
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python -m mlx_lm.server \
--model ./Swift-1.5-4bit-MLX --host 127.0.0.1 --port 8080
Remove an earlier --prompt-cache-size 0 override to enable reuse. Explicit
arguments override the new defaults. Installing/downloading prerequisites still
requires internet; offline operation requires those prerequisites already present.
Behavior and limits
Defaults: two retained cache entries, prompt concurrency 1, decode concurrency 1, and prefill steps of 512 tokens. The byte cap is now passed to the LRU cache, so every insertion enforces it. Previously the CLI byte limit only triggered a trim when adding a batched request; later retained entries could exceed that setting.
The automatic Metal budget uses the reported recommended working-set size R,
the loaded model parameter bytes W, and workspace reserve max(4 GiB, R/8).
It budgets half the remaining space for retained entries:
max(0, (R - W - reserve) / 2). This is a conservative estimate, not a guarantee
that an arbitrary active context fits. It does not change macOS or MLX limits.
When device metadata is unavailable, the entry limit still applies; supply an
explicit byte limit if needed. --prompt-cache-bytes 2GB, for example, sets a
retention budget of 2,000,000,000 bytes; it is not a universal recommended value.
The active request, temporary tensors, allocator and other apps use additional memory. Cache byte accounting is logical and may include shared buffers. Entries that cannot fit are evicted whole; the client's input history is not truncated. Reuse can therefore vary with context size. This patch does not add KV quantization, change attention or alter any checkpoint tensors.
Reproduce the checks
Use the patched environment with Python 3.12, MLX 0.32.2 and MLX-LM 0.32.0.
validation.json and the two offline-*-results.json files record our local
results. No hosted CI status is claimed. The tests use random tiny models and the
local tokenizer; they never open the released Swift weight shards.
python -m pip install 'pytest==9.1.1' 'requests==2.32.5'
python -m pytest -q swift15-mlx-lm/tests/test_server_cache_budget.py swift15-mlx-lm/tests/test_server.py
python Swift-1.5-4bit-MLX/compatibility/cache-tests/test_server_cache_http.py \
--assets ./Swift-1.5-4bit-MLX \
--config Swift-1.5-4bit-MLX/compatibility/cache-tests/synthetic-config.json \
--bits 4 --output cache-http-results.json
The upstream server suite may fetch its small upstream test model if not already cached. The separate HTTP harness sets offline mode and needs only local assets. It covers 12 requests each with automatic budgeting, an explicit byte limit, and disabled retention; streaming and seeded sequential execution are included. Synthetic output is deliberately fixed for transport/cache checks, not quality. Full-model follow-up validation is now recorded in FULL_MAC_VALIDATION.md: both complete checkpoints passed an 86k-token synthetic text conversation and two cached follow-ups on a 48 GiB M4 Pro using the patched server. GUI integration and other memory/context sizes remain outside that test.
Roll back this runtime patch
Stop the server, then reverse only this patch and restart with explicit settings:
git -C swift15-mlx-lm apply --reverse --check ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch
git -C swift15-mlx-lm apply --reverse ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch
This leaves the architecture patch and model files intact. An editable install uses the source change after restart; reinstall if a non-editable copy was used.