Instructions to use ukisai/Swift-1.5-5bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ukisai/Swift-1.5-5bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ukisai/Swift-1.5-5bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ukisai/Swift-1.5-5bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-1.5-5bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ukisai/Swift-1.5-5bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ukisai/Swift-1.5-5bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-1.5-5bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ukisai/Swift-1.5-5bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-1.5-5bit-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-1.5-5bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-1.5-5bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Bound retained server cache while keeping prompt reuse enabled
The pinned server constructed its retained-cache LRU without the configured byte limit, so entries saved during prefill or completion could exceed --prompt-cache-bytes. The patch enforces the limit on insertion and keeps prompt reuse enabled with bounded defaults: two retained entries, one prompt and decode stream, and 512-token prefill steps. Metal memory estimates set the default retained-cache budget; an explicit byte cap overrides it.
Both complete, existing Swift 4-bit and 5-bit checkpoints now passed independent native Metal validation on an AWS M4 Pro with 48 GiB RAM. Each processed our own synthetic 86k-token conversation, generated a reply, and completed two follow-ups with most history reused from cache. Hub offline mode was enabled. The full report and reproduction instructions include measured results and limitations.
Checks passed:
- Full checkpoint SHA-256 verification, runtime identity, three successful HTTP requests per format, retained-cache budget enforcement and substantial warm reuse.
- 44 unit/upstream server tests per clean architecture patch stack.
- 36 tiny-fixture offline HTTP requests per format, including streaming, sequential generation, disabled retention and explicit byte limits.
- Clean patch application and rollback, Black 25.1.0, isort 6.0.0 and Ruff 0.16.6.
These are independently executed checks, not hosted CI badges. The tested workload does not establish arbitrary context lengths, smaller Mac memory sizes, GUI/plugin integration, image/video chat or speculative MTP support. The retention cap excludes active-request and total process memory.
Weights, quantization, model semantics and tokenizer assets are unchanged. Existing installations must apply the runtime patch and restart. This PR is merged into public main at 8aff72b145212e62c15146e41dcbef35ea5fa9ff. The published tree and unchanged checkpoint hashes were verified after merging.
Retrieve only the runtime update
With the complete model and dependencies already local, fetch these three small files:
hf download ukisai/Swift-1.5-5bit-MLX \
compatibility/swift15-server-cache.patch \
compatibility/cache-tests/validation.json \
compatibility/cache-tests/verify_server_patch.py \
--revision 80896f29f71f3f6b8131f4322e9a70262e8e7b51 --local-dir Swift-1.5-5bit-MLX
Then follow the patch, restart and offline-startup instructions. The command does not download any weight shards.