Instructions to use ukisai/Swift-1.5-4bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ukisai/Swift-1.5-4bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ukisai/Swift-1.5-4bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ukisai/Swift-1.5-4bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-1.5-4bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ukisai/Swift-1.5-4bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ukisai/Swift-1.5-4bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-1.5-4bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ukisai/Swift-1.5-4bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-1.5-4bit-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-1.5-4bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-1.5-4bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download FULL_MAC_VALIDATION.md from ukisai/Swift-1.5-4bit-MLX: direct link, hf CLI and curl.
- Browser
- Download file 3.56 kB
-
https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/resolve/main/FULL_MAC_VALIDATION.md
- Command line
-
hf download hf://ukisai/Swift-1.5-4bit-MLX/FULL_MAC_VALIDATION.md
-
curl -L -o FULL_MAC_VALIDATION.md https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/resolve/main/FULL_MAC_VALIDATION.md
3.56 kB
| # Full Swift Mac cache validation — 2026-09-25 | |
| Both complete, existing Swift checkpoints passed native Metal text generation | |
| and two cached follow-up requests on an AWS M4 Pro Mac with 48 GiB RAM. | |
| All weight shards were verified against their recorded SHA-256 hashes before | |
| loading. No quantization, tensor edits, context changes or manual memory-limit overrides | |
| were performed. Hub offline mode was enabled during testing. | |
| The input is our own synthetic repeated-record conversation, followed by a short | |
| request to reply READY. It is a memory/cache regression test, not a quality | |
| benchmark, an exact reproduction of another person's conversation, or a claim | |
| that every context fits. The follow-ups must reuse at least half the prompt to | |
| pass; the test fails if the worker dies or the retained-cache budget is exceeded. | |
| | Checkpoint | Initial prompt tokens | First request, seconds | Follow-ups, seconds | Follow-up cached tokens | Peak MLX allocation, GiB | | |
| |---|---:|---:|---|---|---:| | |
| | 4-bit | 86,004 | 893.64 | 1.46, 1.42 | 86000, 86020 | 31.57 | | |
| | 5-bit | 86,004 | 907.00 | 1.45, 1.45 | 86000, 86020 | 34.79 | | |
| The first request includes model loading and initial prompt processing. MLX peak | |
| allocation is not total process or system memory. Complete request results, | |
| generated replies, cache limits and sampled swap observations are recorded in | |
| `compatibility/cache-tests/full-mac-4bit-results.json` and | |
| `compatibility/cache-tests/full-mac-5bit-results.json`. | |
| Environment: macOS 26.7 (25G229), Python 3.12.13, MLX 0.32.2, MLX-LM 0.32.0, Transformers 5.14.1, | |
| Hugging Face Hub 1.31.0. Official MLX-LM base revision: | |
| `c69d1288440a0dc4e6401fc417098b07598dccd5`. | |
| The shared runtime uses the existing Swift architecture patch, its 5-bit support | |
| extension, and the unchanged server-cache patch. The 5-bit extension also accepts | |
| the original 4-bit format. The cache defaults are two retained entries, automatic | |
| byte budgeting, one prompt/decode stream, and 512-token prefill steps. | |
| To reproduce, install the pinned environment and all three patches in the order | |
| documented by the 5-bit release, then use a complete local snapshot. Run one model | |
| at a time on an Apple Silicon Mac with at least 48 GiB RAM: | |
| ```bash | |
| HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python \ | |
| Swift-1.5-4bit-MLX/compatibility/cache-tests/validate_full_mac_cache.py \ | |
| --snapshot Swift-1.5-4bit-MLX --output full-mac-4bit-results | |
| HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python \ | |
| Swift-1.5-5bit-MLX/compatibility/cache-tests/validate_full_mac_cache.py \ | |
| --snapshot Swift-1.5-5bit-MLX --output full-mac-5bit-results | |
| ``` | |
| Output directories must be new. The harness uses the real server and complete | |
| released weights, without building or quantizing any synthetic model. It suppresses | |
| the server startup helper's wired-limit call. The unchanged stock BatchGenerator | |
| still calls MLX's set_wired_limit with Apple's recommended working-set size; | |
| macOS memory settings are not changed. The legacy result field | |
| `memory_limit_overrides=false` denotes no manual override, not suppression of this | |
| normal batch-generator behavior. The actual Metal device is recorded in the results. | |
| These results establish the tested text workload on the stated 48 GiB machine. | |
| They do not establish 24 GiB operation, arbitrary 262k-token workloads, GUI/plugin | |
| integration, image/video chat, speculative MTP generation, or broad output quality. | |
| Existing installations still need to apply the runtime patch and restart the server. | |
| The recorded checks are independently executed tests, not a hosted CI status. | |