Instructions to use ukisai/Swift-1.5-4bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ukisai/Swift-1.5-4bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ukisai/Swift-1.5-4bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ukisai/Swift-1.5-4bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-1.5-4bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ukisai/Swift-1.5-4bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ukisai/Swift-1.5-4bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-1.5-4bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ukisai/Swift-1.5-4bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-1.5-4bit-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-1.5-4bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-1.5-4bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Add a verified Swift runtime installer and explicit server launcher (#2)
Browse files- Add a verified Swift runtime installer and explicit server launcher (47701522752d40e1a5d0c3ce2ee235a897cc08af)
- QUICKSTART.md +98 -0
- README.md +6 -0
- SERVER_CACHE_UPDATE.md +6 -0
- TROUBLESHOOTING.md +6 -0
- UPLOAD_MANIFEST.json +23 -8
- USAGE.md +6 -0
- compatibility/launcher-validation.json +56 -0
- swift_runtime.py +251 -0
QUICKSTART.md
ADDED
|
@@ -0,0 +1,98 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Start Swift with the required MLX runtime
|
| 2 |
+
|
| 3 |
+
`Received 501 parameters not in model: language_model.visual...` means the
|
| 4 |
+
selected loader does not represent this checkpoint's vision parameter tree.
|
| 5 |
+
For example, Homebrew's MLX-LM 0.31.3 runs from its own Python environment and
|
| 6 |
+
does not acquire the Swift patches when you download this model. This error
|
| 7 |
+
happens before generation; changing prompt-cache settings does not fix it.
|
| 8 |
+
|
| 9 |
+
The helper below installs the same pinned architecture and cache patches used
|
| 10 |
+
in [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md). It creates a separate Python
|
| 11 |
+
environment and an explicit launcher. Your cached weights can stay where they are.
|
| 12 |
+
|
| 13 |
+
## 1. Use Python 3.12 on Apple Silicon
|
| 14 |
+
|
| 15 |
+
```bash
|
| 16 |
+
python3.12 --version
|
| 17 |
+
```
|
| 18 |
+
|
| 19 |
+
If that command is missing and you use Homebrew, install it with
|
| 20 |
+
`brew install python@3.12`. Leave your existing Homebrew MLX-LM installation alone.
|
| 21 |
+
Git is also required. The helper checks the Python version and native Apple
|
| 22 |
+
Silicon platform before installing anything.
|
| 23 |
+
|
| 24 |
+
## 2. Install the runtime using your existing model directory
|
| 25 |
+
|
| 26 |
+
Stop your old server. Download only the small setup helper:
|
| 27 |
+
|
| 28 |
+
```bash
|
| 29 |
+
hf download ukisai/Swift-1.5-4bit-MLX swift_runtime.py --local-dir Swift-MLX-setup
|
| 30 |
+
```
|
| 31 |
+
|
| 32 |
+
Set `SWIFT_MODEL` to the complete model directory printed by your earlier
|
| 33 |
+
`hf download` command. It can be an HF cache snapshot or a local download folder.
|
| 34 |
+
This example uses the previously published snapshot in the default HF cache:
|
| 35 |
+
|
| 36 |
+
```bash
|
| 37 |
+
SWIFT_MODEL="$HOME/.cache/huggingface/hub/models--ukisai--Swift-1.5-4bit-MLX/snapshots/730aab9b0395b26f7d9cf4b0dfa6a4c788aff6fd"
|
| 38 |
+
python3.12 Swift-MLX-setup/swift_runtime.py setup --model "$SWIFT_MODEL"
|
| 39 |
+
```
|
| 40 |
+
|
| 41 |
+
If you downloaded with `--local-dir Swift-1.5-4bit-MLX`, set
|
| 42 |
+
`SWIFT_MODEL="$PWD/Swift-1.5-4bit-MLX"` instead. Use your actual path if you customized
|
| 43 |
+
the HF cache location or downloaded a different complete revision.
|
| 44 |
+
|
| 45 |
+
The helper verifies fixed SHA-256 hashes for the patches, fetches official
|
| 46 |
+
MLX-LM revision `c69d1288440a0dc4e6401fc417098b07598dccd5`, applies the required
|
| 47 |
+
architecture and cache patches, and installs pinned packages into
|
| 48 |
+
`~/.local/share/swift15-mlx/4bit/.venv`. For 5-bit it also applies the existing
|
| 49 |
+
5-bit support patch. It checks the selected architecture, server source and
|
| 50 |
+
package versions before creating the launcher.
|
| 51 |
+
|
| 52 |
+
Installation needs internet for source code and Python dependencies. It never
|
| 53 |
+
downloads, converts, edits or re-quantizes the model weights. Setup checks local
|
| 54 |
+
assets and indexed shard presence; use `check_download.py --hash` from
|
| 55 |
+
[USAGE.md](USAGE.md) if you also need to verify every weight byte.
|
| 56 |
+
|
| 57 |
+
## 3. Start with this command every time
|
| 58 |
+
|
| 59 |
+
```bash
|
| 60 |
+
"$HOME/.local/share/swift15-mlx/4bit/serve"
|
| 61 |
+
```
|
| 62 |
+
|
| 63 |
+
This starts the server at `http://127.0.0.1:8080` using the dedicated Python
|
| 64 |
+
environment and your existing local model path. It enables offline Hub and
|
| 65 |
+
Transformers operation. You do not need to activate a virtual environment.
|
| 66 |
+
Keep using this launcher after closing and reopening Terminal; typing the bare
|
| 67 |
+
`mlx_lm.server` command can select Homebrew again.
|
| 68 |
+
|
| 69 |
+
To check the runtime without loading weights:
|
| 70 |
+
|
| 71 |
+
```bash
|
| 72 |
+
"$HOME/.local/share/swift15-mlx/4bit/serve" --check-only
|
| 73 |
+
```
|
| 74 |
+
|
| 75 |
+
Additional server arguments work, for example `serve --port 8081` using the same
|
| 76 |
+
full launcher path. The tested cache defaults remain enabled. A previous
|
| 77 |
+
`--prompt-cache-size 0` override disables reuse if you add it again.
|
| 78 |
+
|
| 79 |
+
If the runtime directory already exists, use its launcher. The installer refuses
|
| 80 |
+
to overwrite an existing directory; choose a new path with `--runtime-dir` if
|
| 81 |
+
you need a separate installation, then use the launcher path it prints.
|
| 82 |
+
Stop on an installation failure. A successful `--check-only` verifies the runtime
|
| 83 |
+
and local file presence, not full-model inference or available memory.
|
| 84 |
+
|
| 85 |
+
## Validation
|
| 86 |
+
|
| 87 |
+
Both new 4-bit and 5-bit installations passed native Apple Silicon checks using
|
| 88 |
+
existing tiny synthetic checkpoints containing text, vision and MTP entries.
|
| 89 |
+
The actual launchers served two HTTP generation requests each while a conflicting
|
| 90 |
+
`mlx_lm.server` and Python module were placed on the search path. Paths containing
|
| 91 |
+
spaces worked. An unpatched loader was rejected before loading weights, and a
|
| 92 |
+
repeat installation did not overwrite an existing runtime.
|
| 93 |
+
|
| 94 |
+
[Recorded results](compatibility/launcher-validation.json) cover this installer
|
| 95 |
+
and launch path. The earlier complete-model 48 GiB Mac tests remain in
|
| 96 |
+
[FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md). No full-model rerun or hosted CI
|
| 97 |
+
badge is claimed for this packaging update. Memory capacity, GUI integration,
|
| 98 |
+
integrated image/video chat and speculative MTP limitations are unchanged.
|
README.md
CHANGED
|
@@ -34,6 +34,12 @@ converted with the official Apple MLX-LM converter using 4 bits and group size 6
|
|
| 34 |
Swift 1.5 is UkisAI's reasoning-efficient Qwen3.8-27B derivative, focused on stronger
|
| 35 |
long-horizon, agentic and coding performance while using fewer thinking tokens.
|
| 36 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
**Runtime compatibility:** this complete checkpoint requires the supplied architecture patch
|
| 38 |
and the pinned Python installation in [USAGE.md](USAGE.md). The tested unpatched
|
| 39 |
MLX-LM 0.32.0 loader rejects 501 saved vision entries. Downloading the model does
|
|
|
|
| 34 |
Swift 1.5 is UkisAI's reasoning-efficient Qwen3.8-27B derivative, focused on stronger
|
| 35 |
long-horizon, agentic and coding performance while using fewer thinking tokens.
|
| 36 |
|
| 37 |
+
**Starting from Homebrew or seeing `Received 501 parameters not in model`?**
|
| 38 |
+
Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an
|
| 39 |
+
explicit `serve` launcher. It reuses your existing model directory, including an
|
| 40 |
+
HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew
|
| 41 |
+
installation that lacks the Swift architecture and cache patches.
|
| 42 |
+
|
| 43 |
**Runtime compatibility:** this complete checkpoint requires the supplied architecture patch
|
| 44 |
and the pinned Python installation in [USAGE.md](USAGE.md). The tested unpatched
|
| 45 |
MLX-LM 0.32.0 loader rejects 501 saved vision entries. Downloading the model does
|
SERVER_CACHE_UPDATE.md
CHANGED
|
@@ -1,5 +1,11 @@
|
|
| 1 |
# Server cache update
|
| 2 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
This update changes the pinned MLX-LM server, not the checkpoint. It is a Swift
|
| 4 |
runtime patch based on official MLX-LM `c69d1288440a0dc4e6401fc417098b07598dccd5`, not an upstream release.
|
| 5 |
The existing architecture patch remains required.
|
|
|
|
| 1 |
# Server cache update
|
| 2 |
|
| 3 |
+
**Starting from Homebrew or seeing `Received 501 parameters not in model`?**
|
| 4 |
+
Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an
|
| 5 |
+
explicit `serve` launcher. It reuses your existing model directory, including an
|
| 6 |
+
HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew
|
| 7 |
+
installation that lacks the Swift architecture and cache patches.
|
| 8 |
+
|
| 9 |
This update changes the pinned MLX-LM server, not the checkpoint. It is a Swift
|
| 10 |
runtime patch based on official MLX-LM `c69d1288440a0dc4e6401fc417098b07598dccd5`, not an upstream release.
|
| 11 |
The existing architecture patch remains required.
|
TROUBLESHOOTING.md
CHANGED
|
@@ -1,5 +1,11 @@
|
|
| 1 |
# Check this download and its runtime
|
| 2 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
This release needs **all 3 shards (15.83 GB (14.74 GiB))**, its index/config/tokenizer
|
| 4 |
assets, and the supplied architecture patch. Downloading an HF repository does not install
|
| 5 |
the MLX-LM patches into Python or into a GUI app's separate runtime.
|
|
|
|
| 1 |
# Check this download and its runtime
|
| 2 |
|
| 3 |
+
**Starting from Homebrew or seeing `Received 501 parameters not in model`?**
|
| 4 |
+
Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an
|
| 5 |
+
explicit `serve` launcher. It reuses your existing model directory, including an
|
| 6 |
+
HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew
|
| 7 |
+
installation that lacks the Swift architecture and cache patches.
|
| 8 |
+
|
| 9 |
This release needs **all 3 shards (15.83 GB (14.74 GiB))**, its index/config/tokenizer
|
| 10 |
assets, and the supplied architecture patch. Downloading an HF repository does not install
|
| 11 |
the MLX-LM patches into Python or into a GUI app's separate runtime.
|
UPLOAD_MANIFEST.json
CHANGED
|
@@ -34,25 +34,30 @@
|
|
| 34 |
"bytes": 8230,
|
| 35 |
"sha256": "9f061d15042d9368e2c6a406ee20071ba6d78ce8da4c3ca53bfc13f5c2f95dd1"
|
| 36 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
{
|
| 38 |
"path": "README.md",
|
| 39 |
-
"bytes":
|
| 40 |
-
"sha256": "
|
| 41 |
},
|
| 42 |
{
|
| 43 |
"path": "SERVER_CACHE_UPDATE.md",
|
| 44 |
-
"bytes":
|
| 45 |
-
"sha256": "
|
| 46 |
},
|
| 47 |
{
|
| 48 |
"path": "TROUBLESHOOTING.md",
|
| 49 |
-
"bytes":
|
| 50 |
-
"sha256": "
|
| 51 |
},
|
| 52 |
{
|
| 53 |
"path": "USAGE.md",
|
| 54 |
-
"bytes":
|
| 55 |
-
"sha256": "
|
| 56 |
},
|
| 57 |
{
|
| 58 |
"path": "chat_template.jinja",
|
|
@@ -169,6 +174,11 @@
|
|
| 169 |
"bytes": 113,
|
| 170 |
"sha256": "a40772b0d08d9f5ce659ee06cd5e5573ec0697073295d819a5dc9b6b1f4302bd"
|
| 171 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 172 |
{
|
| 173 |
"path": "compatibility/mac-check/README.md",
|
| 174 |
"bytes": 759,
|
|
@@ -314,6 +324,11 @@
|
|
| 314 |
"bytes": 4738825,
|
| 315 |
"sha256": "1cda6924169e8b83e5d6d7299baf8329ed1b1a182a481aa0e06344e0f5c00111"
|
| 316 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 317 |
{
|
| 318 |
"path": "tokenizer.json",
|
| 319 |
"bytes": 12809320,
|
|
|
|
| 34 |
"bytes": 8230,
|
| 35 |
"sha256": "9f061d15042d9368e2c6a406ee20071ba6d78ce8da4c3ca53bfc13f5c2f95dd1"
|
| 36 |
},
|
| 37 |
+
{
|
| 38 |
+
"path": "QUICKSTART.md",
|
| 39 |
+
"bytes": 4671,
|
| 40 |
+
"sha256": "42046f33ae7e07154c89cacb32fe1569e4931dc82177406b03f5da91d28190cc"
|
| 41 |
+
},
|
| 42 |
{
|
| 43 |
"path": "README.md",
|
| 44 |
+
"bytes": 10794,
|
| 45 |
+
"sha256": "e30be618a534d318226726f04f7868242fc19550714dab1d617f48e7e3e35a45"
|
| 46 |
},
|
| 47 |
{
|
| 48 |
"path": "SERVER_CACHE_UPDATE.md",
|
| 49 |
+
"bytes": 5337,
|
| 50 |
+
"sha256": "a6f648d3491eb4e9bd6f9c3291b38bc33aab1e64030e308b96a843f747b1cfc3"
|
| 51 |
},
|
| 52 |
{
|
| 53 |
"path": "TROUBLESHOOTING.md",
|
| 54 |
+
"bytes": 6339,
|
| 55 |
+
"sha256": "8a0c5babebad8cc24ef5f3f7df3189806318d04dee723c145a67f5bc667343ae"
|
| 56 |
},
|
| 57 |
{
|
| 58 |
"path": "USAGE.md",
|
| 59 |
+
"bytes": 8387,
|
| 60 |
+
"sha256": "bbf3b30ad7cf5b6453ab485fbade6018cd736523f92e2de11560e5bb95c4bb39"
|
| 61 |
},
|
| 62 |
{
|
| 63 |
"path": "chat_template.jinja",
|
|
|
|
| 174 |
"bytes": 113,
|
| 175 |
"sha256": "a40772b0d08d9f5ce659ee06cd5e5573ec0697073295d819a5dc9b6b1f4302bd"
|
| 176 |
},
|
| 177 |
+
{
|
| 178 |
+
"path": "compatibility/launcher-validation.json",
|
| 179 |
+
"bytes": 1598,
|
| 180 |
+
"sha256": "30a08c065138836fbf9fa3c4b586e303742b5a7a155294177a15fc23e0efef50"
|
| 181 |
+
},
|
| 182 |
{
|
| 183 |
"path": "compatibility/mac-check/README.md",
|
| 184 |
"bytes": 759,
|
|
|
|
| 324 |
"bytes": 4738825,
|
| 325 |
"sha256": "1cda6924169e8b83e5d6d7299baf8329ed1b1a182a481aa0e06344e0f5c00111"
|
| 326 |
},
|
| 327 |
+
{
|
| 328 |
+
"path": "swift_runtime.py",
|
| 329 |
+
"bytes": 9029,
|
| 330 |
+
"sha256": "7a459033c6c419e3f7e8cabf59b303767d002039ef5e0a2e2999521951716cbf"
|
| 331 |
+
},
|
| 332 |
{
|
| 333 |
"path": "tokenizer.json",
|
| 334 |
"bytes": 12809320,
|
USAGE.md
CHANGED
|
@@ -1,5 +1,11 @@
|
|
| 1 |
# Load Swift 1.5 with its complete MLX architecture
|
| 2 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
Use the included patch with the pinned official Apple MLX-LM revision. Unpatched
|
| 4 |
text-only Qwen support does not preserve this checkpoint's complete parameter tree.
|
| 5 |
|
|
|
|
| 1 |
# Load Swift 1.5 with its complete MLX architecture
|
| 2 |
|
| 3 |
+
**Starting from Homebrew or seeing `Received 501 parameters not in model`?**
|
| 4 |
+
Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an
|
| 5 |
+
explicit `serve` launcher. It reuses your existing model directory, including an
|
| 6 |
+
HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew
|
| 7 |
+
installation that lacks the Swift architecture and cache patches.
|
| 8 |
+
|
| 9 |
Use the included patch with the pinned official Apple MLX-LM revision. Unpatched
|
| 10 |
text-only Qwen support does not preserve this checkpoint's complete parameter tree.
|
| 11 |
|
compatibility/launcher-validation.json
ADDED
|
@@ -0,0 +1,56 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"status": "PASS_LAUNCHER_INSTALL_AND_SYNTHETIC_HTTP",
|
| 3 |
+
"date": "2026-09-26",
|
| 4 |
+
"scope": "New isolated installation and real CLI launch on Apple Silicon with existing tiny synthetic 4-bit and 5-bit checkpoints. No full Swift weights loaded or requantized.",
|
| 5 |
+
"results": [
|
| 6 |
+
{
|
| 7 |
+
"bits": 4,
|
| 8 |
+
"clean_install": "PASS",
|
| 9 |
+
"runtime_identity": "PASS",
|
| 10 |
+
"path_shadowing_bypassed": "PASS",
|
| 11 |
+
"pythonpath_shadowing_bypassed": "PASS",
|
| 12 |
+
"paths_with_spaces": "PASS",
|
| 13 |
+
"fixture_tensor_entries": 190,
|
| 14 |
+
"vision_entries": 37,
|
| 15 |
+
"mtp_entries": 31,
|
| 16 |
+
"http_requests": [
|
| 17 |
+
{
|
| 18 |
+
"http_status": 200,
|
| 19 |
+
"completion_tokens": 4
|
| 20 |
+
},
|
| 21 |
+
{
|
| 22 |
+
"http_status": 200,
|
| 23 |
+
"completion_tokens": 4
|
| 24 |
+
}
|
| 25 |
+
],
|
| 26 |
+
"worker_alive": true
|
| 27 |
+
},
|
| 28 |
+
{
|
| 29 |
+
"bits": 5,
|
| 30 |
+
"clean_install": "PASS",
|
| 31 |
+
"runtime_identity": "PASS",
|
| 32 |
+
"path_shadowing_bypassed": "PASS",
|
| 33 |
+
"pythonpath_shadowing_bypassed": "PASS",
|
| 34 |
+
"paths_with_spaces": "PASS",
|
| 35 |
+
"fixture_tensor_entries": 190,
|
| 36 |
+
"vision_entries": 37,
|
| 37 |
+
"mtp_entries": 31,
|
| 38 |
+
"http_requests": [
|
| 39 |
+
{
|
| 40 |
+
"http_status": 200,
|
| 41 |
+
"completion_tokens": 4
|
| 42 |
+
},
|
| 43 |
+
{
|
| 44 |
+
"http_status": 200,
|
| 45 |
+
"completion_tokens": 4
|
| 46 |
+
}
|
| 47 |
+
],
|
| 48 |
+
"worker_alive": true
|
| 49 |
+
}
|
| 50 |
+
],
|
| 51 |
+
"negative_checks": {
|
| 52 |
+
"stock_loader_rejected_before_loading": "PASS",
|
| 53 |
+
"existing_install_not_overwritten": "PASS"
|
| 54 |
+
},
|
| 55 |
+
"full_model_validation": "Previously recorded separately in FULL_MAC_VALIDATION.md; not rerun for this launcher."
|
| 56 |
+
}
|
swift_runtime.py
ADDED
|
@@ -0,0 +1,251 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Install and launch the pinned Swift MLX runtime using an existing local model."""
|
| 2 |
+
|
| 3 |
+
import argparse
|
| 4 |
+
import hashlib
|
| 5 |
+
import inspect
|
| 6 |
+
import json
|
| 7 |
+
import os
|
| 8 |
+
import platform
|
| 9 |
+
import shlex
|
| 10 |
+
import shutil
|
| 11 |
+
import subprocess
|
| 12 |
+
import sys
|
| 13 |
+
import urllib.request
|
| 14 |
+
import venv
|
| 15 |
+
from pathlib import Path
|
| 16 |
+
|
| 17 |
+
BASE = "c69d1288440a0dc4e6401fc417098b07598dccd5"
|
| 18 |
+
REVISIONS = {
|
| 19 |
+
4: "730aab9b0395b26f7d9cf4b0dfa6a4c788aff6fd",
|
| 20 |
+
5: "8aff72b145212e62c15146e41dcbef35ea5fa9ff",
|
| 21 |
+
}
|
| 22 |
+
PATCHES = {
|
| 23 |
+
"swift15-mlx-lm.patch": "f6f1d0bdafa45863bfbf93dac0398c481c993ea04fdf38b9bae98c643f89eaec",
|
| 24 |
+
"enable-5bit.patch": "b985961eac3035e05ca4c9f3a8b283c26e6dd69bab4d113997fc04fe2a4f99cc",
|
| 25 |
+
"swift15-server-cache.patch": "a7fbc0f0524ee7864d9f41a98a2e35acf0e63f53e369ee5b9a87ca12926a9eab",
|
| 26 |
+
}
|
| 27 |
+
MODEL_HASHES = {
|
| 28 |
+
4: "f559e4559946bfad4ce20b93906028449b76e007f42f36bb83d2c2c5f6cb95a9",
|
| 29 |
+
5: "ecda666d6d5f1059388a2ce4267d9ff8625398bb925d6c0840d0059ce17b676f",
|
| 30 |
+
}
|
| 31 |
+
SERVER_HASH = "791f8c1eb3431c4d24b4f9084f67b665c38e3d547ee4a307e52c5b2b6f12aba0"
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
def digest(path):
|
| 35 |
+
return hashlib.sha256(Path(path).read_bytes()).hexdigest()
|
| 36 |
+
|
| 37 |
+
|
| 38 |
+
def local_config(model):
|
| 39 |
+
config = json.loads((model / "config.json").read_text())
|
| 40 |
+
quant = config.get("quantization", {})
|
| 41 |
+
bits = quant.get("bits")
|
| 42 |
+
if bits not in REVISIONS or quant != {
|
| 43 |
+
"bits": bits,
|
| 44 |
+
"group_size": 64,
|
| 45 |
+
"mode": "affine",
|
| 46 |
+
}:
|
| 47 |
+
raise ValueError(
|
| 48 |
+
"Expected the Swift 1.5 affine 4-bit or 5-bit / group-size-64 checkpoint."
|
| 49 |
+
)
|
| 50 |
+
if config.get("language_model_only") is not False:
|
| 51 |
+
raise ValueError("Expected the complete Swift architecture configuration.")
|
| 52 |
+
for name in ("tokenizer.json", "tokenizer_config.json", "chat_template.jinja"):
|
| 53 |
+
if not (model / name).is_file():
|
| 54 |
+
raise ValueError(f"Missing local model asset: {name}")
|
| 55 |
+
index = json.loads((model / "model.safetensors.index.json").read_text())
|
| 56 |
+
shards = set(index["weight_map"].values())
|
| 57 |
+
if not shards:
|
| 58 |
+
raise ValueError("The checkpoint index is empty.")
|
| 59 |
+
for name in shards:
|
| 60 |
+
if Path(name).name != name or not name.endswith(".safetensors"):
|
| 61 |
+
raise ValueError("The checkpoint index contains an invalid shard path.")
|
| 62 |
+
if not (model / name).is_file():
|
| 63 |
+
raise ValueError(
|
| 64 |
+
f"Missing local weight shard: {name}; complete the existing download first."
|
| 65 |
+
)
|
| 66 |
+
return config, bits
|
| 67 |
+
|
| 68 |
+
|
| 69 |
+
def runtime_check(config, bits):
|
| 70 |
+
from importlib.metadata import version
|
| 71 |
+
|
| 72 |
+
from mlx_lm import server
|
| 73 |
+
from mlx_lm.utils import _get_classes
|
| 74 |
+
|
| 75 |
+
cls, args = _get_classes(config)
|
| 76 |
+
args.from_dict(config)
|
| 77 |
+
if cls.__module__ != "mlx_lm.models.qwen3_5_full":
|
| 78 |
+
raise ValueError(
|
| 79 |
+
"Swift architecture patch is missing. Use the serve launcher created by setup."
|
| 80 |
+
)
|
| 81 |
+
if digest(inspect.getsourcefile(cls)) != MODEL_HASHES[bits]:
|
| 82 |
+
raise ValueError(
|
| 83 |
+
"The installed architecture differs from the pinned Swift implementation."
|
| 84 |
+
)
|
| 85 |
+
if digest(server.__file__) != SERVER_HASH:
|
| 86 |
+
raise ValueError("The tested server cache patch is missing or has changed.")
|
| 87 |
+
versions = {
|
| 88 |
+
name: version(name)
|
| 89 |
+
for name in ("mlx", "mlx-metal", "mlx-lm", "transformers", "huggingface_hub")
|
| 90 |
+
}
|
| 91 |
+
expected = {
|
| 92 |
+
"mlx": "0.32.2",
|
| 93 |
+
"mlx-metal": "0.32.2",
|
| 94 |
+
"mlx-lm": "0.32.0",
|
| 95 |
+
"transformers": "5.14.1",
|
| 96 |
+
"huggingface_hub": "1.31.0",
|
| 97 |
+
}
|
| 98 |
+
if versions != expected:
|
| 99 |
+
raise ValueError(
|
| 100 |
+
f"Installed package versions differ from the tested runtime: {versions}"
|
| 101 |
+
)
|
| 102 |
+
print(
|
| 103 |
+
json.dumps(
|
| 104 |
+
{
|
| 105 |
+
"status": "PASS_RUNTIME_IDENTITY",
|
| 106 |
+
"python": sys.executable,
|
| 107 |
+
"model_class": cls.__module__ + "." + cls.__name__,
|
| 108 |
+
"server": server.__file__,
|
| 109 |
+
"versions": versions,
|
| 110 |
+
"weights_loaded": False,
|
| 111 |
+
}
|
| 112 |
+
),
|
| 113 |
+
flush=True,
|
| 114 |
+
)
|
| 115 |
+
|
| 116 |
+
|
| 117 |
+
def setup(model, bits, destination):
|
| 118 |
+
if shutil.which("git") is None:
|
| 119 |
+
raise ValueError(
|
| 120 |
+
"Git is required. Install the macOS command-line tools, then retry."
|
| 121 |
+
)
|
| 122 |
+
if destination.exists():
|
| 123 |
+
raise ValueError(
|
| 124 |
+
f"Runtime directory already exists: {destination}. Use its serve launcher, or choose a new --runtime-dir."
|
| 125 |
+
)
|
| 126 |
+
destination.mkdir(parents=True)
|
| 127 |
+
source = destination / "mlx-lm"
|
| 128 |
+
envdir = destination / ".venv"
|
| 129 |
+
python = envdir / "bin/python"
|
| 130 |
+
patchdir = destination / "patches"
|
| 131 |
+
patchdir.mkdir()
|
| 132 |
+
patch_names = ["swift15-mlx-lm.patch"]
|
| 133 |
+
if bits == 5:
|
| 134 |
+
patch_names.append("enable-5bit.patch")
|
| 135 |
+
patch_names.append("swift15-server-cache.patch")
|
| 136 |
+
for name in patch_names:
|
| 137 |
+
url = f"https://huggingface.co/ukisai/Swift-1.5-{bits}bit-MLX/resolve/{REVISIONS[bits]}/compatibility/{name}"
|
| 138 |
+
with urllib.request.urlopen(url, timeout=60) as response:
|
| 139 |
+
data = response.read()
|
| 140 |
+
if hashlib.sha256(data).hexdigest() != PATCHES[name]:
|
| 141 |
+
raise ValueError(f"Patch checksum mismatch: {name}")
|
| 142 |
+
(patchdir / name).write_bytes(data)
|
| 143 |
+
subprocess.run(["git", "init", "--quiet", str(source)], check=True)
|
| 144 |
+
git = ["git", "-C", str(source)]
|
| 145 |
+
subprocess.run(
|
| 146 |
+
git + ["remote", "add", "origin", "https://github.com/ml-explore/mlx-lm.git"],
|
| 147 |
+
check=True,
|
| 148 |
+
)
|
| 149 |
+
subprocess.run(git + ["fetch", "--depth", "1", "origin", BASE], check=True)
|
| 150 |
+
subprocess.run(git + ["checkout", "--detach", BASE], check=True)
|
| 151 |
+
for name in patch_names:
|
| 152 |
+
subprocess.run(git + ["apply", "--check", str(patchdir / name)], check=True)
|
| 153 |
+
subprocess.run(git + ["apply", str(patchdir / name)], check=True)
|
| 154 |
+
venv.EnvBuilder(with_pip=True).create(envdir)
|
| 155 |
+
subprocess.run(
|
| 156 |
+
[
|
| 157 |
+
str(python),
|
| 158 |
+
"-I",
|
| 159 |
+
"-m",
|
| 160 |
+
"pip",
|
| 161 |
+
"install",
|
| 162 |
+
"mlx==0.32.2",
|
| 163 |
+
"mlx-metal==0.32.2",
|
| 164 |
+
"transformers==5.14.1",
|
| 165 |
+
"huggingface_hub==1.31.0",
|
| 166 |
+
"pillow==12.3.0",
|
| 167 |
+
"safetensors==0.8.0",
|
| 168 |
+
"-e",
|
| 169 |
+
str(source),
|
| 170 |
+
],
|
| 171 |
+
check=True,
|
| 172 |
+
)
|
| 173 |
+
subprocess.run([str(python), "-I", "-m", "pip", "check"], check=True)
|
| 174 |
+
installed = destination / "swift_runtime.py"
|
| 175 |
+
shutil.copyfile(__file__, installed)
|
| 176 |
+
command = [str(python), "-I", str(installed), "start", "--model", str(model)]
|
| 177 |
+
subprocess.run(command + ["--check-only"], check=True)
|
| 178 |
+
launcher = destination / "serve"
|
| 179 |
+
launcher.write_text("#!/bin/sh\nexec " + shlex.join(command) + ' "$@"\n')
|
| 180 |
+
launcher.chmod(0o755)
|
| 181 |
+
print(
|
| 182 |
+
f"\nSetup complete. Start or restart with this exact command:\n{shlex.quote(str(launcher))}"
|
| 183 |
+
)
|
| 184 |
+
|
| 185 |
+
|
| 186 |
+
def main():
|
| 187 |
+
parser = argparse.ArgumentParser(description=__doc__)
|
| 188 |
+
sub = parser.add_subparsers(dest="action", required=True)
|
| 189 |
+
install = sub.add_parser(
|
| 190 |
+
"setup",
|
| 191 |
+
help="Install into a new isolated directory; reuse existing model files.",
|
| 192 |
+
)
|
| 193 |
+
install.add_argument("--model", required=True, type=Path)
|
| 194 |
+
install.add_argument("--runtime-dir", type=Path)
|
| 195 |
+
start = sub.add_parser(
|
| 196 |
+
"start", help="Verify the selected Python runtime and serve a local checkpoint."
|
| 197 |
+
)
|
| 198 |
+
start.add_argument("--model", required=True, type=Path)
|
| 199 |
+
start.add_argument("--check-only", action="store_true")
|
| 200 |
+
args, extra = parser.parse_known_args()
|
| 201 |
+
if args.action == "setup" and extra:
|
| 202 |
+
parser.error("Unknown setup arguments: " + " ".join(extra))
|
| 203 |
+
if sys.version_info[:2] != (3, 12):
|
| 204 |
+
parser.error(
|
| 205 |
+
"Use python3.12 for this tested runtime. Homebrew's Python 3.14 environment is separate."
|
| 206 |
+
)
|
| 207 |
+
if platform.system() != "Darwin" or platform.machine() != "arm64":
|
| 208 |
+
parser.error("This installer and launcher target native Apple Silicon macOS.")
|
| 209 |
+
model = args.model.expanduser().resolve()
|
| 210 |
+
try:
|
| 211 |
+
config, bits = local_config(model)
|
| 212 |
+
if args.action == "setup":
|
| 213 |
+
target = (
|
| 214 |
+
args.runtime_dir
|
| 215 |
+
or Path.home() / ".local/share/swift15-mlx" / f"{bits}bit"
|
| 216 |
+
)
|
| 217 |
+
setup(model, bits, target.expanduser().resolve())
|
| 218 |
+
return
|
| 219 |
+
os.environ["HF_HUB_OFFLINE"] = "1"
|
| 220 |
+
os.environ["TRANSFORMERS_OFFLINE"] = "1"
|
| 221 |
+
runtime_check(config, bits)
|
| 222 |
+
if args.check_only:
|
| 223 |
+
return
|
| 224 |
+
os.execv(
|
| 225 |
+
sys.executable,
|
| 226 |
+
[
|
| 227 |
+
sys.executable,
|
| 228 |
+
"-I",
|
| 229 |
+
"-m",
|
| 230 |
+
"mlx_lm.server",
|
| 231 |
+
"--model",
|
| 232 |
+
str(model),
|
| 233 |
+
"--host",
|
| 234 |
+
"127.0.0.1",
|
| 235 |
+
"--port",
|
| 236 |
+
"8080",
|
| 237 |
+
*extra,
|
| 238 |
+
],
|
| 239 |
+
)
|
| 240 |
+
except (
|
| 241 |
+
OSError,
|
| 242 |
+
ValueError,
|
| 243 |
+
KeyError,
|
| 244 |
+
ImportError,
|
| 245 |
+
subprocess.CalledProcessError,
|
| 246 |
+
) as exc:
|
| 247 |
+
raise SystemExit(f"Swift runtime setup/start failed: {exc}") from None
|
| 248 |
+
|
| 249 |
+
|
| 250 |
+
if __name__ == "__main__":
|
| 251 |
+
main()
|