Add a verified Swift runtime installer and explicit server launcher

#2
by ukisai - opened
QUICKSTART.md ADDED
@@ -0,0 +1,98 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Start Swift with the required MLX runtime
2
+
3
+ `Received 501 parameters not in model: language_model.visual...` means the
4
+ selected loader does not represent this checkpoint's vision parameter tree.
5
+ For example, Homebrew's MLX-LM 0.31.3 runs from its own Python environment and
6
+ does not acquire the Swift patches when you download this model. This error
7
+ happens before generation; changing prompt-cache settings does not fix it.
8
+
9
+ The helper below installs the same pinned architecture and cache patches used
10
+ in [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md). It creates a separate Python
11
+ environment and an explicit launcher. Your cached weights can stay where they are.
12
+
13
+ ## 1. Use Python 3.12 on Apple Silicon
14
+
15
+ ```bash
16
+ python3.12 --version
17
+ ```
18
+
19
+ If that command is missing and you use Homebrew, install it with
20
+ `brew install python@3.12`. Leave your existing Homebrew MLX-LM installation alone.
21
+ Git is also required. The helper checks the Python version and native Apple
22
+ Silicon platform before installing anything.
23
+
24
+ ## 2. Install the runtime using your existing model directory
25
+
26
+ Stop your old server. Download only the small setup helper:
27
+
28
+ ```bash
29
+ hf download ukisai/Swift-1.5-4bit-MLX swift_runtime.py --local-dir Swift-MLX-setup
30
+ ```
31
+
32
+ Set `SWIFT_MODEL` to the complete model directory printed by your earlier
33
+ `hf download` command. It can be an HF cache snapshot or a local download folder.
34
+ This example uses the previously published snapshot in the default HF cache:
35
+
36
+ ```bash
37
+ SWIFT_MODEL="$HOME/.cache/huggingface/hub/models--ukisai--Swift-1.5-4bit-MLX/snapshots/730aab9b0395b26f7d9cf4b0dfa6a4c788aff6fd"
38
+ python3.12 Swift-MLX-setup/swift_runtime.py setup --model "$SWIFT_MODEL"
39
+ ```
40
+
41
+ If you downloaded with `--local-dir Swift-1.5-4bit-MLX`, set
42
+ `SWIFT_MODEL="$PWD/Swift-1.5-4bit-MLX"` instead. Use your actual path if you customized
43
+ the HF cache location or downloaded a different complete revision.
44
+
45
+ The helper verifies fixed SHA-256 hashes for the patches, fetches official
46
+ MLX-LM revision `c69d1288440a0dc4e6401fc417098b07598dccd5`, applies the required
47
+ architecture and cache patches, and installs pinned packages into
48
+ `~/.local/share/swift15-mlx/4bit/.venv`. For 5-bit it also applies the existing
49
+ 5-bit support patch. It checks the selected architecture, server source and
50
+ package versions before creating the launcher.
51
+
52
+ Installation needs internet for source code and Python dependencies. It never
53
+ downloads, converts, edits or re-quantizes the model weights. Setup checks local
54
+ assets and indexed shard presence; use `check_download.py --hash` from
55
+ [USAGE.md](USAGE.md) if you also need to verify every weight byte.
56
+
57
+ ## 3. Start with this command every time
58
+
59
+ ```bash
60
+ "$HOME/.local/share/swift15-mlx/4bit/serve"
61
+ ```
62
+
63
+ This starts the server at `http://127.0.0.1:8080` using the dedicated Python
64
+ environment and your existing local model path. It enables offline Hub and
65
+ Transformers operation. You do not need to activate a virtual environment.
66
+ Keep using this launcher after closing and reopening Terminal; typing the bare
67
+ `mlx_lm.server` command can select Homebrew again.
68
+
69
+ To check the runtime without loading weights:
70
+
71
+ ```bash
72
+ "$HOME/.local/share/swift15-mlx/4bit/serve" --check-only
73
+ ```
74
+
75
+ Additional server arguments work, for example `serve --port 8081` using the same
76
+ full launcher path. The tested cache defaults remain enabled. A previous
77
+ `--prompt-cache-size 0` override disables reuse if you add it again.
78
+
79
+ If the runtime directory already exists, use its launcher. The installer refuses
80
+ to overwrite an existing directory; choose a new path with `--runtime-dir` if
81
+ you need a separate installation, then use the launcher path it prints.
82
+ Stop on an installation failure. A successful `--check-only` verifies the runtime
83
+ and local file presence, not full-model inference or available memory.
84
+
85
+ ## Validation
86
+
87
+ Both new 4-bit and 5-bit installations passed native Apple Silicon checks using
88
+ existing tiny synthetic checkpoints containing text, vision and MTP entries.
89
+ The actual launchers served two HTTP generation requests each while a conflicting
90
+ `mlx_lm.server` and Python module were placed on the search path. Paths containing
91
+ spaces worked. An unpatched loader was rejected before loading weights, and a
92
+ repeat installation did not overwrite an existing runtime.
93
+
94
+ [Recorded results](compatibility/launcher-validation.json) cover this installer
95
+ and launch path. The earlier complete-model 48 GiB Mac tests remain in
96
+ [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md). No full-model rerun or hosted CI
97
+ badge is claimed for this packaging update. Memory capacity, GUI integration,
98
+ integrated image/video chat and speculative MTP limitations are unchanged.
README.md CHANGED
@@ -34,6 +34,12 @@ converted with the official Apple MLX-LM converter using 4 bits and group size 6
34
  Swift 1.5 is UkisAI's reasoning-efficient Qwen3.8-27B derivative, focused on stronger
35
  long-horizon, agentic and coding performance while using fewer thinking tokens.
36
 
 
 
 
 
 
 
37
  **Runtime compatibility:** this complete checkpoint requires the supplied architecture patch
38
  and the pinned Python installation in [USAGE.md](USAGE.md). The tested unpatched
39
  MLX-LM 0.32.0 loader rejects 501 saved vision entries. Downloading the model does
 
34
  Swift 1.5 is UkisAI's reasoning-efficient Qwen3.8-27B derivative, focused on stronger
35
  long-horizon, agentic and coding performance while using fewer thinking tokens.
36
 
37
+ **Starting from Homebrew or seeing `Received 501 parameters not in model`?**
38
+ Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an
39
+ explicit `serve` launcher. It reuses your existing model directory, including an
40
+ HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew
41
+ installation that lacks the Swift architecture and cache patches.
42
+
43
  **Runtime compatibility:** this complete checkpoint requires the supplied architecture patch
44
  and the pinned Python installation in [USAGE.md](USAGE.md). The tested unpatched
45
  MLX-LM 0.32.0 loader rejects 501 saved vision entries. Downloading the model does
SERVER_CACHE_UPDATE.md CHANGED
@@ -1,5 +1,11 @@
1
  # Server cache update
2
 
 
 
 
 
 
 
3
  This update changes the pinned MLX-LM server, not the checkpoint. It is a Swift
4
  runtime patch based on official MLX-LM `c69d1288440a0dc4e6401fc417098b07598dccd5`, not an upstream release.
5
  The existing architecture patch remains required.
 
1
  # Server cache update
2
 
3
+ **Starting from Homebrew or seeing `Received 501 parameters not in model`?**
4
+ Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an
5
+ explicit `serve` launcher. It reuses your existing model directory, including an
6
+ HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew
7
+ installation that lacks the Swift architecture and cache patches.
8
+
9
  This update changes the pinned MLX-LM server, not the checkpoint. It is a Swift
10
  runtime patch based on official MLX-LM `c69d1288440a0dc4e6401fc417098b07598dccd5`, not an upstream release.
11
  The existing architecture patch remains required.
TROUBLESHOOTING.md CHANGED
@@ -1,5 +1,11 @@
1
  # Check this download and its runtime
2
 
 
 
 
 
 
 
3
  This release needs **all 3 shards (15.83 GB (14.74 GiB))**, its index/config/tokenizer
4
  assets, and the supplied architecture patch. Downloading an HF repository does not install
5
  the MLX-LM patches into Python or into a GUI app's separate runtime.
 
1
  # Check this download and its runtime
2
 
3
+ **Starting from Homebrew or seeing `Received 501 parameters not in model`?**
4
+ Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an
5
+ explicit `serve` launcher. It reuses your existing model directory, including an
6
+ HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew
7
+ installation that lacks the Swift architecture and cache patches.
8
+
9
  This release needs **all 3 shards (15.83 GB (14.74 GiB))**, its index/config/tokenizer
10
  assets, and the supplied architecture patch. Downloading an HF repository does not install
11
  the MLX-LM patches into Python or into a GUI app's separate runtime.
UPLOAD_MANIFEST.json CHANGED
@@ -34,25 +34,30 @@
34
  "bytes": 8230,
35
  "sha256": "9f061d15042d9368e2c6a406ee20071ba6d78ce8da4c3ca53bfc13f5c2f95dd1"
36
  },
 
 
 
 
 
37
  {
38
  "path": "README.md",
39
- "bytes": 10407,
40
- "sha256": "763ade1ad454b132b019879763b41b897cd324519a24fb16b85ee742d9736bce"
41
  },
42
  {
43
  "path": "SERVER_CACHE_UPDATE.md",
44
- "bytes": 4950,
45
- "sha256": "d85f05bc0bd339947b06af7d608b44898ac4b400ef680cfd6a4869910524eba1"
46
  },
47
  {
48
  "path": "TROUBLESHOOTING.md",
49
- "bytes": 5952,
50
- "sha256": "ff92cf87f0bf35caf7bf8962be6eaaea0a132d781ee26fde83d8eae4e9b306d7"
51
  },
52
  {
53
  "path": "USAGE.md",
54
- "bytes": 8000,
55
- "sha256": "8925823cdb938fcc5a597eef121486fbe2b3a593aa2753ec9028c46f9e3f51e2"
56
  },
57
  {
58
  "path": "chat_template.jinja",
@@ -169,6 +174,11 @@
169
  "bytes": 113,
170
  "sha256": "a40772b0d08d9f5ce659ee06cd5e5573ec0697073295d819a5dc9b6b1f4302bd"
171
  },
 
 
 
 
 
172
  {
173
  "path": "compatibility/mac-check/README.md",
174
  "bytes": 759,
@@ -314,6 +324,11 @@
314
  "bytes": 4738825,
315
  "sha256": "1cda6924169e8b83e5d6d7299baf8329ed1b1a182a481aa0e06344e0f5c00111"
316
  },
 
 
 
 
 
317
  {
318
  "path": "tokenizer.json",
319
  "bytes": 12809320,
 
34
  "bytes": 8230,
35
  "sha256": "9f061d15042d9368e2c6a406ee20071ba6d78ce8da4c3ca53bfc13f5c2f95dd1"
36
  },
37
+ {
38
+ "path": "QUICKSTART.md",
39
+ "bytes": 4671,
40
+ "sha256": "42046f33ae7e07154c89cacb32fe1569e4931dc82177406b03f5da91d28190cc"
41
+ },
42
  {
43
  "path": "README.md",
44
+ "bytes": 10794,
45
+ "sha256": "e30be618a534d318226726f04f7868242fc19550714dab1d617f48e7e3e35a45"
46
  },
47
  {
48
  "path": "SERVER_CACHE_UPDATE.md",
49
+ "bytes": 5337,
50
+ "sha256": "a6f648d3491eb4e9bd6f9c3291b38bc33aab1e64030e308b96a843f747b1cfc3"
51
  },
52
  {
53
  "path": "TROUBLESHOOTING.md",
54
+ "bytes": 6339,
55
+ "sha256": "8a0c5babebad8cc24ef5f3f7df3189806318d04dee723c145a67f5bc667343ae"
56
  },
57
  {
58
  "path": "USAGE.md",
59
+ "bytes": 8387,
60
+ "sha256": "bbf3b30ad7cf5b6453ab485fbade6018cd736523f92e2de11560e5bb95c4bb39"
61
  },
62
  {
63
  "path": "chat_template.jinja",
 
174
  "bytes": 113,
175
  "sha256": "a40772b0d08d9f5ce659ee06cd5e5573ec0697073295d819a5dc9b6b1f4302bd"
176
  },
177
+ {
178
+ "path": "compatibility/launcher-validation.json",
179
+ "bytes": 1598,
180
+ "sha256": "30a08c065138836fbf9fa3c4b586e303742b5a7a155294177a15fc23e0efef50"
181
+ },
182
  {
183
  "path": "compatibility/mac-check/README.md",
184
  "bytes": 759,
 
324
  "bytes": 4738825,
325
  "sha256": "1cda6924169e8b83e5d6d7299baf8329ed1b1a182a481aa0e06344e0f5c00111"
326
  },
327
+ {
328
+ "path": "swift_runtime.py",
329
+ "bytes": 9029,
330
+ "sha256": "7a459033c6c419e3f7e8cabf59b303767d002039ef5e0a2e2999521951716cbf"
331
+ },
332
  {
333
  "path": "tokenizer.json",
334
  "bytes": 12809320,
USAGE.md CHANGED
@@ -1,5 +1,11 @@
1
  # Load Swift 1.5 with its complete MLX architecture
2
 
 
 
 
 
 
 
3
  Use the included patch with the pinned official Apple MLX-LM revision. Unpatched
4
  text-only Qwen support does not preserve this checkpoint's complete parameter tree.
5
 
 
1
  # Load Swift 1.5 with its complete MLX architecture
2
 
3
+ **Starting from Homebrew or seeing `Received 501 parameters not in model`?**
4
+ Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an
5
+ explicit `serve` launcher. It reuses your existing model directory, including an
6
+ HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew
7
+ installation that lacks the Swift architecture and cache patches.
8
+
9
  Use the included patch with the pinned official Apple MLX-LM revision. Unpatched
10
  text-only Qwen support does not preserve this checkpoint's complete parameter tree.
11
 
compatibility/launcher-validation.json ADDED
@@ -0,0 +1,56 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "status": "PASS_LAUNCHER_INSTALL_AND_SYNTHETIC_HTTP",
3
+ "date": "2026-09-26",
4
+ "scope": "New isolated installation and real CLI launch on Apple Silicon with existing tiny synthetic 4-bit and 5-bit checkpoints. No full Swift weights loaded or requantized.",
5
+ "results": [
6
+ {
7
+ "bits": 4,
8
+ "clean_install": "PASS",
9
+ "runtime_identity": "PASS",
10
+ "path_shadowing_bypassed": "PASS",
11
+ "pythonpath_shadowing_bypassed": "PASS",
12
+ "paths_with_spaces": "PASS",
13
+ "fixture_tensor_entries": 190,
14
+ "vision_entries": 37,
15
+ "mtp_entries": 31,
16
+ "http_requests": [
17
+ {
18
+ "http_status": 200,
19
+ "completion_tokens": 4
20
+ },
21
+ {
22
+ "http_status": 200,
23
+ "completion_tokens": 4
24
+ }
25
+ ],
26
+ "worker_alive": true
27
+ },
28
+ {
29
+ "bits": 5,
30
+ "clean_install": "PASS",
31
+ "runtime_identity": "PASS",
32
+ "path_shadowing_bypassed": "PASS",
33
+ "pythonpath_shadowing_bypassed": "PASS",
34
+ "paths_with_spaces": "PASS",
35
+ "fixture_tensor_entries": 190,
36
+ "vision_entries": 37,
37
+ "mtp_entries": 31,
38
+ "http_requests": [
39
+ {
40
+ "http_status": 200,
41
+ "completion_tokens": 4
42
+ },
43
+ {
44
+ "http_status": 200,
45
+ "completion_tokens": 4
46
+ }
47
+ ],
48
+ "worker_alive": true
49
+ }
50
+ ],
51
+ "negative_checks": {
52
+ "stock_loader_rejected_before_loading": "PASS",
53
+ "existing_install_not_overwritten": "PASS"
54
+ },
55
+ "full_model_validation": "Previously recorded separately in FULL_MAC_VALIDATION.md; not rerun for this launcher."
56
+ }
swift_runtime.py ADDED
@@ -0,0 +1,251 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Install and launch the pinned Swift MLX runtime using an existing local model."""
2
+
3
+ import argparse
4
+ import hashlib
5
+ import inspect
6
+ import json
7
+ import os
8
+ import platform
9
+ import shlex
10
+ import shutil
11
+ import subprocess
12
+ import sys
13
+ import urllib.request
14
+ import venv
15
+ from pathlib import Path
16
+
17
+ BASE = "c69d1288440a0dc4e6401fc417098b07598dccd5"
18
+ REVISIONS = {
19
+ 4: "730aab9b0395b26f7d9cf4b0dfa6a4c788aff6fd",
20
+ 5: "8aff72b145212e62c15146e41dcbef35ea5fa9ff",
21
+ }
22
+ PATCHES = {
23
+ "swift15-mlx-lm.patch": "f6f1d0bdafa45863bfbf93dac0398c481c993ea04fdf38b9bae98c643f89eaec",
24
+ "enable-5bit.patch": "b985961eac3035e05ca4c9f3a8b283c26e6dd69bab4d113997fc04fe2a4f99cc",
25
+ "swift15-server-cache.patch": "a7fbc0f0524ee7864d9f41a98a2e35acf0e63f53e369ee5b9a87ca12926a9eab",
26
+ }
27
+ MODEL_HASHES = {
28
+ 4: "f559e4559946bfad4ce20b93906028449b76e007f42f36bb83d2c2c5f6cb95a9",
29
+ 5: "ecda666d6d5f1059388a2ce4267d9ff8625398bb925d6c0840d0059ce17b676f",
30
+ }
31
+ SERVER_HASH = "791f8c1eb3431c4d24b4f9084f67b665c38e3d547ee4a307e52c5b2b6f12aba0"
32
+
33
+
34
+ def digest(path):
35
+ return hashlib.sha256(Path(path).read_bytes()).hexdigest()
36
+
37
+
38
+ def local_config(model):
39
+ config = json.loads((model / "config.json").read_text())
40
+ quant = config.get("quantization", {})
41
+ bits = quant.get("bits")
42
+ if bits not in REVISIONS or quant != {
43
+ "bits": bits,
44
+ "group_size": 64,
45
+ "mode": "affine",
46
+ }:
47
+ raise ValueError(
48
+ "Expected the Swift 1.5 affine 4-bit or 5-bit / group-size-64 checkpoint."
49
+ )
50
+ if config.get("language_model_only") is not False:
51
+ raise ValueError("Expected the complete Swift architecture configuration.")
52
+ for name in ("tokenizer.json", "tokenizer_config.json", "chat_template.jinja"):
53
+ if not (model / name).is_file():
54
+ raise ValueError(f"Missing local model asset: {name}")
55
+ index = json.loads((model / "model.safetensors.index.json").read_text())
56
+ shards = set(index["weight_map"].values())
57
+ if not shards:
58
+ raise ValueError("The checkpoint index is empty.")
59
+ for name in shards:
60
+ if Path(name).name != name or not name.endswith(".safetensors"):
61
+ raise ValueError("The checkpoint index contains an invalid shard path.")
62
+ if not (model / name).is_file():
63
+ raise ValueError(
64
+ f"Missing local weight shard: {name}; complete the existing download first."
65
+ )
66
+ return config, bits
67
+
68
+
69
+ def runtime_check(config, bits):
70
+ from importlib.metadata import version
71
+
72
+ from mlx_lm import server
73
+ from mlx_lm.utils import _get_classes
74
+
75
+ cls, args = _get_classes(config)
76
+ args.from_dict(config)
77
+ if cls.__module__ != "mlx_lm.models.qwen3_5_full":
78
+ raise ValueError(
79
+ "Swift architecture patch is missing. Use the serve launcher created by setup."
80
+ )
81
+ if digest(inspect.getsourcefile(cls)) != MODEL_HASHES[bits]:
82
+ raise ValueError(
83
+ "The installed architecture differs from the pinned Swift implementation."
84
+ )
85
+ if digest(server.__file__) != SERVER_HASH:
86
+ raise ValueError("The tested server cache patch is missing or has changed.")
87
+ versions = {
88
+ name: version(name)
89
+ for name in ("mlx", "mlx-metal", "mlx-lm", "transformers", "huggingface_hub")
90
+ }
91
+ expected = {
92
+ "mlx": "0.32.2",
93
+ "mlx-metal": "0.32.2",
94
+ "mlx-lm": "0.32.0",
95
+ "transformers": "5.14.1",
96
+ "huggingface_hub": "1.31.0",
97
+ }
98
+ if versions != expected:
99
+ raise ValueError(
100
+ f"Installed package versions differ from the tested runtime: {versions}"
101
+ )
102
+ print(
103
+ json.dumps(
104
+ {
105
+ "status": "PASS_RUNTIME_IDENTITY",
106
+ "python": sys.executable,
107
+ "model_class": cls.__module__ + "." + cls.__name__,
108
+ "server": server.__file__,
109
+ "versions": versions,
110
+ "weights_loaded": False,
111
+ }
112
+ ),
113
+ flush=True,
114
+ )
115
+
116
+
117
+ def setup(model, bits, destination):
118
+ if shutil.which("git") is None:
119
+ raise ValueError(
120
+ "Git is required. Install the macOS command-line tools, then retry."
121
+ )
122
+ if destination.exists():
123
+ raise ValueError(
124
+ f"Runtime directory already exists: {destination}. Use its serve launcher, or choose a new --runtime-dir."
125
+ )
126
+ destination.mkdir(parents=True)
127
+ source = destination / "mlx-lm"
128
+ envdir = destination / ".venv"
129
+ python = envdir / "bin/python"
130
+ patchdir = destination / "patches"
131
+ patchdir.mkdir()
132
+ patch_names = ["swift15-mlx-lm.patch"]
133
+ if bits == 5:
134
+ patch_names.append("enable-5bit.patch")
135
+ patch_names.append("swift15-server-cache.patch")
136
+ for name in patch_names:
137
+ url = f"https://huggingface.co/ukisai/Swift-1.5-{bits}bit-MLX/resolve/{REVISIONS[bits]}/compatibility/{name}"
138
+ with urllib.request.urlopen(url, timeout=60) as response:
139
+ data = response.read()
140
+ if hashlib.sha256(data).hexdigest() != PATCHES[name]:
141
+ raise ValueError(f"Patch checksum mismatch: {name}")
142
+ (patchdir / name).write_bytes(data)
143
+ subprocess.run(["git", "init", "--quiet", str(source)], check=True)
144
+ git = ["git", "-C", str(source)]
145
+ subprocess.run(
146
+ git + ["remote", "add", "origin", "https://github.com/ml-explore/mlx-lm.git"],
147
+ check=True,
148
+ )
149
+ subprocess.run(git + ["fetch", "--depth", "1", "origin", BASE], check=True)
150
+ subprocess.run(git + ["checkout", "--detach", BASE], check=True)
151
+ for name in patch_names:
152
+ subprocess.run(git + ["apply", "--check", str(patchdir / name)], check=True)
153
+ subprocess.run(git + ["apply", str(patchdir / name)], check=True)
154
+ venv.EnvBuilder(with_pip=True).create(envdir)
155
+ subprocess.run(
156
+ [
157
+ str(python),
158
+ "-I",
159
+ "-m",
160
+ "pip",
161
+ "install",
162
+ "mlx==0.32.2",
163
+ "mlx-metal==0.32.2",
164
+ "transformers==5.14.1",
165
+ "huggingface_hub==1.31.0",
166
+ "pillow==12.3.0",
167
+ "safetensors==0.8.0",
168
+ "-e",
169
+ str(source),
170
+ ],
171
+ check=True,
172
+ )
173
+ subprocess.run([str(python), "-I", "-m", "pip", "check"], check=True)
174
+ installed = destination / "swift_runtime.py"
175
+ shutil.copyfile(__file__, installed)
176
+ command = [str(python), "-I", str(installed), "start", "--model", str(model)]
177
+ subprocess.run(command + ["--check-only"], check=True)
178
+ launcher = destination / "serve"
179
+ launcher.write_text("#!/bin/sh\nexec " + shlex.join(command) + ' "$@"\n')
180
+ launcher.chmod(0o755)
181
+ print(
182
+ f"\nSetup complete. Start or restart with this exact command:\n{shlex.quote(str(launcher))}"
183
+ )
184
+
185
+
186
+ def main():
187
+ parser = argparse.ArgumentParser(description=__doc__)
188
+ sub = parser.add_subparsers(dest="action", required=True)
189
+ install = sub.add_parser(
190
+ "setup",
191
+ help="Install into a new isolated directory; reuse existing model files.",
192
+ )
193
+ install.add_argument("--model", required=True, type=Path)
194
+ install.add_argument("--runtime-dir", type=Path)
195
+ start = sub.add_parser(
196
+ "start", help="Verify the selected Python runtime and serve a local checkpoint."
197
+ )
198
+ start.add_argument("--model", required=True, type=Path)
199
+ start.add_argument("--check-only", action="store_true")
200
+ args, extra = parser.parse_known_args()
201
+ if args.action == "setup" and extra:
202
+ parser.error("Unknown setup arguments: " + " ".join(extra))
203
+ if sys.version_info[:2] != (3, 12):
204
+ parser.error(
205
+ "Use python3.12 for this tested runtime. Homebrew's Python 3.14 environment is separate."
206
+ )
207
+ if platform.system() != "Darwin" or platform.machine() != "arm64":
208
+ parser.error("This installer and launcher target native Apple Silicon macOS.")
209
+ model = args.model.expanduser().resolve()
210
+ try:
211
+ config, bits = local_config(model)
212
+ if args.action == "setup":
213
+ target = (
214
+ args.runtime_dir
215
+ or Path.home() / ".local/share/swift15-mlx" / f"{bits}bit"
216
+ )
217
+ setup(model, bits, target.expanduser().resolve())
218
+ return
219
+ os.environ["HF_HUB_OFFLINE"] = "1"
220
+ os.environ["TRANSFORMERS_OFFLINE"] = "1"
221
+ runtime_check(config, bits)
222
+ if args.check_only:
223
+ return
224
+ os.execv(
225
+ sys.executable,
226
+ [
227
+ sys.executable,
228
+ "-I",
229
+ "-m",
230
+ "mlx_lm.server",
231
+ "--model",
232
+ str(model),
233
+ "--host",
234
+ "127.0.0.1",
235
+ "--port",
236
+ "8080",
237
+ *extra,
238
+ ],
239
+ )
240
+ except (
241
+ OSError,
242
+ ValueError,
243
+ KeyError,
244
+ ImportError,
245
+ subprocess.CalledProcessError,
246
+ ) as exc:
247
+ raise SystemExit(f"Swift runtime setup/start failed: {exc}") from None
248
+
249
+
250
+ if __name__ == "__main__":
251
+ main()