Swift-1.5-4bit-MLX / QUICKSTART.md
ukisai's picture
Add a verified Swift runtime installer and explicit server launcher (#2)
8227673
|
Raw History Blame Contribute Delete
4.67 kB
# Start Swift with the required MLX runtime
`Received 501 parameters not in model: language_model.visual...` means the
selected loader does not represent this checkpoint's vision parameter tree.
For example, Homebrew's MLX-LM 0.31.3 runs from its own Python environment and
does not acquire the Swift patches when you download this model. This error
happens before generation; changing prompt-cache settings does not fix it.
The helper below installs the same pinned architecture and cache patches used
in [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md). It creates a separate Python
environment and an explicit launcher. Your cached weights can stay where they are.
## 1. Use Python 3.12 on Apple Silicon
```bash
python3.12 --version
```
If that command is missing and you use Homebrew, install it with
`brew install python@3.12`. Leave your existing Homebrew MLX-LM installation alone.
Git is also required. The helper checks the Python version and native Apple
Silicon platform before installing anything.
## 2. Install the runtime using your existing model directory
Stop your old server. Download only the small setup helper:
```bash
hf download ukisai/Swift-1.5-4bit-MLX swift_runtime.py --local-dir Swift-MLX-setup
```
Set `SWIFT_MODEL` to the complete model directory printed by your earlier
`hf download` command. It can be an HF cache snapshot or a local download folder.
This example uses the previously published snapshot in the default HF cache:
```bash
SWIFT_MODEL="$HOME/.cache/huggingface/hub/models--ukisai--Swift-1.5-4bit-MLX/snapshots/730aab9b0395b26f7d9cf4b0dfa6a4c788aff6fd"
python3.12 Swift-MLX-setup/swift_runtime.py setup --model "$SWIFT_MODEL"
```
If you downloaded with `--local-dir Swift-1.5-4bit-MLX`, set
`SWIFT_MODEL="$PWD/Swift-1.5-4bit-MLX"` instead. Use your actual path if you customized
the HF cache location or downloaded a different complete revision.
The helper verifies fixed SHA-256 hashes for the patches, fetches official
MLX-LM revision `c69d1288440a0dc4e6401fc417098b07598dccd5`, applies the required
architecture and cache patches, and installs pinned packages into
`~/.local/share/swift15-mlx/4bit/.venv`. For 5-bit it also applies the existing
5-bit support patch. It checks the selected architecture, server source and
package versions before creating the launcher.
Installation needs internet for source code and Python dependencies. It never
downloads, converts, edits or re-quantizes the model weights. Setup checks local
assets and indexed shard presence; use `check_download.py --hash` from
[USAGE.md](USAGE.md) if you also need to verify every weight byte.
## 3. Start with this command every time
```bash
"$HOME/.local/share/swift15-mlx/4bit/serve"
```
This starts the server at `http://127.0.0.1:8080` using the dedicated Python
environment and your existing local model path. It enables offline Hub and
Transformers operation. You do not need to activate a virtual environment.
Keep using this launcher after closing and reopening Terminal; typing the bare
`mlx_lm.server` command can select Homebrew again.
To check the runtime without loading weights:
```bash
"$HOME/.local/share/swift15-mlx/4bit/serve" --check-only
```
Additional server arguments work, for example `serve --port 8081` using the same
full launcher path. The tested cache defaults remain enabled. A previous
`--prompt-cache-size 0` override disables reuse if you add it again.
If the runtime directory already exists, use its launcher. The installer refuses
to overwrite an existing directory; choose a new path with `--runtime-dir` if
you need a separate installation, then use the launcher path it prints.
Stop on an installation failure. A successful `--check-only` verifies the runtime
and local file presence, not full-model inference or available memory.
## Validation
Both new 4-bit and 5-bit installations passed native Apple Silicon checks using
existing tiny synthetic checkpoints containing text, vision and MTP entries.
The actual launchers served two HTTP generation requests each while a conflicting
`mlx_lm.server` and Python module were placed on the search path. Paths containing
spaces worked. An unpatched loader was rejected before loading weights, and a
repeat installation did not overwrite an existing runtime.
[Recorded results](compatibility/launcher-validation.json) cover this installer
and launch path. The earlier complete-model 48 GiB Mac tests remain in
[FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md). No full-model rerun or hosted CI
badge is claimed for this packaging update. Memory capacity, GUI integration,
integrated image/video chat and speculative MTP limitations are unchanged.