Swift-1.5-4bit-MLX / QUICKSTART.md
ukisai's picture
Add a verified Swift runtime installer and explicit server launcher (#2)
8227673
|
Raw History Blame Contribute Delete
4.67 kB

Start Swift with the required MLX runtime

Received 501 parameters not in model: language_model.visual... means the selected loader does not represent this checkpoint's vision parameter tree. For example, Homebrew's MLX-LM 0.31.3 runs from its own Python environment and does not acquire the Swift patches when you download this model. This error happens before generation; changing prompt-cache settings does not fix it.

The helper below installs the same pinned architecture and cache patches used in FULL_MAC_VALIDATION.md. It creates a separate Python environment and an explicit launcher. Your cached weights can stay where they are.

1. Use Python 3.12 on Apple Silicon

python3.12 --version

If that command is missing and you use Homebrew, install it with brew install python@3.12. Leave your existing Homebrew MLX-LM installation alone. Git is also required. The helper checks the Python version and native Apple Silicon platform before installing anything.

2. Install the runtime using your existing model directory

Stop your old server. Download only the small setup helper:

hf download ukisai/Swift-1.5-4bit-MLX swift_runtime.py --local-dir Swift-MLX-setup

Set SWIFT_MODEL to the complete model directory printed by your earlier hf download command. It can be an HF cache snapshot or a local download folder. This example uses the previously published snapshot in the default HF cache:

SWIFT_MODEL="$HOME/.cache/huggingface/hub/models--ukisai--Swift-1.5-4bit-MLX/snapshots/730aab9b0395b26f7d9cf4b0dfa6a4c788aff6fd"
python3.12 Swift-MLX-setup/swift_runtime.py setup --model "$SWIFT_MODEL"

If you downloaded with --local-dir Swift-1.5-4bit-MLX, set SWIFT_MODEL="$PWD/Swift-1.5-4bit-MLX" instead. Use your actual path if you customized the HF cache location or downloaded a different complete revision.

The helper verifies fixed SHA-256 hashes for the patches, fetches official MLX-LM revision c69d1288440a0dc4e6401fc417098b07598dccd5, applies the required architecture and cache patches, and installs pinned packages into ~/.local/share/swift15-mlx/4bit/.venv. For 5-bit it also applies the existing 5-bit support patch. It checks the selected architecture, server source and package versions before creating the launcher.

Installation needs internet for source code and Python dependencies. It never downloads, converts, edits or re-quantizes the model weights. Setup checks local assets and indexed shard presence; use check_download.py --hash from USAGE.md if you also need to verify every weight byte.

3. Start with this command every time

"$HOME/.local/share/swift15-mlx/4bit/serve"

This starts the server at http://127.0.0.1:8080 using the dedicated Python environment and your existing local model path. It enables offline Hub and Transformers operation. You do not need to activate a virtual environment. Keep using this launcher after closing and reopening Terminal; typing the bare mlx_lm.server command can select Homebrew again.

To check the runtime without loading weights:

"$HOME/.local/share/swift15-mlx/4bit/serve" --check-only

Additional server arguments work, for example serve --port 8081 using the same full launcher path. The tested cache defaults remain enabled. A previous --prompt-cache-size 0 override disables reuse if you add it again.

If the runtime directory already exists, use its launcher. The installer refuses to overwrite an existing directory; choose a new path with --runtime-dir if you need a separate installation, then use the launcher path it prints. Stop on an installation failure. A successful --check-only verifies the runtime and local file presence, not full-model inference or available memory.

Validation

Both new 4-bit and 5-bit installations passed native Apple Silicon checks using existing tiny synthetic checkpoints containing text, vision and MTP entries. The actual launchers served two HTTP generation requests each while a conflicting mlx_lm.server and Python module were placed on the search path. Paths containing spaces worked. An unpatched loader was rejected before loading weights, and a repeat installation did not overwrite an existing runtime.

Recorded results cover this installer and launch path. The earlier complete-model 48 GiB Mac tests remain in FULL_MAC_VALIDATION.md. No full-model rerun or hosted CI badge is claimed for this packaging update. Memory capacity, GUI integration, integrated image/video chat and speculative MTP limitations are unchanged.