Swift-1.5-4bit-MLX / SERVER_CACHE_UPDATE.md
ukisai's picture
Add a verified Swift runtime installer and explicit server launcher (#2)
8227673
|
Raw History Blame Contribute Delete
5.34 kB

Server cache update

Starting from Homebrew or seeing Received 501 parameters not in model? Use QUICKSTART.md to install the required runtime and create an explicit serve launcher. It reuses your existing model directory, including an HF cache snapshot. A bare mlx_lm.server command may select a separate Homebrew installation that lacks the Swift architecture and cache patches.

This update changes the pinned MLX-LM server, not the checkpoint. It is a Swift runtime patch based on official MLX-LM c69d1288440a0dc4e6401fc417098b07598dccd5, not an upstream release. The existing architecture patch remains required.

Apply to an existing installation

Stop the server. Activate the same Python environment used for serving. From the directory containing swift15-mlx-lm and Swift-1.5-4bit-MLX, after downloading this patch:

git -C swift15-mlx-lm rev-parse HEAD
git -C swift15-mlx-lm apply --check ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch
git -C swift15-mlx-lm apply ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch
python -m pip install --no-deps -e ./swift15-mlx-lm
python Swift-1.5-4bit-MLX/compatibility/cache-tests/verify_server_patch.py

The revision must equal the pinned revision above. Stop if the patch check fails; do not force it onto different server code or apply it twice. For a fresh install, follow USAGE.md, which applies the architecture patches first. The server verifier checks the active Python installation without loading weights.

Once dependencies and the complete checkpoint are already local, startup works without Hub access:

HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python -m mlx_lm.server \
  --model ./Swift-1.5-4bit-MLX --host 127.0.0.1 --port 8080

Remove an earlier --prompt-cache-size 0 override to enable reuse. Explicit arguments override the new defaults. Installing/downloading prerequisites still requires internet; offline operation requires those prerequisites already present.

Behavior and limits

Defaults: two retained cache entries, prompt concurrency 1, decode concurrency 1, and prefill steps of 512 tokens. The byte cap is now passed to the LRU cache, so every insertion enforces it. Previously the CLI byte limit only triggered a trim when adding a batched request; later retained entries could exceed that setting.

The automatic Metal budget uses the reported recommended working-set size R, the loaded model parameter bytes W, and workspace reserve max(4 GiB, R/8). It budgets half the remaining space for retained entries: max(0, (R - W - reserve) / 2). This is a conservative estimate, not a guarantee that an arbitrary active context fits. It does not change macOS or MLX limits. When device metadata is unavailable, the entry limit still applies; supply an explicit byte limit if needed. --prompt-cache-bytes 2GB, for example, sets a retention budget of 2,000,000,000 bytes; it is not a universal recommended value.

The active request, temporary tensors, allocator and other apps use additional memory. Cache byte accounting is logical and may include shared buffers. Entries that cannot fit are evicted whole; the client's input history is not truncated. Reuse can therefore vary with context size. This patch does not add KV quantization, change attention or alter any checkpoint tensors.

Reproduce the checks

Use the patched environment with Python 3.12, MLX 0.32.2 and MLX-LM 0.32.0. validation.json and the two offline-*-results.json files record our local results. No hosted CI status is claimed. The tests use random tiny models and the local tokenizer; they never open the released Swift weight shards.

python -m pip install 'pytest==9.1.1' 'requests==2.32.5'
python -m pytest -q swift15-mlx-lm/tests/test_server_cache_budget.py swift15-mlx-lm/tests/test_server.py
python Swift-1.5-4bit-MLX/compatibility/cache-tests/test_server_cache_http.py \
  --assets ./Swift-1.5-4bit-MLX \
  --config Swift-1.5-4bit-MLX/compatibility/cache-tests/synthetic-config.json \
  --bits 4 --output cache-http-results.json

The upstream server suite may fetch its small upstream test model if not already cached. The separate HTTP harness sets offline mode and needs only local assets. It covers 12 requests each with automatic budgeting, an explicit byte limit, and disabled retention; streaming and seeded sequential execution are included. Synthetic output is deliberately fixed for transport/cache checks, not quality. Full-model follow-up validation is now recorded in FULL_MAC_VALIDATION.md: both complete checkpoints passed an 86k-token synthetic text conversation and two cached follow-ups on a 48 GiB M4 Pro using the patched server. GUI integration and other memory/context sizes remain outside that test.

Roll back this runtime patch

Stop the server, then reverse only this patch and restart with explicit settings:

git -C swift15-mlx-lm apply --reverse --check ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch
git -C swift15-mlx-lm apply --reverse ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch

This leaves the architecture patch and model files intact. An editable install uses the source change after restart; reinstall if a non-editable copy was used.