# Load Swift 1.5 with its complete MLX architecture **Starting from Homebrew or seeing `Received 501 parameters not in model`?** Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an explicit `serve` launcher. It reuses your existing model directory, including an HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew installation that lacks the Swift architecture and cache patches. Use the included patch with the pinned official Apple MLX-LM revision. Unpatched text-only Qwen support does not preserve this checkpoint's complete parameter tree. Use Python 3.12 in a new working directory. The model repository is public; authentication is optional. Download the complete repository, including all three root-level `model-0000*-of-00003.safetensors` shards (15,826,764,635 bytes total), the index, config and tokenizer assets. The archived 3.87 MB diagnostic sample under `compatibility/mac-check/` cannot be used as the model. The included MLX-LM patch is required. This is the supported installation path; GUI apps and stock runtimes that cannot use that patch have not been validated. The commands resolve the current public revision once, then download and verify that exact snapshot. They do not load the model into memory. ```bash python3.12 -m venv .venv source .venv/bin/activate python -m pip install 'huggingface_hub==1.31.0' SWIFT_MLX_REVISION="$(python -c 'import json, urllib.request; print(json.load(urllib.request.urlopen("https://huggingface.co/api/models/ukisai/Swift-1.5-4bit-MLX"))["sha"])')" hf download ukisai/Swift-1.5-4bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-4bit-MLX hf cache verify ukisai/Swift-1.5-4bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-4bit-MLX --fail-on-missing-files python Swift-1.5-4bit-MLX/check_download.py Swift-1.5-4bit-MLX git clone https://github.com/ml-explore/mlx-lm.git swift15-mlx-lm git -C swift15-mlx-lm checkout --detach c69d1288440a0dc4e6401fc417098b07598dccd5 git -C swift15-mlx-lm apply --check ../Swift-1.5-4bit-MLX/compatibility/swift15-mlx-lm.patch git -C swift15-mlx-lm apply ../Swift-1.5-4bit-MLX/compatibility/swift15-mlx-lm.patch git -C swift15-mlx-lm apply --check ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch git -C swift15-mlx-lm apply ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch ``` Verify that the download contains all three root weight shards. If a prior app download gave you only `real-checkpoint-samples.safetensors`, run the repository download above in a fresh directory. That file is a test fixture, not a model shard. Stop after any missing-file or checksum failure. The checkpoint contains about 15.83 GB of tensor data, before runtime, cache and OS overhead. Do not force the full model onto a 16 GiB Mac or increase system memory limits. Full-model generation on 24 GB Macs remains unverified. Full-model follow-up validation is now recorded in [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md): both complete checkpoints passed an 86k-token synthetic text conversation and two cached follow-ups on a 48 GiB M4 Pro using the patched server. GUI integration and other memory/context sizes remain outside that test. On Apple Silicon: ```bash pip install 'mlx==0.32.2' 'transformers==5.14.1' 'huggingface_hub==1.31.0' pillow pip install -e ./swift15-mlx-lm ``` On Linux CPU (Python 3.12 and glibc 2.35 or newer): ```bash pip install 'mlx[cpu]==0.32.2' 'transformers==5.14.1' 'huggingface_hub==1.31.0' pillow pip install -e ./swift15-mlx-lm ``` Before generating, check the Python environment that will actually run the model: ```bash python Swift-1.5-4bit-MLX/check_download.py Swift-1.5-4bit-MLX --runtime ``` This checks the complete download and selects the pinned patched architecture without loading weights. It does not check a separate GUI app's environment, available inference memory, or generated quality. See [TROUBLESHOOTING.md](TROUBLESHOOTING.md) if it fails or a server reports `generation thread died`. Text generation: ```python import mlx.core as mx from mlx_lm import load, generate model, tokenizer = load("Swift-1.5-4bit-MLX") if mx.default_device() == mx.cpu: model.apply( lambda value: value.astype(mx.float32) if mx.issubdtype(value.dtype, mx.floating) else value ) prompt = tokenizer.apply_chat_template( [{"role": "user", "content": "Say hello."}], tokenize=False, add_generation_prompt=True, enable_thinking=False, ) print(generate(model, tokenizer, prompt=prompt, max_tokens=32)) ``` The Linux CPU branch promotes only in-memory floating parameters to FP32. Packed 4-bit weights and all files remain unchanged. This avoids the official MLX 0.32.2 Linux scalar BF16 quantized-matmul accumulation bug reproduced in `compatibility/cpu-quantized-matmul-diagnostic.json` (8,192 exact ones summed to 256 in BF16, versus the correct 8,192 in FP32). The release's CPU generation, MTP and vision smoke tests use this FP32 runtime. Apple Silicon inference does not use this CPU workaround; the separate full-model Metal test is recorded in [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md). The original chat template also accepts `reasoning_effort="low"`, `"medium"`, and `"xhigh"`; this release validates the original low and xhigh formats. This structural smoke test does not establish long-context or benchmark accuracy. The patch implements an explicit MTP step (`model.mtp_logits`) and the vision encoder (`model.visual`). Their weights are retained and the release validation records their component execution. Speculative generation and integrated image/video chat are not implemented. Unsupported multimodal generation calls raise an error. ## Server cache update The supplied server patch keeps prompt reuse enabled with at most two retained entries, enforces a retained-cache byte budget at insertion, and defaults to one prompt and one decode stream with 512-token prefill steps. On Metal the default budget estimates headroom after model weights and workspace; an explicit `--prompt-cache-bytes` overrides it. `--prompt-cache-size 2` means two entries, not two GB. See [SERVER_CACHE_UPDATE.md](SERVER_CACHE_UPDATE.md) for installation, offline startup, validation and rollback. Existing installations must apply this runtime update and restart. No model-weight download or re-quantization is needed. The retained-cache limit does not cap active-request or total process memory. An entry larger than the budget is evicted, so reuse depends on what fits. Use `--prompt-cache-size 0` only when intentionally disabling reuse. The model's configured context length is not a promise that it fits every Mac. Validation: 44 unit/upstream server tests and 72 offline HTTP requests passed across tiny 4-bit and 5-bit fixtures. These checks cover retention, warm reuse, streaming and sequential generation. They do not establish full-model long-context capacity or GUI integration. Reproduce conversion only from the complete original Swift BF16 export identified in `QUANTIZATION_MANIFEST.json`, after verifying its 18 shards and original assets: ```bash mlx_lm.convert --hf-path /path/to/Swift-1.5-BF16 \ --mlx-path Swift-1.5-4bit-MLX \ --quantize --q-mode affine --q-bits 4 --q-group-size 64 ``` The converter refuses an existing output directory. It uses official MLX-LM lazy loading, quantization, sharding, and saving; the patch supplies the complete model and strict parameter/asset mapping. ## Reproduce the small Metal checks The diagnostic files are stored in `compatibility/mac-check/fixtures.zip` so model downloaders do not mistake a small test checkpoint for the real model. The updated validator reads that archive automatically, without downloading or evaluating the full 27B model: ```bash python Swift-1.5-4bit-MLX/compatibility/validate-mac-format.py ``` This checks format, the complete parameter tree and three real weight samples. It does not establish full-model Apple Silicon generation or GUI compatibility. ## Full-model Mac follow-up Full-model follow-up validation is now recorded in [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md): both complete checkpoints passed an 86k-token synthetic text conversation and two cached follow-ups on a 48 GiB M4 Pro using the patched server. GUI integration and other memory/context sizes remain outside that test.