# Check this download and its runtime **Starting from Homebrew or seeing `Received 501 parameters not in model`?** Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an explicit `serve` launcher. It reuses your existing model directory, including an HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew installation that lacks the Swift architecture and cache patches. This release needs **all 3 shards (15.83 GB (14.74 GiB))**, its index/config/tokenizer assets, and the supplied architecture patch. Downloading an HF repository does not install the MLX-LM patches into Python or into a GUI app's separate runtime. ## 1. Verify the files without loading the model Run from the directory containing `Swift-1.5-4bit-MLX`: ```bash python Swift-1.5-4bit-MLX/check_download.py Swift-1.5-4bit-MLX --hash ``` This uses standard Python, reads the headers and stream-checks the original shard SHA-256 hashes with bounded memory. It does not load the model or contact a server. A successful result verifies the listed download checks, not inference. If files are missing or incomplete, resume the full download using the pinned snapshot instructions in [USAGE.md](USAGE.md), then verify again. Keep all shards in one directory alongside the index, config and tokenizer. Do not load a single shard or a diagnostic sample as a standalone model. Do not combine 4-bit and 5-bit files. The complete 4-bit payload is 15.83 GB; the complete 5-bit payload is 19.28 GB. A reported 5.7 GB download is insufficient for either complete checkpoint, although that number alone does not tell us what an app has downloaded or is displaying. ## 2. Verify the runtime Follow [USAGE.md](USAGE.md) to install the supplied architecture patch in the pinned Python environment. Then, in that same environment: ```bash python Swift-1.5-4bit-MLX/check_download.py Swift-1.5-4bit-MLX --runtime ``` The checker confirms that MLX-LM dispatches to `qwen3_5_full` and that its architecture source matches the audited supplied patch. It records the MLX and MLX-LM versions. It deliberately rejects other architecture code until checked separately. For 5-bit, the base architecture patch alone is insufficient: `enable-5bit.patch` must also be applied. A passing check in a terminal does not change an app's separate runtime. GUI integration has not been validated for these full checkpoints. Do not remove vision/MTP parameters or change the config to force a stock text loader to accept the file. All 2,379 saved tensor entries belong to this release. ## Server cache update The supplied server patch keeps prompt reuse enabled with at most two retained entries, enforces a retained-cache byte budget at insertion, and defaults to one prompt and one decode stream with 512-token prefill steps. On Metal the default budget estimates headroom after model weights and workspace; an explicit `--prompt-cache-bytes` overrides it. `--prompt-cache-size 2` means two entries, not two GB. See [SERVER_CACHE_UPDATE.md](SERVER_CACHE_UPDATE.md) for installation, offline startup, validation and rollback. Existing installations must apply this runtime update and restart. No model-weight download or re-quantization is needed. The retained-cache limit does not cap active-request or total process memory. An entry larger than the budget is evicted, so reuse depends on what fits. Use `--prompt-cache-size 0` only when intentionally disabling reuse. The model's configured context length is not a promise that it fits every Mac. Validation: 44 unit/upstream server tests and 72 offline HTTP requests passed across tiny 4-bit and 5-bit fixtures. These checks cover retention, warm reuse, streaming and sequential generation. They do not establish full-model long-context capacity or GUI integration. ## 3. If a server reports `404 generation thread died` In the pinned official MLX-LM server, this message means its generation worker failed. The HTTP 404 is not enough to diagnose a missing web page, a broken model, or a plugin problem. The original exception is in the server's terminal/log, before the later generic response. After the file and runtime checks pass, run the short direct Python generation example in [USAGE.md](USAGE.md) in the same environment, independently of any plugin. Use a Mac with sufficient RAM for the complete weights, runtime, cache and macOS. Full-model generation on a 24 GB Mac has not been verified. Neither full release has been validated on the available 16 GiB test Mac; do not raise memory limits. If it fails, retain the original exception and traceback. A useful report contains the app/command and version, Mac RAM, model revision, checker output, and that first exception. Do not include access tokens or private prompts. Repeating a request against a dead worker will not recover it; restart only after addressing the cause. ## What was checked on 2026-09-25 - Both public releases contain all indexed shards and 2,379 matching tensor headers. - The unpatched official MLX-LM 0.32.0 model at the revision pinned in USAGE rejects 501 saved vision entries and removes 31 saved MTP entries during sanitization. The required patched parameter trees accept all 2,379 entries with no shape mismatch. These comparisons use real headers and unevaluated zero fixtures, not full payloads. - An incomplete checkpoint reproduced the exact HTTP 404 / `generation thread died` response. That demonstrates one possible cause, not the cause of any specific report. - Both patched 4-bit and 5-bit runtimes passed small synthetic HTTP tests for non-thinking chat, default thinking format, seeded generation, and streaming with a tool definition. These are architecture/server checks, not full Swift generation or a test of the reported third-party plugin. These checks do not establish full-model Mac execution or GUI compatibility. Diagnosing an individual crash requires its original exception and runtime configuration. ## Full-model Mac follow-up Full-model follow-up validation is now recorded in [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md): both complete checkpoints passed an 86k-token synthetic text conversation and two cached follow-ups on a 48 GiB M4 Pro using the patched server. GUI integration and other memory/context sizes remain outside that test.