File size: 5,337 Bytes
730aab9
 
8227673
 
 
 
 
 
730aab9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
# Server cache update

**Starting from Homebrew or seeing `Received 501 parameters not in model`?**
Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an
explicit `serve` launcher. It reuses your existing model directory, including an
HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew
installation that lacks the Swift architecture and cache patches.

This update changes the pinned MLX-LM server, not the checkpoint. It is a Swift
runtime patch based on official MLX-LM `c69d1288440a0dc4e6401fc417098b07598dccd5`, not an upstream release.
The existing architecture patch remains required.

## Apply to an existing installation

Stop the server. Activate the same Python environment used for serving. From the
directory containing `swift15-mlx-lm` and `Swift-1.5-4bit-MLX`, after downloading this patch:

```bash
git -C swift15-mlx-lm rev-parse HEAD
git -C swift15-mlx-lm apply --check ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch
git -C swift15-mlx-lm apply ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch
python -m pip install --no-deps -e ./swift15-mlx-lm
python Swift-1.5-4bit-MLX/compatibility/cache-tests/verify_server_patch.py
```

The revision must equal the pinned revision above. Stop if the patch check fails;
do not force it onto different server code or apply it twice. For a fresh install,
follow [USAGE.md](USAGE.md), which applies the architecture patches first.
The server verifier checks the active Python installation without loading weights.

Once dependencies and the complete checkpoint are already local, startup works
without Hub access:

```bash
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python -m mlx_lm.server \
  --model ./Swift-1.5-4bit-MLX --host 127.0.0.1 --port 8080
```

Remove an earlier `--prompt-cache-size 0` override to enable reuse. Explicit
arguments override the new defaults. Installing/downloading prerequisites still
requires internet; offline operation requires those prerequisites already present.

## Behavior and limits

Defaults: two retained cache entries, prompt concurrency 1, decode concurrency 1,
and prefill steps of 512 tokens. The byte cap is now passed to the LRU cache, so
every insertion enforces it. Previously the CLI byte limit only triggered a trim
when adding a batched request; later retained entries could exceed that setting.

The automatic Metal budget uses the reported recommended working-set size `R`,
the loaded model parameter bytes `W`, and workspace reserve `max(4 GiB, R/8)`.
It budgets half the remaining space for retained entries:
`max(0, (R - W - reserve) / 2)`. This is a conservative estimate, not a guarantee
that an arbitrary active context fits. It does not change macOS or MLX limits.
When device metadata is unavailable, the entry limit still applies; supply an
explicit byte limit if needed. `--prompt-cache-bytes 2GB`, for example, sets a
retention budget of 2,000,000,000 bytes; it is not a universal recommended value.

The active request, temporary tensors, allocator and other apps use additional
memory. Cache byte accounting is logical and may include shared buffers. Entries
that cannot fit are evicted whole; the client's input history is not truncated.
Reuse can therefore vary with context size. This patch does not add KV
quantization, change attention or alter any checkpoint tensors.

## Reproduce the checks

Use the patched environment with Python 3.12, MLX 0.32.2 and MLX-LM 0.32.0.
`validation.json` and the two `offline-*-results.json` files record our local
results. No hosted CI status is claimed. The tests use random tiny models and the
local tokenizer; they never open the released Swift weight shards.

```bash
python -m pip install 'pytest==9.1.1' 'requests==2.32.5'
python -m pytest -q swift15-mlx-lm/tests/test_server_cache_budget.py swift15-mlx-lm/tests/test_server.py
python Swift-1.5-4bit-MLX/compatibility/cache-tests/test_server_cache_http.py \
  --assets ./Swift-1.5-4bit-MLX \
  --config Swift-1.5-4bit-MLX/compatibility/cache-tests/synthetic-config.json \
  --bits 4 --output cache-http-results.json
```

The upstream server suite may fetch its small upstream test model if not already
cached. The separate HTTP harness sets offline mode and needs only local assets.
It covers 12 requests each with automatic budgeting, an explicit byte limit, and
disabled retention; streaming and seeded sequential execution are included.
Synthetic output is deliberately fixed for transport/cache checks, not quality.
Full-model follow-up validation is now recorded in
[FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md): both complete checkpoints passed
an 86k-token synthetic text conversation and two cached follow-ups on a 48 GiB
M4 Pro using the patched server. GUI integration and other memory/context sizes
remain outside that test.

## Roll back this runtime patch

Stop the server, then reverse only this patch and restart with explicit settings:

```bash
git -C swift15-mlx-lm apply --reverse --check ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch
git -C swift15-mlx-lm apply --reverse ../Swift-1.5-4bit-MLX/compatibility/swift15-server-cache.patch
```

This leaves the architecture patch and model files intact. An editable install
uses the source change after restart; reinstall if a non-editable copy was used.