ukisai commited on
Commit
5359607
·
verified ·
1 Parent(s): ff55db2

Bound retained server cache while keeping prompt reuse enabled

Browse files

The pinned server constructed its retained-cache LRU without the configured byte limit. Entries saved during prefill or completion could therefore exceed `--prompt-cache-bytes` after the earlier request-admission trim. This patch enforces the retention limit on each insertion and keeps prompt reuse enabled with bounded defaults.

The Swift server profile defaults to two retained entries, one prompt and one decode stream, and 512-token prefill steps. A Metal memory estimate sets the default retained-cache budget; an explicit byte limit overrides it. The budget excludes active-request and total process memory, so this is not a promise that every context length fits.

Validation completed on clean copies of the pinned source for both architecture patch stacks:

- PASS: 44 unit/upstream server tests per patch stack.
- PASS: 36 offline HTTP requests per format, including warm cache reuse, explicit byte-cap enforcement, streaming, sequential generation and disabled retention.
- PASS: clean patch application, runtime source identity, unchanged model implementation, reverse-apply check, Black 25.1.0, isort 6.0.0 and Ruff 0.16.6.

These are reproducible local Apple Silicon checks, not hosted CI statuses. Results and the offline regression harness are included under `compatibility/cache-tests/`. They use tiny random model fixtures; full Swift long-context capacity remains unverified by this change.

Only runtime patch, tests, instructions and file manifest change. Checkpoint weights, quantization, architecture code and tokenizer assets are unchanged. Existing installations must apply the patch and restart; fetching model files alone does not update an installed server. `SERVER_CACHE_UPDATE.md` includes installation, offline startup, validation and rollback instructions.

This PR is for review and has not been merged into main.

SERVER_CACHE_UPDATE.md ADDED
@@ -0,0 +1,93 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Server cache update
2
+
3
+ This update changes the pinned MLX-LM server, not the checkpoint. It is a Swift
4
+ runtime patch based on official MLX-LM `c69d1288440a0dc4e6401fc417098b07598dccd5`, not an upstream release.
5
+ The existing architecture patch remains required.
6
+
7
+ ## Apply to an existing installation
8
+
9
+ Stop the server. Activate the same Python environment used for serving. From the
10
+ directory containing `swift5-mlx-lm` and `Swift-1.5-5bit-MLX`, after downloading this patch:
11
+
12
+ ```bash
13
+ git -C swift5-mlx-lm rev-parse HEAD
14
+ git -C swift5-mlx-lm apply --check ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
15
+ git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
16
+ python -m pip install --no-deps -e ./swift5-mlx-lm
17
+ python Swift-1.5-5bit-MLX/compatibility/cache-tests/verify_server_patch.py
18
+ ```
19
+
20
+ The revision must equal the pinned revision above. Stop if the patch check fails;
21
+ do not force it onto different server code or apply it twice. For a fresh install,
22
+ follow [USAGE.md](USAGE.md), which applies the architecture patches first.
23
+ The server verifier checks the active Python installation without loading weights.
24
+
25
+ Once dependencies and the complete checkpoint are already local, startup works
26
+ without Hub access:
27
+
28
+ ```bash
29
+ HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python -m mlx_lm.server \
30
+ --model ./Swift-1.5-5bit-MLX --host 127.0.0.1 --port 8080
31
+ ```
32
+
33
+ Remove an earlier `--prompt-cache-size 0` override to enable reuse. Explicit
34
+ arguments override the new defaults. Installing/downloading prerequisites still
35
+ requires internet; offline operation requires those prerequisites already present.
36
+
37
+ ## Behavior and limits
38
+
39
+ Defaults: two retained cache entries, prompt concurrency 1, decode concurrency 1,
40
+ and prefill steps of 512 tokens. The byte cap is now passed to the LRU cache, so
41
+ every insertion enforces it. Previously the CLI byte limit only triggered a trim
42
+ when adding a batched request; later retained entries could exceed that setting.
43
+
44
+ The automatic Metal budget uses the reported recommended working-set size `R`,
45
+ the loaded model parameter bytes `W`, and workspace reserve `max(4 GiB, R/8)`.
46
+ It budgets half the remaining space for retained entries:
47
+ `max(0, (R - W - reserve) / 2)`. This is a conservative estimate, not a guarantee
48
+ that an arbitrary active context fits. It does not change macOS or MLX limits.
49
+ When device metadata is unavailable, the entry limit still applies; supply an
50
+ explicit byte limit if needed. `--prompt-cache-bytes 2GB`, for example, sets a
51
+ retention budget of 2,000,000,000 bytes; it is not a universal recommended value.
52
+
53
+ The active request, temporary tensors, allocator and other apps use additional
54
+ memory. Cache byte accounting is logical and may include shared buffers. Entries
55
+ that cannot fit are evicted whole; the client's input history is not truncated.
56
+ Reuse can therefore vary with context size. This patch does not add KV
57
+ quantization, change attention or alter any checkpoint tensors.
58
+
59
+ ## Reproduce the checks
60
+
61
+ Use the patched environment with Python 3.12, MLX 0.32.2 and MLX-LM 0.32.0.
62
+ `validation.json` and the two `offline-*-results.json` files record our local
63
+ results. No hosted CI status is claimed. The tests use random tiny models and the
64
+ local tokenizer; they never open the released Swift weight shards.
65
+
66
+ ```bash
67
+ python -m pip install 'pytest==9.1.1' 'requests==2.32.5'
68
+ python -m pytest -q swift5-mlx-lm/tests/test_server_cache_budget.py swift5-mlx-lm/tests/test_server.py
69
+ python Swift-1.5-5bit-MLX/compatibility/cache-tests/test_server_cache_http.py \
70
+ --assets ./Swift-1.5-5bit-MLX \
71
+ --config Swift-1.5-5bit-MLX/compatibility/cache-tests/synthetic-config.json \
72
+ --bits 5 --output cache-http-results.json
73
+ ```
74
+
75
+ The upstream server suite may fetch its small upstream test model if not already
76
+ cached. The separate HTTP harness sets offline mode and needs only local assets.
77
+ It covers 12 requests each with automatic budgeting, an explicit byte limit, and
78
+ disabled retention; streaming and seeded sequential execution are included.
79
+ Synthetic output is deliberately fixed for transport/cache checks, not quality.
80
+ Full Swift long-context execution on a larger Mac has not been independently
81
+ retested by this change. GUI runtimes need their own integration.
82
+
83
+ ## Roll back this runtime patch
84
+
85
+ Stop the server, then reverse only this patch and restart with explicit settings:
86
+
87
+ ```bash
88
+ git -C swift5-mlx-lm apply --reverse --check ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
89
+ git -C swift5-mlx-lm apply --reverse ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
90
+ ```
91
+
92
+ This leaves the architecture patch and model files intact. An editable install
93
+ uses the source change after restart; reinstall if a non-editable copy was used.
TROUBLESHOOTING.md CHANGED
@@ -43,35 +43,28 @@ separate runtime. GUI integration has not been validated for these full checkpoi
43
  Do not remove vision/MTP parameters or change the config to force a stock text
44
  loader to accept the file. All 2,379 saved tensor entries belong to this release.
45
 
46
- ## Server settings for limited memory
47
-
48
- The pinned MLX-LM server retains prompt caches for reuse by default. For a setup
49
- that favors lower memory use over reuse speed, stop the existing server and start
50
- it from the same patched Python environment with:
51
-
52
- ```bash
53
- mlx_lm.server --model ukisai/Swift-1.5-5bit-MLX \
54
- --host 127.0.0.1 --port 8080 \
55
- --prompt-cache-size 0 \
56
- --prompt-concurrency 1 --decode-concurrency 1 \
57
- --prefill-step-size 512
58
- ```
59
-
60
- This disables retained prompt caches, limits concurrency to one, and processes
61
- prefill in smaller chunks. The model weights and their quantization are unchanged.
62
- The active request still needs its own cache and temporary memory; these options
63
- do not set a total process-memory cap or guarantee that any context length fits.
64
-
65
- Disabling reuse can slow later turns because their history must be processed
66
- again. It does not remove history that the client includes in the next request.
67
- Start with a fresh short conversation and increase history while monitoring RAM.
68
- The model's configured context limit is distinct from the memory needed to run it.
69
-
70
- Validation scope: three successive HTTP requests, including a streaming request,
71
- passed on each small synthetic 4-bit and 5-bit architecture with these settings.
72
- Retained-cache accounting stayed at zero. These tests did not load the complete
73
- Swift checkpoint or establish full-model long-context capacity on any Mac.
74
- No macOS or MLX memory-limit overrides were used in those tests.
75
 
76
  ## 3. If a server reports `404 generation thread died`
77
 
 
43
  Do not remove vision/MTP parameters or change the config to force a stock text
44
  loader to accept the file. All 2,379 saved tensor entries belong to this release.
45
 
46
+ ## Server cache update
47
+
48
+ The supplied server patch keeps prompt reuse enabled with at most two retained
49
+ entries, enforces a retained-cache byte budget at insertion, and defaults to one
50
+ prompt and one decode stream with 512-token prefill steps. On Metal the default
51
+ budget estimates headroom after model weights and workspace; an explicit
52
+ `--prompt-cache-bytes` overrides it. `--prompt-cache-size 2` means two entries,
53
+ not two GB.
54
+
55
+ See [SERVER_CACHE_UPDATE.md](SERVER_CACHE_UPDATE.md) for installation, offline
56
+ startup, validation and rollback. Existing installations must apply this runtime
57
+ update and restart. No model-weight download or re-quantization is needed.
58
+
59
+ The retained-cache limit does not cap active-request or total process memory.
60
+ An entry larger than the budget is evicted, so reuse depends on what fits. Use
61
+ `--prompt-cache-size 0` only when intentionally disabling reuse. The model's
62
+ configured context length is not a promise that it fits every Mac.
63
+
64
+ Validation: 44 unit/upstream server tests and 72 offline HTTP requests passed
65
+ across tiny 4-bit and 5-bit fixtures. These checks cover retention, warm reuse,
66
+ streaming and sequential generation. They do not establish full-model
67
+ long-context capacity or GUI integration.
 
 
 
 
 
 
 
68
 
69
  ## 3. If a server reports `404 generation thread died`
70
 
UPLOAD_MANIFEST.json CHANGED
@@ -36,17 +36,20 @@
36
  "sha256": "e0ce3327646a3dde90ee001a3bb819c8cb24e550a83d28340671f2542e7ad55b",
37
  "git_blob_sha1": "008e4ab150bc188e372cad12fda89317afda850e"
38
  },
 
 
 
 
 
39
  {
40
  "path": "TROUBLESHOOTING.md",
41
- "bytes": 5865,
42
- "sha256": "d94c7c741398c55b8f9abaa1f1542d7046dc6fab1df7e9ea4b13b59af3945dee",
43
- "git_blob_sha1": "3af3c505aff07e4e527e1d2d01b7b8f4cdb5288f"
44
  },
45
  {
46
  "path": "USAGE.md",
47
- "bytes": 6912,
48
- "sha256": "5c401d9f19fdd22c7bdc495bfdc925bc4ef6f9f829827a22918e26e0ba67d0e4",
49
- "git_blob_sha1": "8532eaa9e4c661635164a08e3fb00c597191b1a0"
50
  },
51
  {
52
  "path": "chat_template.jinja",
@@ -66,6 +69,36 @@
66
  "sha256": "ccfab7ccb2ea306f71531c8ca77bb55507606cd90768b1e32b8b52ab5b48cf01",
67
  "git_blob_sha1": "98ff47b9ef9d4ac9f1a4bde3db13dd27456c5ea1"
68
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
69
  {
70
  "path": "compatibility/conversion-result.json",
71
  "bytes": 694,
@@ -126,6 +159,11 @@
126
  "sha256": "f6f1d0bdafa45863bfbf93dac0398c481c993ea04fdf38b9bae98c643f89eaec",
127
  "git_blob_sha1": "b739acc2e3b4cb82f073b7280255bcb898bb2a33"
128
  },
 
 
 
 
 
129
  {
130
  "path": "config.json",
131
  "bytes": 4577,
@@ -218,5 +256,7 @@
218
  ],
219
  "file_count": 37,
220
  "total_bytes": 19310880490,
221
- "self_excluded": true
 
 
222
  }
 
36
  "sha256": "e0ce3327646a3dde90ee001a3bb819c8cb24e550a83d28340671f2542e7ad55b",
37
  "git_blob_sha1": "008e4ab150bc188e372cad12fda89317afda850e"
38
  },
39
+ {
40
+ "path": "SERVER_CACHE_UPDATE.md",
41
+ "bytes": 4765,
42
+ "sha256": "31ae8115dcf0ccba9e91dd03f9eab30dc56a6a752fa26955c05555cc655bdcd7"
43
+ },
44
  {
45
  "path": "TROUBLESHOOTING.md",
46
+ "bytes": 5623,
47
+ "sha256": "8d631dcdacfd6f98d0ac37c4d2532f1a8139ab2e336c8ae27bfe03c251444327"
 
48
  },
49
  {
50
  "path": "USAGE.md",
51
+ "bytes": 6858,
52
+ "sha256": "946ba223dd2a7cf1aceb069cf4fd82f4bf609366abacba61017f08e2bf9a29b4"
 
53
  },
54
  {
55
  "path": "chat_template.jinja",
 
69
  "sha256": "ccfab7ccb2ea306f71531c8ca77bb55507606cd90768b1e32b8b52ab5b48cf01",
70
  "git_blob_sha1": "98ff47b9ef9d4ac9f1a4bde3db13dd27456c5ea1"
71
  },
72
+ {
73
+ "path": "compatibility/cache-tests/offline-4bit-results.json",
74
+ "bytes": 5014,
75
+ "sha256": "ddcfa5f200aa57d1cdadc6225299329b592e858bea2502e9bf07d313f4cf149a"
76
+ },
77
+ {
78
+ "path": "compatibility/cache-tests/offline-5bit-results.json",
79
+ "bytes": 5014,
80
+ "sha256": "f50ca1c21eafc43623f1461999e33e99b9e684a990d926aa77fbe59cb919fe25"
81
+ },
82
+ {
83
+ "path": "compatibility/cache-tests/synthetic-config.json",
84
+ "bytes": 1814,
85
+ "sha256": "852fa793054753245559ffeae66d4367b6009f3ebdb7ca611747ec7564d785c7"
86
+ },
87
+ {
88
+ "path": "compatibility/cache-tests/test_server_cache_http.py",
89
+ "bytes": 9689,
90
+ "sha256": "e7aff50b441145338575a50c7abc5fb6d40fdac3d11bbf6433d909ecfe22483b"
91
+ },
92
+ {
93
+ "path": "compatibility/cache-tests/validation.json",
94
+ "bytes": 1984,
95
+ "sha256": "288450e5a3b34d4c45ce16489de749f49c7bd03cd2d28d4940d4bd94953022a5"
96
+ },
97
+ {
98
+ "path": "compatibility/cache-tests/verify_server_patch.py",
99
+ "bytes": 504,
100
+ "sha256": "4c3acf33b4b9ffe0131c45ad0584785767a81cf455f9f60d9ec814e4b7eaf61a"
101
+ },
102
  {
103
  "path": "compatibility/conversion-result.json",
104
  "bytes": 694,
 
159
  "sha256": "f6f1d0bdafa45863bfbf93dac0398c481c993ea04fdf38b9bae98c643f89eaec",
160
  "git_blob_sha1": "b739acc2e3b4cb82f073b7280255bcb898bb2a33"
161
  },
162
+ {
163
+ "path": "compatibility/swift15-server-cache.patch",
164
+ "bytes": 9356,
165
+ "sha256": "a7fbc0f0524ee7864d9f41a98a2e35acf0e63f53e369ee5b9a87ca12926a9eab"
166
+ },
167
  {
168
  "path": "config.json",
169
  "bytes": 4577,
 
256
  ],
257
  "file_count": 37,
258
  "total_bytes": 19310880490,
259
+ "self_excluded": true,
260
+ "total_bytes_excluding_manifest": 19310918334,
261
+ "server_cache_update": "2026-09-25: bounded cache reuse runtime patch and recorded local tests; weights unchanged"
262
  }
USAGE.md CHANGED
@@ -38,6 +38,8 @@ git -C swift5-mlx-lm apply --check ../Swift-1.5-5bit-MLX/compatibility/swift15-m
38
  git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/swift15-mlx-lm.patch
39
  git -C swift5-mlx-lm apply --check ../Swift-1.5-5bit-MLX/compatibility/enable-5bit.patch
40
  git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/enable-5bit.patch
 
 
41
  ```
42
 
43
  Apple Silicon:
@@ -102,35 +104,28 @@ The explicit `model.mtp_logits` step and `model.visual` encoder have historical
102
  component evidence. Integrated image/video chat and speculative generation are
103
  not implemented; unsupported multimodal generation must not be reported as working.
104
 
105
- ## Server settings for limited memory
106
 
107
- The pinned MLX-LM server retains prompt caches for reuse by default. For a setup
108
- that favors lower memory use over reuse speed, stop the existing server and start
109
- it from the same patched Python environment with:
 
 
 
110
 
111
- ```bash
112
- mlx_lm.server --model ukisai/Swift-1.5-5bit-MLX \
113
- --host 127.0.0.1 --port 8080 \
114
- --prompt-cache-size 0 \
115
- --prompt-concurrency 1 --decode-concurrency 1 \
116
- --prefill-step-size 512
117
- ```
 
118
 
119
- This disables retained prompt caches, limits concurrency to one, and processes
120
- prefill in smaller chunks. The model weights and their quantization are unchanged.
121
- The active request still needs its own cache and temporary memory; these options
122
- do not set a total process-memory cap or guarantee that any context length fits.
123
-
124
- Disabling reuse can slow later turns because their history must be processed
125
- again. It does not remove history that the client includes in the next request.
126
- Start with a fresh short conversation and increase history while monitoring RAM.
127
- The model's configured context limit is distinct from the memory needed to run it.
128
-
129
- Validation scope: three successive HTTP requests, including a streaming request,
130
- passed on each small synthetic 4-bit and 5-bit architecture with these settings.
131
- Retained-cache accounting stayed at zero. These tests did not load the complete
132
- Swift checkpoint or establish full-model long-context capacity on any Mac.
133
- No macOS or MLX memory-limit overrides were used in those tests.
134
 
135
  ## Conversion provenance
136
 
 
38
  git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/swift15-mlx-lm.patch
39
  git -C swift5-mlx-lm apply --check ../Swift-1.5-5bit-MLX/compatibility/enable-5bit.patch
40
  git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/enable-5bit.patch
41
+ git -C swift5-mlx-lm apply --check ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
42
+ git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
43
  ```
44
 
45
  Apple Silicon:
 
104
  component evidence. Integrated image/video chat and speculative generation are
105
  not implemented; unsupported multimodal generation must not be reported as working.
106
 
107
+ ## Server cache update
108
 
109
+ The supplied server patch keeps prompt reuse enabled with at most two retained
110
+ entries, enforces a retained-cache byte budget at insertion, and defaults to one
111
+ prompt and one decode stream with 512-token prefill steps. On Metal the default
112
+ budget estimates headroom after model weights and workspace; an explicit
113
+ `--prompt-cache-bytes` overrides it. `--prompt-cache-size 2` means two entries,
114
+ not two GB.
115
 
116
+ See [SERVER_CACHE_UPDATE.md](SERVER_CACHE_UPDATE.md) for installation, offline
117
+ startup, validation and rollback. Existing installations must apply this runtime
118
+ update and restart. No model-weight download or re-quantization is needed.
119
+
120
+ The retained-cache limit does not cap active-request or total process memory.
121
+ An entry larger than the budget is evicted, so reuse depends on what fits. Use
122
+ `--prompt-cache-size 0` only when intentionally disabling reuse. The model's
123
+ configured context length is not a promise that it fits every Mac.
124
 
125
+ Validation: 44 unit/upstream server tests and 72 offline HTTP requests passed
126
+ across tiny 4-bit and 5-bit fixtures. These checks cover retention, warm reuse,
127
+ streaming and sequential generation. They do not establish full-model
128
+ long-context capacity or GUI integration.
 
 
 
 
 
 
 
 
 
 
 
129
 
130
  ## Conversion provenance
131
 
compatibility/cache-tests/offline-4bit-results.json ADDED
@@ -0,0 +1,222 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "scope": "Tiny random weights with real local tokenizer assets. No released model weights loaded or changed; no full-model capacity claim.",
3
+ "bits": 4,
4
+ "offline_hub": true,
5
+ "memory_limit_overrides": false,
6
+ "machine": "arm64",
7
+ "python": "3.12.13",
8
+ "packages": {
9
+ "mlx": "0.32.2",
10
+ "mlx-lm": "0.32.0",
11
+ "transformers": "5.14.1",
12
+ "huggingface_hub": "1.31.0"
13
+ },
14
+ "profiles": [
15
+ {
16
+ "name": "defaults",
17
+ "requests": [
18
+ {
19
+ "http_status": 200,
20
+ "prompt_tokens": 225,
21
+ "cached_tokens": 0
22
+ },
23
+ {
24
+ "http_status": 200,
25
+ "prompt_tokens": 246,
26
+ "cached_tokens": 221
27
+ },
28
+ {
29
+ "http_status": 200,
30
+ "prompt_tokens": 267,
31
+ "cached_tokens": 242
32
+ },
33
+ {
34
+ "http_status": 200,
35
+ "prompt_tokens": 288,
36
+ "cached_tokens": 263
37
+ },
38
+ {
39
+ "http_status": 200,
40
+ "prompt_tokens": 309,
41
+ "cached_tokens": 284
42
+ },
43
+ {
44
+ "http_status": 200,
45
+ "prompt_tokens": 330,
46
+ "cached_tokens": 305
47
+ },
48
+ {
49
+ "http_status": 200,
50
+ "prompt_tokens": 351,
51
+ "cached_tokens": 326
52
+ },
53
+ {
54
+ "http_status": 200,
55
+ "prompt_tokens": 372,
56
+ "cached_tokens": 347
57
+ },
58
+ {
59
+ "http_status": 200,
60
+ "prompt_tokens": 393,
61
+ "cached_tokens": 368
62
+ },
63
+ {
64
+ "http_status": 200,
65
+ "prompt_tokens": 414,
66
+ "cached_tokens": 389
67
+ },
68
+ {
69
+ "http_status": 200,
70
+ "stream_complete": true
71
+ },
72
+ {
73
+ "http_status": 200,
74
+ "prompt_tokens": 457,
75
+ "cached_tokens": 431
76
+ }
77
+ ],
78
+ "byte_limit": 4190923688,
79
+ "peak_retained_bytes": 246528,
80
+ "retained_sequences": 2,
81
+ "insertions": 35
82
+ },
83
+ {
84
+ "name": "explicit_bytes",
85
+ "requests": [
86
+ {
87
+ "http_status": 200,
88
+ "prompt_tokens": 225,
89
+ "cached_tokens": 0
90
+ },
91
+ {
92
+ "http_status": 200,
93
+ "prompt_tokens": 246,
94
+ "cached_tokens": 112
95
+ },
96
+ {
97
+ "http_status": 200,
98
+ "prompt_tokens": 267,
99
+ "cached_tokens": 0
100
+ },
101
+ {
102
+ "http_status": 200,
103
+ "prompt_tokens": 288,
104
+ "cached_tokens": 112
105
+ },
106
+ {
107
+ "http_status": 200,
108
+ "prompt_tokens": 309,
109
+ "cached_tokens": 0
110
+ },
111
+ {
112
+ "http_status": 200,
113
+ "prompt_tokens": 330,
114
+ "cached_tokens": 112
115
+ },
116
+ {
117
+ "http_status": 200,
118
+ "prompt_tokens": 351,
119
+ "cached_tokens": 0
120
+ },
121
+ {
122
+ "http_status": 200,
123
+ "prompt_tokens": 372,
124
+ "cached_tokens": 112
125
+ },
126
+ {
127
+ "http_status": 200,
128
+ "prompt_tokens": 393,
129
+ "cached_tokens": 0
130
+ },
131
+ {
132
+ "http_status": 200,
133
+ "prompt_tokens": 414,
134
+ "cached_tokens": 112
135
+ },
136
+ {
137
+ "http_status": 200,
138
+ "stream_complete": true
139
+ },
140
+ {
141
+ "http_status": 200,
142
+ "prompt_tokens": 457,
143
+ "cached_tokens": 112
144
+ }
145
+ ],
146
+ "byte_limit": 100000,
147
+ "peak_retained_bytes": 82432,
148
+ "retained_sequences": 1,
149
+ "insertions": 40
150
+ },
151
+ {
152
+ "name": "no_retention",
153
+ "requests": [
154
+ {
155
+ "http_status": 200,
156
+ "prompt_tokens": 225,
157
+ "cached_tokens": 0
158
+ },
159
+ {
160
+ "http_status": 200,
161
+ "prompt_tokens": 246,
162
+ "cached_tokens": 0
163
+ },
164
+ {
165
+ "http_status": 200,
166
+ "prompt_tokens": 267,
167
+ "cached_tokens": 0
168
+ },
169
+ {
170
+ "http_status": 200,
171
+ "prompt_tokens": 288,
172
+ "cached_tokens": 0
173
+ },
174
+ {
175
+ "http_status": 200,
176
+ "prompt_tokens": 309,
177
+ "cached_tokens": 0
178
+ },
179
+ {
180
+ "http_status": 200,
181
+ "prompt_tokens": 330,
182
+ "cached_tokens": 0
183
+ },
184
+ {
185
+ "http_status": 200,
186
+ "prompt_tokens": 351,
187
+ "cached_tokens": 0
188
+ },
189
+ {
190
+ "http_status": 200,
191
+ "prompt_tokens": 372,
192
+ "cached_tokens": 0
193
+ },
194
+ {
195
+ "http_status": 200,
196
+ "prompt_tokens": 393,
197
+ "cached_tokens": 0
198
+ },
199
+ {
200
+ "http_status": 200,
201
+ "prompt_tokens": 414,
202
+ "cached_tokens": 0
203
+ },
204
+ {
205
+ "http_status": 200,
206
+ "stream_complete": true
207
+ },
208
+ {
209
+ "http_status": 200,
210
+ "prompt_tokens": 457,
211
+ "cached_tokens": 0
212
+ }
213
+ ],
214
+ "byte_limit": 4190923688,
215
+ "peak_retained_bytes": 0,
216
+ "retained_sequences": 0,
217
+ "insertions": 45
218
+ }
219
+ ],
220
+ "peak_mlx_bytes": 338846824,
221
+ "status": "PASS"
222
+ }
compatibility/cache-tests/offline-5bit-results.json ADDED
@@ -0,0 +1,222 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "scope": "Tiny random weights with real local tokenizer assets. No released model weights loaded or changed; no full-model capacity claim.",
3
+ "bits": 5,
4
+ "offline_hub": true,
5
+ "memory_limit_overrides": false,
6
+ "machine": "arm64",
7
+ "python": "3.12.13",
8
+ "packages": {
9
+ "mlx": "0.32.2",
10
+ "mlx-lm": "0.32.0",
11
+ "transformers": "5.14.1",
12
+ "huggingface_hub": "1.31.0"
13
+ },
14
+ "profiles": [
15
+ {
16
+ "name": "defaults",
17
+ "requests": [
18
+ {
19
+ "http_status": 200,
20
+ "prompt_tokens": 225,
21
+ "cached_tokens": 0
22
+ },
23
+ {
24
+ "http_status": 200,
25
+ "prompt_tokens": 246,
26
+ "cached_tokens": 221
27
+ },
28
+ {
29
+ "http_status": 200,
30
+ "prompt_tokens": 267,
31
+ "cached_tokens": 242
32
+ },
33
+ {
34
+ "http_status": 200,
35
+ "prompt_tokens": 288,
36
+ "cached_tokens": 263
37
+ },
38
+ {
39
+ "http_status": 200,
40
+ "prompt_tokens": 309,
41
+ "cached_tokens": 284
42
+ },
43
+ {
44
+ "http_status": 200,
45
+ "prompt_tokens": 330,
46
+ "cached_tokens": 305
47
+ },
48
+ {
49
+ "http_status": 200,
50
+ "prompt_tokens": 351,
51
+ "cached_tokens": 326
52
+ },
53
+ {
54
+ "http_status": 200,
55
+ "prompt_tokens": 372,
56
+ "cached_tokens": 347
57
+ },
58
+ {
59
+ "http_status": 200,
60
+ "prompt_tokens": 393,
61
+ "cached_tokens": 368
62
+ },
63
+ {
64
+ "http_status": 200,
65
+ "prompt_tokens": 414,
66
+ "cached_tokens": 389
67
+ },
68
+ {
69
+ "http_status": 200,
70
+ "stream_complete": true
71
+ },
72
+ {
73
+ "http_status": 200,
74
+ "prompt_tokens": 457,
75
+ "cached_tokens": 431
76
+ }
77
+ ],
78
+ "byte_limit": 4186895080,
79
+ "peak_retained_bytes": 246528,
80
+ "retained_sequences": 2,
81
+ "insertions": 35
82
+ },
83
+ {
84
+ "name": "explicit_bytes",
85
+ "requests": [
86
+ {
87
+ "http_status": 200,
88
+ "prompt_tokens": 225,
89
+ "cached_tokens": 0
90
+ },
91
+ {
92
+ "http_status": 200,
93
+ "prompt_tokens": 246,
94
+ "cached_tokens": 112
95
+ },
96
+ {
97
+ "http_status": 200,
98
+ "prompt_tokens": 267,
99
+ "cached_tokens": 0
100
+ },
101
+ {
102
+ "http_status": 200,
103
+ "prompt_tokens": 288,
104
+ "cached_tokens": 112
105
+ },
106
+ {
107
+ "http_status": 200,
108
+ "prompt_tokens": 309,
109
+ "cached_tokens": 0
110
+ },
111
+ {
112
+ "http_status": 200,
113
+ "prompt_tokens": 330,
114
+ "cached_tokens": 112
115
+ },
116
+ {
117
+ "http_status": 200,
118
+ "prompt_tokens": 351,
119
+ "cached_tokens": 0
120
+ },
121
+ {
122
+ "http_status": 200,
123
+ "prompt_tokens": 372,
124
+ "cached_tokens": 112
125
+ },
126
+ {
127
+ "http_status": 200,
128
+ "prompt_tokens": 393,
129
+ "cached_tokens": 0
130
+ },
131
+ {
132
+ "http_status": 200,
133
+ "prompt_tokens": 414,
134
+ "cached_tokens": 112
135
+ },
136
+ {
137
+ "http_status": 200,
138
+ "stream_complete": true
139
+ },
140
+ {
141
+ "http_status": 200,
142
+ "prompt_tokens": 457,
143
+ "cached_tokens": 112
144
+ }
145
+ ],
146
+ "byte_limit": 100000,
147
+ "peak_retained_bytes": 82432,
148
+ "retained_sequences": 1,
149
+ "insertions": 40
150
+ },
151
+ {
152
+ "name": "no_retention",
153
+ "requests": [
154
+ {
155
+ "http_status": 200,
156
+ "prompt_tokens": 225,
157
+ "cached_tokens": 0
158
+ },
159
+ {
160
+ "http_status": 200,
161
+ "prompt_tokens": 246,
162
+ "cached_tokens": 0
163
+ },
164
+ {
165
+ "http_status": 200,
166
+ "prompt_tokens": 267,
167
+ "cached_tokens": 0
168
+ },
169
+ {
170
+ "http_status": 200,
171
+ "prompt_tokens": 288,
172
+ "cached_tokens": 0
173
+ },
174
+ {
175
+ "http_status": 200,
176
+ "prompt_tokens": 309,
177
+ "cached_tokens": 0
178
+ },
179
+ {
180
+ "http_status": 200,
181
+ "prompt_tokens": 330,
182
+ "cached_tokens": 0
183
+ },
184
+ {
185
+ "http_status": 200,
186
+ "prompt_tokens": 351,
187
+ "cached_tokens": 0
188
+ },
189
+ {
190
+ "http_status": 200,
191
+ "prompt_tokens": 372,
192
+ "cached_tokens": 0
193
+ },
194
+ {
195
+ "http_status": 200,
196
+ "prompt_tokens": 393,
197
+ "cached_tokens": 0
198
+ },
199
+ {
200
+ "http_status": 200,
201
+ "prompt_tokens": 414,
202
+ "cached_tokens": 0
203
+ },
204
+ {
205
+ "http_status": 200,
206
+ "stream_complete": true
207
+ },
208
+ {
209
+ "http_status": 200,
210
+ "prompt_tokens": 457,
211
+ "cached_tokens": 0
212
+ }
213
+ ],
214
+ "byte_limit": 4186895080,
215
+ "peak_retained_bytes": 0,
216
+ "retained_sequences": 0,
217
+ "insertions": 45
218
+ }
219
+ ],
220
+ "peak_mlx_bytes": 430070044,
221
+ "status": "PASS"
222
+ }
compatibility/cache-tests/synthetic-config.json ADDED
@@ -0,0 +1,71 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "qwen3_5",
3
+ "architectures": [
4
+ "Qwen3_5ForConditionalGeneration"
5
+ ],
6
+ "language_model_only": false,
7
+ "tie_word_embeddings": false,
8
+ "image_token_id": 248056,
9
+ "video_token_id": 248057,
10
+ "vision_start_token_id": 248053,
11
+ "text_config": {
12
+ "model_type": "qwen3_5_text",
13
+ "hidden_size": 128,
14
+ "intermediate_size": 256,
15
+ "num_hidden_layers": 4,
16
+ "num_attention_heads": 4,
17
+ "num_key_value_heads": 2,
18
+ "head_dim": 32,
19
+ "vocab_size": 248320,
20
+ "full_attention_interval": 4,
21
+ "layer_types": [
22
+ "linear_attention",
23
+ "linear_attention",
24
+ "linear_attention",
25
+ "full_attention"
26
+ ],
27
+ "linear_num_key_heads": 2,
28
+ "linear_num_value_heads": 4,
29
+ "linear_key_head_dim": 32,
30
+ "linear_value_head_dim": 32,
31
+ "linear_conv_kernel_dim": 4,
32
+ "hidden_act": "silu",
33
+ "attn_output_gate": true,
34
+ "output_gate_type": "swish",
35
+ "mamba_ssm_dtype": "float32",
36
+ "rms_norm_eps": 1e-06,
37
+ "max_position_embeddings": 1024,
38
+ "tie_word_embeddings": false,
39
+ "attention_bias": false,
40
+ "attention_dropout": 0.0,
41
+ "mtp_num_hidden_layers": 1,
42
+ "mtp_use_dedicated_embeddings": false,
43
+ "rope_parameters": {
44
+ "rope_type": "default",
45
+ "rope_theta": 10000000,
46
+ "partial_rotary_factor": 0.5,
47
+ "mrope_interleaved": true,
48
+ "mrope_section": [
49
+ 3,
50
+ 3,
51
+ 2
52
+ ]
53
+ }
54
+ },
55
+ "vision_config": {
56
+ "model_type": "qwen3_5",
57
+ "depth": 2,
58
+ "hidden_size": 32,
59
+ "intermediate_size": 48,
60
+ "num_heads": 4,
61
+ "out_hidden_size": 128,
62
+ "num_position_embeddings": 16,
63
+ "patch_size": 2,
64
+ "temporal_patch_size": 2,
65
+ "spatial_merge_size": 2,
66
+ "in_channels": 3,
67
+ "hidden_act": "gelu_pytorch_tanh",
68
+ "deepstack_visual_indexes": []
69
+ },
70
+ "vision_end_token_id": 248054
71
+ }
compatibility/cache-tests/test_server_cache_http.py ADDED
@@ -0,0 +1,271 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Copyright © 2026 Swift contributors.
2
+
3
+ """Offline HTTP regression checks on tiny random models, never Swift weights."""
4
+
5
+ import argparse
6
+ import contextlib
7
+ import copy
8
+ import functools
9
+ import importlib.metadata
10
+ import json
11
+ import logging
12
+ import os
13
+ import platform
14
+ import shutil
15
+ import sys
16
+ import tempfile
17
+ import threading
18
+ import urllib.request
19
+ from http.server import ThreadingHTTPServer
20
+ from pathlib import Path
21
+ from unittest.mock import patch
22
+
23
+ os.environ["HF_HUB_OFFLINE"] = "1"
24
+ os.environ["TRANSFORMERS_OFFLINE"] = "1"
25
+
26
+ import mlx.core as mx
27
+ import mlx.nn as nn
28
+ from mlx.utils import tree_flatten
29
+
30
+ from mlx_lm import server
31
+ from mlx_lm.models.cache import LRUPromptCache
32
+ from mlx_lm.models.qwen3_5_full import Model, ModelArgs
33
+
34
+
35
+ class RecordingCache(LRUPromptCache):
36
+ def __init__(self, *args, **kwargs):
37
+ super().__init__(*args, **kwargs)
38
+ self.peak_bytes = 0
39
+ self.insertions = 0
40
+
41
+ def insert_cache(self, *args, **kwargs):
42
+ super().insert_cache(*args, **kwargs)
43
+ self.insertions += 1
44
+ self.peak_bytes = max(self.peak_bytes, self.nbytes)
45
+ assert self.nbytes <= self.max_bytes
46
+ assert len(self) <= self.max_size
47
+
48
+
49
+ def make_fixture(folder, assets, config_path, bits):
50
+ config = json.loads(config_path.read_text())
51
+ source = json.loads((assets / "config.json").read_text())
52
+ config["text_config"]["vocab_size"] = source["text_config"]["vocab_size"]
53
+ config.pop("quantization", None)
54
+ mx.random.seed(25)
55
+ model = Model(ModelArgs.from_dict(copy.deepcopy(config)))
56
+ model.language_model.lm_head.weight = mx.zeros_like(
57
+ model.language_model.lm_head.weight
58
+ )
59
+ model.apply(
60
+ lambda value: (
61
+ value.astype(mx.bfloat16)
62
+ if mx.issubdtype(value.dtype, mx.floating)
63
+ else value
64
+ )
65
+ )
66
+ nn.quantize(
67
+ model,
68
+ bits=bits,
69
+ group_size=64,
70
+ mode="affine",
71
+ class_predicate=lambda path, layer: hasattr(layer, "to_quantized")
72
+ and layer.weight.shape[-1] % 64 == 0,
73
+ )
74
+ mx.eval(model.parameters())
75
+ weights = dict(tree_flatten(model.parameters()))
76
+ mx.save_safetensors(
77
+ str(folder / "model.safetensors"),
78
+ weights,
79
+ metadata={"purpose": "synthetic regression only"},
80
+ )
81
+ index = {
82
+ "metadata": {"total_size": sum(w.nbytes for w in weights.values())},
83
+ "weight_map": {name: "model.safetensors" for name in weights},
84
+ }
85
+ (folder / "model.safetensors.index.json").write_text(json.dumps(index))
86
+ config["quantization"] = {"bits": bits, "group_size": 64, "mode": "affine"}
87
+ (folder / "config.json").write_text(json.dumps(config))
88
+ for name in (
89
+ "tokenizer.json",
90
+ "tokenizer_config.json",
91
+ "vocab.json",
92
+ "merges.txt",
93
+ "chat_template.jinja",
94
+ "generation_config.json",
95
+ ):
96
+ shutil.copyfile(assets / name, folder / name)
97
+ del model, weights
98
+ mx.clear_cache()
99
+
100
+
101
+ @contextlib.contextmanager
102
+ def running(folder, options):
103
+ captured = []
104
+ argv = ["mlx_lm.server", "--model", str(folder), "--port", "0", *options]
105
+ with patch.object(sys, "argv", argv), patch.object(
106
+ server, "run", side_effect=lambda h, p, m: captured.append(m)
107
+ ), patch.object(server, "maybe_set_recommended_wired_limit", return_value=None):
108
+ server.main()
109
+ provider = captured[0]
110
+ instances = []
111
+
112
+ def start_http(host, port, generator):
113
+ handler = functools.partial(
114
+ server.APIHandler, generator, system_fingerprint="synthetic-offline-test"
115
+ )
116
+ httpd = ThreadingHTTPServer((host, port), handler)
117
+ thread = threading.Thread(target=httpd.serve_forever, daemon=True)
118
+ instances.append((httpd, thread, generator))
119
+ thread.start()
120
+
121
+ with patch.object(server, "LRUPromptCache", RecordingCache), patch.object(
122
+ server, "_run_http_server", start_http
123
+ ):
124
+ server.run("127.0.0.1", 0, provider)
125
+ httpd, thread, generator = instances[0]
126
+ try:
127
+ yield f"http://127.0.0.1:{httpd.server_port}", generator, provider.cli_args
128
+ finally:
129
+ httpd.shutdown()
130
+ httpd.server_close()
131
+ thread.join(timeout=5)
132
+ generator.stop_and_join()
133
+
134
+
135
+ def request(url, messages, stream=False, seeded=False):
136
+ data = {
137
+ "model": "default_model",
138
+ "messages": messages,
139
+ "max_tokens": 4,
140
+ "temperature": 0,
141
+ "chat_template_kwargs": {"enable_thinking": False},
142
+ "stream": stream,
143
+ }
144
+ if seeded:
145
+ data["seed"] = 25
146
+ req = urllib.request.Request(
147
+ url + "/v1/chat/completions",
148
+ data=json.dumps(data).encode(),
149
+ headers={"Content-Type": "application/json"},
150
+ )
151
+ with urllib.request.urlopen(req, timeout=60) as response:
152
+ body = response.read().decode()
153
+ assert response.status == 200
154
+ if stream:
155
+ assert "data: [DONE]" in body
156
+ chunks = [
157
+ json.loads(line[6:])
158
+ for line in body.splitlines()
159
+ if line.startswith("data: ") and line != "data: [DONE]"
160
+ ]
161
+ text = "".join(
162
+ choice.get("delta", {}).get("content", "") or ""
163
+ for chunk in chunks
164
+ for choice in chunk.get("choices", [])
165
+ )
166
+ assert text == "!!!!", text
167
+ return {"http_status": 200, "stream_complete": True}
168
+ result = json.loads(body)
169
+ assert result["choices"][0]["message"]["content"] == "!!!!"
170
+ usage = result["usage"]
171
+ return {
172
+ "http_status": 200,
173
+ "prompt_tokens": usage["prompt_tokens"],
174
+ "cached_tokens": usage["prompt_tokens_details"]["cached_tokens"],
175
+ }
176
+
177
+
178
+ def main():
179
+ parser = argparse.ArgumentParser(description=__doc__)
180
+ parser.add_argument(
181
+ "--assets",
182
+ type=Path,
183
+ required=True,
184
+ help="Local Swift snapshot; only tokenizer and config assets are read",
185
+ )
186
+ parser.add_argument("--config", type=Path, required=True)
187
+ parser.add_argument("--bits", type=int, choices=(4, 5), required=True)
188
+ parser.add_argument("--output", type=Path, required=True)
189
+ args = parser.parse_args()
190
+ logging.basicConfig(level=logging.WARNING)
191
+ report = {
192
+ "scope": "Tiny random weights with real local tokenizer assets. "
193
+ "No released model weights loaded or changed; no full-model capacity claim.",
194
+ "bits": args.bits,
195
+ "offline_hub": True,
196
+ "memory_limit_overrides": False,
197
+ "machine": platform.machine(),
198
+ "python": platform.python_version(),
199
+ "packages": {
200
+ name: importlib.metadata.version(name)
201
+ for name in ("mlx", "mlx-lm", "transformers", "huggingface_hub")
202
+ },
203
+ "profiles": [],
204
+ }
205
+ profiles = [
206
+ ("defaults", []),
207
+ ("explicit_bytes", ["--prompt-cache-bytes", "100000"]),
208
+ ("no_retention", ["--prompt-cache-size", "0"]),
209
+ ]
210
+ with tempfile.TemporaryDirectory(prefix="swift-cache-test-") as temp:
211
+ folder = Path(temp)
212
+ make_fixture(folder, args.assets, args.config, args.bits)
213
+ for name, options in profiles:
214
+ with running(folder, options) as (url, generator, cli):
215
+ assert cli.prompt_concurrency == cli.decode_concurrency == 1
216
+ assert cli.prefill_step_size == 512
217
+ messages = [
218
+ {
219
+ "role": "system",
220
+ "content": "Help with coding. "
221
+ + "Read the session carefully. " * 20,
222
+ },
223
+ {
224
+ "role": "user",
225
+ "content": "Remember this context. " + "sample words " * 40,
226
+ },
227
+ {"role": "assistant", "content": "Previous response."},
228
+ {"role": "user", "content": "Continue briefly."},
229
+ ]
230
+ row = {"name": name, "requests": []}
231
+ for index in range(12):
232
+ result = request(
233
+ url, messages, stream=index == 10, seeded=index == 11
234
+ )
235
+ assert generator.generation_available()
236
+ row["requests"].append(result)
237
+ messages.extend(
238
+ [
239
+ {"role": "assistant", "content": "!!!!"},
240
+ {
241
+ "role": "user",
242
+ "content": f"Another short answer {index}.",
243
+ },
244
+ ]
245
+ )
246
+ cache = generator.prompt_cache
247
+ assert cache.insertions > 0
248
+ row.update(
249
+ byte_limit=cache.max_bytes,
250
+ peak_retained_bytes=cache.peak_bytes,
251
+ retained_sequences=len(cache),
252
+ insertions=cache.insertions,
253
+ )
254
+ nonstream = [r for r in row["requests"] if "cached_tokens" in r]
255
+ if name == "no_retention":
256
+ assert cache.peak_bytes == 0
257
+ assert all(r["cached_tokens"] == 0 for r in nonstream)
258
+ else:
259
+ assert any(r["cached_tokens"] > 0 for r in nonstream[1:])
260
+ assert 0 < cache.max_bytes < 1 << 63
261
+ if name == "explicit_bytes":
262
+ assert cache.max_bytes == 100000
263
+ report["profiles"].append(row)
264
+ report["peak_mlx_bytes"] = mx.get_peak_memory()
265
+ report["status"] = "PASS"
266
+ args.output.write_text(json.dumps(report, indent=2) + "\n")
267
+ print(json.dumps(report, indent=2))
268
+
269
+
270
+ if __name__ == "__main__":
271
+ main()
compatibility/cache-tests/validation.json ADDED
@@ -0,0 +1,55 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "status": "PASS_LOCAL_REGRESSION_CHECKS",
3
+ "base_commit": "c69d1288440a0dc4e6401fc417098b07598dccd5",
4
+ "mlx_version": "0.32.2",
5
+ "mlx_lm_version": "0.32.0",
6
+ "server_sha256_before": "57392576de114af80ede723976b9dd156dc234f856b9943b8b44b8b78bc27f79",
7
+ "server_sha256_after": "791f8c1eb3431c4d24b4f9084f67b665c38e3d547ee4a307e52c5b2b6f12aba0",
8
+ "patch_sha256": "a7fbc0f0524ee7864d9f41a98a2e35acf0e63f53e369ee5b9a87ca12926a9eab",
9
+ "unit_and_upstream_server_tests": {
10
+ "passed": 44,
11
+ "failed": 0
12
+ },
13
+ "offline_http_tests": {
14
+ "passed_requests": 72,
15
+ "failed_requests": 0
16
+ },
17
+ "scope": "Local Apple Silicon tests using small fixtures. Full Swift long-context capacity is not established.",
18
+ "hosted_ci": "NOT_RUN; these are recorded local checks, not a hosted CI status",
19
+ "weights_changed": false,
20
+ "quantization_changed": false,
21
+ "fresh_patch_replay": {
22
+ "base_commit": "c69d1288440a0dc4e6401fc417098b07598dccd5",
23
+ "status": "PASS",
24
+ "releases": [
25
+ {
26
+ "bits": 4,
27
+ "clean_patch_apply": "PASS",
28
+ "rollback_check": "PASS",
29
+ "runtime_identity": "PASS",
30
+ "model_implementation_unchanged": "PASS",
31
+ "unit_and_upstream_tests_passed": 44,
32
+ "offline_http_requests_passed": 36,
33
+ "black": "25.1.0 PASS",
34
+ "isort": "6.0.0 PASS",
35
+ "ruff": "0.16.6 PASS",
36
+ "patch_sha256": "a7fbc0f0524ee7864d9f41a98a2e35acf0e63f53e369ee5b9a87ca12926a9eab",
37
+ "status": "PASS"
38
+ },
39
+ {
40
+ "bits": 5,
41
+ "clean_patch_apply": "PASS",
42
+ "rollback_check": "PASS",
43
+ "runtime_identity": "PASS",
44
+ "model_implementation_unchanged": "PASS",
45
+ "unit_and_upstream_tests_passed": 44,
46
+ "offline_http_requests_passed": 36,
47
+ "black": "25.1.0 PASS",
48
+ "isort": "6.0.0 PASS",
49
+ "ruff": "0.16.6 PASS",
50
+ "patch_sha256": "a7fbc0f0524ee7864d9f41a98a2e35acf0e63f53e369ee5b9a87ca12926a9eab",
51
+ "status": "PASS"
52
+ }
53
+ ]
54
+ }
55
+ }
compatibility/cache-tests/verify_server_patch.py ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Verify the installed server patch without loading weights."""
2
+ import hashlib
3
+ import json
4
+ from pathlib import Path
5
+ from mlx_lm import server
6
+
7
+ manifest = json.loads((Path(__file__).parent / "validation.json").read_text())
8
+ actual = hashlib.sha256(Path(server.__file__).read_bytes()).hexdigest()
9
+ if actual != manifest["server_sha256_after"]:
10
+ raise SystemExit("FAIL: this Python environment does not use the checked server patch")
11
+ print("PASS: installed server source matches the checked cache patch")
compatibility/swift15-server-cache.patch ADDED
@@ -0,0 +1,253 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ diff --git a/mlx_lm/server.py b/mlx_lm/server.py
2
+ index 9462d6e..36f7787 100644
3
+ --- a/mlx_lm/server.py
4
+ +++ b/mlx_lm/server.py
5
+ @@ -29,6 +29,7 @@ from typing import (
6
+
7
+ import mlx.core as mx
8
+ from huggingface_hub import scan_cache_dir
9
+ +from mlx.utils import tree_flatten
10
+
11
+ from ._version import __version__
12
+ from .generate import (
13
+ @@ -428,6 +429,26 @@ def _format_top_logprobs(logprobs, top_n, tokenizer) -> Tuple[Dict[str, Any]]:
14
+ )
15
+
16
+
17
+ +def _prompt_cache_byte_limit(cli_args, model):
18
+ + limit = getattr(cli_args, "prompt_cache_bytes", None)
19
+ + if limit is not None:
20
+ + if limit < 0:
21
+ + raise ValueError("--prompt-cache-bytes must be non-negative")
22
+ + return limit
23
+ +
24
+ + # Reserve workspace and half the remaining capacity for the active request.
25
+ + try:
26
+ + recommended = mx.device_info().get("max_recommended_working_set_size")
27
+ + except (RuntimeError, ValueError):
28
+ + recommended = None
29
+ + if not recommended:
30
+ + return 1 << 63
31
+ + parameters = {id(p): p for _, p in tree_flatten(model.parameters())}
32
+ + model_bytes = sum(p.nbytes for p in parameters.values())
33
+ + reserve = max(4 * 1024**3, recommended // 8)
34
+ + return max(0, (recommended - model_bytes - reserve) // 2)
35
+ +
36
+ +
37
+ class ResponseGenerator:
38
+ def __init__(self, model_provider: ModelProvider, prompt_cache: LRUPromptCache):
39
+ self.model_provider = model_provider
40
+ @@ -443,6 +464,13 @@ class ResponseGenerator:
41
+ self._generation_thread = Thread(target=self._run_generate)
42
+ self._generation_thread.start()
43
+
44
+ + def _configure_prompt_cache(self, model):
45
+ + limit = _prompt_cache_byte_limit(self.cli_args, model)
46
+ + if self.prompt_cache.max_bytes != limit:
47
+ + self.prompt_cache.max_bytes = limit
48
+ + self.prompt_cache.trim_to(n_bytes=limit)
49
+ + logging.info("Retained prompt-cache limit: %.2f GB", limit / 1e9)
50
+ +
51
+ def _run_generate(self):
52
+ try:
53
+ self._generate()
54
+ @@ -772,6 +800,7 @@ class ResponseGenerator:
55
+ model, tokenizer = self.model_provider.load(
56
+ args.model.model, args.model.adapter, args.model.draft
57
+ )
58
+ + self._configure_prompt_cache(model)
59
+ except Exception as e:
60
+ rqueue.put(e)
61
+ continue
62
+ @@ -1762,7 +1791,13 @@ def run(
63
+ handler_class=APIHandler,
64
+ ):
65
+ group = mx.distributed.init()
66
+ - prompt_cache = LRUPromptCache(model_provider.cli_args.prompt_cache_size)
67
+ + cache_bytes = model_provider.cli_args.prompt_cache_bytes
68
+ + if cache_bytes is not None and cache_bytes < 0:
69
+ + raise ValueError("--prompt-cache-bytes must be non-negative")
70
+ + prompt_cache = LRUPromptCache(
71
+ + model_provider.cli_args.prompt_cache_size,
72
+ + max_bytes=cache_bytes if cache_bytes is not None else 1 << 63,
73
+ + )
74
+ response_generator = ResponseGenerator(model_provider, prompt_cache)
75
+ if group.rank() == 0:
76
+ _run_http_server(host, port, response_generator)
77
+ @@ -1875,31 +1910,32 @@ def main():
78
+ parser.add_argument(
79
+ "--decode-concurrency",
80
+ type=int,
81
+ - default=32,
82
+ + default=1,
83
+ help="When a request is batchable then decode that many requests in parallel",
84
+ )
85
+ parser.add_argument(
86
+ "--prompt-concurrency",
87
+ type=int,
88
+ - default=8,
89
+ + default=1,
90
+ help="When a request is batchable then process that many prompts in parallel",
91
+ )
92
+ parser.add_argument(
93
+ "--prefill-step-size",
94
+ type=int,
95
+ - default=2048,
96
+ - help="Step size for prefill processing (default: 2048)",
97
+ + default=512,
98
+ + help="Step size for prefill processing (default: 512)",
99
+ )
100
+ parser.add_argument(
101
+ "--prompt-cache-size",
102
+ type=int,
103
+ - default=10,
104
+ - help="Maximum number of distinct KV caches to hold in the prompt cache",
105
+ + default=2,
106
+ + help="Maximum retained prompt-cache entries (default: 2)",
107
+ )
108
+ parser.add_argument(
109
+ "--prompt-cache-bytes",
110
+ type=_parse_size,
111
+ - help="Maximum size in bytes of the KV caches",
112
+ + help="Maximum retained prompt-cache bytes. Default: automatic on Metal. "
113
+ + "This does not cap active-request or total process memory.",
114
+ )
115
+ parser.add_argument(
116
+ "--kv-bits",
117
+ diff --git a/tests/test_server_cache_budget.py b/tests/test_server_cache_budget.py
118
+ new file mode 100644
119
+ --- /dev/null
120
+ +++ b/tests/test_server_cache_budget.py
121
+ @@ -0,0 +1,132 @@
122
+ +# Copyright © 2026 Swift contributors.
123
+ +
124
+ +import sys
125
+ +from types import SimpleNamespace
126
+ +from unittest.mock import Mock
127
+ +
128
+ +import pytest
129
+ +
130
+ +from mlx_lm import server
131
+ +from mlx_lm.models.cache import LRUPromptCache
132
+ +
133
+ +
134
+ +class CacheState:
135
+ + def __init__(self, nbytes):
136
+ + self.nbytes = nbytes
137
+ +
138
+ + def is_trimmable(self):
139
+ + return False
140
+ +
141
+ +
142
+ +@pytest.mark.parametrize("limit", [0, 100])
143
+ +def test_server_factory_enforces_configured_bytes(monkeypatch, limit):
144
+ + caches = []
145
+ + provider = SimpleNamespace(
146
+ + cli_args=SimpleNamespace(prompt_cache_size=10, prompt_cache_bytes=limit)
147
+ + )
148
+ + monkeypatch.setattr(
149
+ + server, "ResponseGenerator", lambda provider, cache: caches.append(cache)
150
+ + )
151
+ + monkeypatch.setattr(server, "_run_http_server", lambda *args: None)
152
+ + server.run("127.0.0.1", 0, provider)
153
+ + cache = caches[0]
154
+ + for i in range(5):
155
+ + cache.insert_cache("model", [i, 1], [CacheState(80)])
156
+ + assert cache.nbytes <= limit
157
+ + if limit:
158
+ + reused, remaining = cache.fetch_nearest_cache("model", [4, 1, 9])
159
+ + assert reused is not None
160
+ + assert remaining == [9]
161
+ +
162
+ +
163
+ +def test_negative_limit_rejected_before_worker_start(monkeypatch):
164
+ + worker = Mock()
165
+ + monkeypatch.setattr(server, "ResponseGenerator", worker)
166
+ + provider = SimpleNamespace(
167
+ + cli_args=SimpleNamespace(prompt_cache_size=2, prompt_cache_bytes=-1)
168
+ + )
169
+ + with pytest.raises(ValueError, match="non-negative"):
170
+ + server.run("127.0.0.1", 0, provider)
171
+ + worker.assert_not_called()
172
+ +
173
+ +
174
+ +def test_explicit_limit_does_not_probe_model(monkeypatch):
175
+ + probe = Mock(side_effect=AssertionError("device probe was not needed"))
176
+ + monkeypatch.setattr(server.mx, "device_info", probe)
177
+ + model = Mock()
178
+ + for limit in (0, 1024, 8 * 1024**3):
179
+ + assert (
180
+ + server._prompt_cache_byte_limit(
181
+ + SimpleNamespace(prompt_cache_bytes=limit), model
182
+ + )
183
+ + == limit
184
+ + )
185
+ + model.parameters.assert_not_called()
186
+ +
187
+ +
188
+ +def test_auto_limit_reserves_room_for_active_request(monkeypatch):
189
+ + gib = 1024**3
190
+ + monkeypatch.setattr(
191
+ + server.mx,
192
+ + "device_info",
193
+ + lambda: {"max_recommended_working_set_size": 36 * gib},
194
+ + )
195
+ + model = SimpleNamespace(parameters=lambda: {"weight": CacheState(15 * gib)})
196
+ + args = SimpleNamespace(prompt_cache_bytes=None)
197
+ + limit = server._prompt_cache_byte_limit(args, model)
198
+ + assert 6 * gib <= limit < (36 - 15) * gib // 2
199
+ + larger = SimpleNamespace(parameters=lambda: {"weight": CacheState(30 * gib)})
200
+ + assert server._prompt_cache_byte_limit(args, larger) < limit
201
+ + full = SimpleNamespace(parameters=lambda: {"weight": CacheState(36 * gib)})
202
+ + assert server._prompt_cache_byte_limit(args, full) == 0
203
+ +
204
+ +
205
+ +def test_model_swap_trims_retained_states(monkeypatch):
206
+ + obj = server.ResponseGenerator.__new__(server.ResponseGenerator)
207
+ + obj.model_provider = SimpleNamespace(
208
+ + cli_args=SimpleNamespace(prompt_cache_bytes=100)
209
+ + )
210
+ + obj.prompt_cache = LRUPromptCache(max_size=10)
211
+ + obj.prompt_cache.insert_cache("old", [1], [CacheState(200)])
212
+ + obj._configure_prompt_cache(Mock())
213
+ + assert obj.prompt_cache.nbytes == 0
214
+ + assert obj.prompt_cache.max_bytes == 100
215
+ +
216
+ +
217
+ +@pytest.mark.parametrize("device_info", [{}, {"max_recommended_working_set_size": 0}])
218
+ +def test_auto_limit_without_metal_metadata(monkeypatch, device_info):
219
+ + monkeypatch.setattr(server.mx, "device_info", lambda: device_info)
220
+ + model = Mock()
221
+ + assert (
222
+ + server._prompt_cache_byte_limit(SimpleNamespace(prompt_cache_bytes=None), model)
223
+ + == 1 << 63
224
+ + )
225
+ + model.parameters.assert_not_called()
226
+ +
227
+ +
228
+ +def test_auto_limit_counts_tied_parameters_once(monkeypatch):
229
+ + gib = 1024**3
230
+ + monkeypatch.setattr(
231
+ + server.mx,
232
+ + "device_info",
233
+ + lambda: {"max_recommended_working_set_size": 32 * gib},
234
+ + )
235
+ + weight = CacheState(8 * gib)
236
+ + model = SimpleNamespace(parameters=lambda: {"embed": weight, "head": weight})
237
+ + assert (
238
+ + server._prompt_cache_byte_limit(SimpleNamespace(prompt_cache_bytes=None), model)
239
+ + == 10 * gib
240
+ + )
241
+ +
242
+ +
243
+ +def test_defaults_keep_reuse_enabled_and_limit_concurrency(monkeypatch):
244
+ + captured = []
245
+ + monkeypatch.setattr(sys, "argv", ["mlx_lm.server"])
246
+ + monkeypatch.setattr(server, "maybe_set_recommended_wired_limit", lambda: None)
247
+ + monkeypatch.setattr(server, "run", lambda h, p, m: captured.append(m.cli_args))
248
+ + server.main()
249
+ + args = captured[0]
250
+ + assert args.prompt_cache_size == 2
251
+ + assert args.prompt_cache_bytes is None
252
+ + assert args.prompt_concurrency == args.decode_concurrency == 1
253
+ + assert args.prefill_step_size == 512