ukisai commited on
Commit
c62a35c
·
verified ·
1 Parent(s): 5359607

Record full 4-bit and 5-bit Mac cache validation

Browse files

Existing complete checkpoints tested on a 48 GiB M4 Pro with native Metal, an 86k-token synthetic text prompt and cached follow-ups. No weights or runtime code changed.

FULL_MAC_VALIDATION.md ADDED
@@ -0,0 +1,55 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Full Swift Mac cache validation — 2026-09-25
2
+
3
+ Both complete, existing Swift checkpoints passed native Metal text generation
4
+ and two cached follow-up requests on an AWS M4 Pro Mac with 48 GiB RAM.
5
+ All weight shards were verified against their recorded SHA-256 hashes before
6
+ loading. No quantization, tensor edits, context changes or memory-limit overrides
7
+ were performed. Hub offline mode was enabled during testing.
8
+
9
+ The input is our own synthetic repeated-record conversation, followed by a short
10
+ request to reply READY. It is a memory/cache regression test, not a quality
11
+ benchmark, an exact reproduction of another person's conversation, or a claim
12
+ that every context fits. The follow-ups must reuse at least half the prompt to
13
+ pass; the test fails if the worker dies or the retained-cache budget is exceeded.
14
+
15
+ | Checkpoint | Initial prompt tokens | First request, seconds | Follow-ups, seconds | Follow-up cached tokens | Peak MLX allocation, GiB |
16
+ |---|---:|---:|---|---|---:|
17
+ | 4-bit | 86,004 | 893.64 | 1.46, 1.42 | 86000, 86020 | 31.57 |
18
+ | 5-bit | 86,004 | 907.00 | 1.45, 1.45 | 86000, 86020 | 34.79 |
19
+
20
+ The first request includes model loading and initial prompt processing. MLX peak
21
+ allocation is not total process or system memory. Complete request results,
22
+ generated replies, cache limits and sampled swap observations are recorded in
23
+ `compatibility/cache-tests/full-mac-4bit-results.json` and
24
+ `compatibility/cache-tests/full-mac-5bit-results.json`.
25
+
26
+ Environment: macOS 26.7 (25G229), Python 3.12.13, MLX 0.32.2, MLX-LM 0.32.0, Transformers 5.14.1,
27
+ Hugging Face Hub 1.31.0. Official MLX-LM base revision:
28
+ `c69d1288440a0dc4e6401fc417098b07598dccd5`.
29
+ The shared runtime uses the existing Swift architecture patch, its 5-bit support
30
+ extension, and the unchanged server-cache patch. The 5-bit extension also accepts
31
+ the original 4-bit format. The cache defaults are two retained entries, automatic
32
+ byte budgeting, one prompt/decode stream, and 512-token prefill steps.
33
+
34
+ To reproduce, install the pinned environment and all three patches in the order
35
+ documented by the 5-bit release, then use a complete local snapshot. Run one model
36
+ at a time on an Apple Silicon Mac with at least 48 GiB RAM:
37
+
38
+ ```bash
39
+ HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python \
40
+ Swift-1.5-4bit-MLX/compatibility/cache-tests/validate_full_mac_cache.py \
41
+ --snapshot Swift-1.5-4bit-MLX --output full-mac-4bit-results
42
+ HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python \
43
+ Swift-1.5-5bit-MLX/compatibility/cache-tests/validate_full_mac_cache.py \
44
+ --snapshot Swift-1.5-5bit-MLX --output full-mac-5bit-results
45
+ ```
46
+
47
+ Output directories must be new. The harness uses the real server and complete
48
+ released weights, without building or quantizing any synthetic model. It disables
49
+ the server's optional wired-limit adjustment and reports the actual Metal device.
50
+
51
+ These results establish the tested text workload on the stated 48 GiB machine.
52
+ They do not establish 24 GiB operation, arbitrary 262k-token workloads, GUI/plugin
53
+ integration, image/video chat, speculative MTP generation, or broad output quality.
54
+ Existing installations still need to apply the runtime patch and restart the server.
55
+ The recorded checks are independently executed tests, not a hosted CI status.
README.md CHANGED
@@ -35,8 +35,13 @@ its supported generation interface is text-only with the included MLX-LM patches
35
  **Runtime compatibility:** this complete checkpoint requires both supplied patches, in the order shown
36
  and the pinned Python installation in [USAGE.md](USAGE.md). The tested unpatched
37
  MLX-LM 0.32.0 loader rejects 501 saved vision entries. Downloading the model does
38
- not install these patches into an app's inference engine. GUI compatibility and
39
- full 27B Apple Silicon generation remain unverified.
 
 
 
 
 
40
 
41
  The complete weights total **19.28 GB (17.96 GiB) across 4 required shards**, plus
42
  config, index and tokenizer files. A single 5–6 GB file is not the complete model.
@@ -120,7 +125,7 @@ The CPU example promotes in-memory floating values to FP32 while retaining packe
120
  [diagnostic](compatibility/macos-quantized-matmul-diagnostic.json) and
121
  [package checks](compatibility/package-checks.json).
122
 
123
- Full 27B Apple Silicon generation is unverified. Integrated image/video chat and
124
  speculative MTP generation are **not implemented** by the patch. Vision/MTP weights
125
  and component checks do not establish those end-to-end capabilities. Runtime/cache
126
  and OS memory must be budgeted in addition to the tensor payload.
 
35
  **Runtime compatibility:** this complete checkpoint requires both supplied patches, in the order shown
36
  and the pinned Python installation in [USAGE.md](USAGE.md). The tested unpatched
37
  MLX-LM 0.32.0 loader rejects 501 saved vision entries. Downloading the model does
38
+ not install these patches into an app's inference engine. GUI compatibility remains unverified.
39
+
40
+ Full-model follow-up validation is now recorded in
41
+ [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md): both complete checkpoints passed
42
+ an 86k-token synthetic text conversation and two cached follow-ups on a 48 GiB
43
+ M4 Pro using the patched server. GUI integration and other memory/context sizes
44
+ remain outside that test.
45
 
46
  The complete weights total **19.28 GB (17.96 GiB) across 4 required shards**, plus
47
  config, index and tokenizer files. A single 5–6 GB file is not the complete model.
 
125
  [diagnostic](compatibility/macos-quantized-matmul-diagnostic.json) and
126
  [package checks](compatibility/package-checks.json).
127
 
128
+ Full-model text generation is covered in [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md). Integrated image/video chat and
129
  speculative MTP generation are **not implemented** by the patch. Vision/MTP weights
130
  and component checks do not establish those end-to-end capabilities. Runtime/cache
131
  and OS memory must be budgeted in addition to the tensor payload.
SERVER_CACHE_UPDATE.md CHANGED
@@ -77,8 +77,11 @@ cached. The separate HTTP harness sets offline mode and needs only local assets.
77
  It covers 12 requests each with automatic budgeting, an explicit byte limit, and
78
  disabled retention; streaming and seeded sequential execution are included.
79
  Synthetic output is deliberately fixed for transport/cache checks, not quality.
80
- Full Swift long-context execution on a larger Mac has not been independently
81
- retested by this change. GUI runtimes need their own integration.
 
 
 
82
 
83
  ## Roll back this runtime patch
84
 
 
77
  It covers 12 requests each with automatic budgeting, an explicit byte limit, and
78
  disabled retention; streaming and seeded sequential execution are included.
79
  Synthetic output is deliberately fixed for transport/cache checks, not quality.
80
+ Full-model follow-up validation is now recorded in
81
+ [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md): both complete checkpoints passed
82
+ an 86k-token synthetic text conversation and two cached follow-ups on a 48 GiB
83
+ M4 Pro using the patched server. GUI integration and other memory/context sizes
84
+ remain outside that test.
85
 
86
  ## Roll back this runtime patch
87
 
TROUBLESHOOTING.md CHANGED
@@ -100,3 +100,11 @@ against a dead worker will not recover it; restart only after addressing the cau
100
 
101
  These checks do not establish full-model Mac execution or GUI compatibility.
102
  Diagnosing an individual crash requires its original exception and runtime configuration.
 
 
 
 
 
 
 
 
 
100
 
101
  These checks do not establish full-model Mac execution or GUI compatibility.
102
  Diagnosing an individual crash requires its original exception and runtime configuration.
103
+
104
+ ## Full-model Mac follow-up
105
+
106
+ Full-model follow-up validation is now recorded in
107
+ [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md): both complete checkpoints passed
108
+ an 86k-token synthetic text conversation and two cached follow-ups on a 48 GiB
109
+ M4 Pro using the patched server. GUI integration and other memory/context sizes
110
+ remain outside that test.
UPLOAD_MANIFEST.json CHANGED
@@ -6,6 +6,11 @@
6
  "sha256": "8a9fd1f95733fb82a6cba50f5ea2b1e26ca7578bea7a7cc21b4301ff83b1ce3b",
7
  "git_blob_sha1": "3727b14f485a05787b6cda82f0346eaeaa2e751b"
8
  },
 
 
 
 
 
9
  {
10
  "path": "LICENSE",
11
  "bytes": 13306,
@@ -32,24 +37,23 @@
32
  },
33
  {
34
  "path": "README.md",
35
- "bytes": 8067,
36
- "sha256": "e0ce3327646a3dde90ee001a3bb819c8cb24e550a83d28340671f2542e7ad55b",
37
- "git_blob_sha1": "008e4ab150bc188e372cad12fda89317afda850e"
38
  },
39
  {
40
  "path": "SERVER_CACHE_UPDATE.md",
41
- "bytes": 4765,
42
- "sha256": "31ae8115dcf0ccba9e91dd03f9eab30dc56a6a752fa26955c05555cc655bdcd7"
43
  },
44
  {
45
  "path": "TROUBLESHOOTING.md",
46
- "bytes": 5623,
47
- "sha256": "8d631dcdacfd6f98d0ac37c4d2532f1a8139ab2e336c8ae27bfe03c251444327"
48
  },
49
  {
50
  "path": "USAGE.md",
51
- "bytes": 6858,
52
- "sha256": "946ba223dd2a7cf1aceb069cf4fd82f4bf609366abacba61017f08e2bf9a29b4"
53
  },
54
  {
55
  "path": "chat_template.jinja",
@@ -69,6 +73,16 @@
69
  "sha256": "ccfab7ccb2ea306f71531c8ca77bb55507606cd90768b1e32b8b52ab5b48cf01",
70
  "git_blob_sha1": "98ff47b9ef9d4ac9f1a4bde3db13dd27456c5ea1"
71
  },
 
 
 
 
 
 
 
 
 
 
72
  {
73
  "path": "compatibility/cache-tests/offline-4bit-results.json",
74
  "bytes": 5014,
@@ -89,6 +103,11 @@
89
  "bytes": 9689,
90
  "sha256": "e7aff50b441145338575a50c7abc5fb6d40fdac3d11bbf6433d909ecfe22483b"
91
  },
 
 
 
 
 
92
  {
93
  "path": "compatibility/cache-tests/validation.json",
94
  "bytes": 1984,
 
6
  "sha256": "8a9fd1f95733fb82a6cba50f5ea2b1e26ca7578bea7a7cc21b4301ff83b1ce3b",
7
  "git_blob_sha1": "3727b14f485a05787b6cda82f0346eaeaa2e751b"
8
  },
9
+ {
10
+ "path": "FULL_MAC_VALIDATION.md",
11
+ "bytes": 3246,
12
+ "sha256": "a3a164e3affcfbe8a506b0205f4d4e63a69807ea56f4a8baa06b2201b20b1faa"
13
+ },
14
  {
15
  "path": "LICENSE",
16
  "bytes": 13306,
 
37
  },
38
  {
39
  "path": "README.md",
40
+ "bytes": 8392,
41
+ "sha256": "25708f8dc77158d099de1e1155f64d3eef58bfed06c892d7d65f21125b6381b5"
 
42
  },
43
  {
44
  "path": "SERVER_CACHE_UPDATE.md",
45
+ "bytes": 4941,
46
+ "sha256": "aad61a1a3bf8352ebf214c14f232869f62b282f4fd16f510de63375d7a04d5c7"
47
  },
48
  {
49
  "path": "TROUBLESHOOTING.md",
50
+ "bytes": 5972,
51
+ "sha256": "79d58bcc5929fe4ea3a2becefedf189d2d105102dce07eb894ef08212aa9be6d"
52
  },
53
  {
54
  "path": "USAGE.md",
55
+ "bytes": 7207,
56
+ "sha256": "e8590662c34f04462a5da73cb2e0572a26a5a18c385a2dcd281daaacbf47e963"
57
  },
58
  {
59
  "path": "chat_template.jinja",
 
73
  "sha256": "ccfab7ccb2ea306f71531c8ca77bb55507606cd90768b1e32b8b52ab5b48cf01",
74
  "git_blob_sha1": "98ff47b9ef9d4ac9f1a4bde3db13dd27456c5ea1"
75
  },
76
+ {
77
+ "path": "compatibility/cache-tests/full-mac-4bit-results.json",
78
+ "bytes": 2699,
79
+ "sha256": "6087feec8158cd3f1df6e7c3394a51a5091ee96649066ad875f6b6ac866dd76b"
80
+ },
81
+ {
82
+ "path": "compatibility/cache-tests/full-mac-5bit-results.json",
83
+ "bytes": 2695,
84
+ "sha256": "274b30cdac2d75d0258ff79a91ec53ce9b7029f36d1d64edf8783037a9c67459"
85
+ },
86
  {
87
  "path": "compatibility/cache-tests/offline-4bit-results.json",
88
  "bytes": 5014,
 
103
  "bytes": 9689,
104
  "sha256": "e7aff50b441145338575a50c7abc5fb6d40fdac3d11bbf6433d909ecfe22483b"
105
  },
106
+ {
107
+ "path": "compatibility/cache-tests/validate_full_mac_cache.py",
108
+ "bytes": 8156,
109
+ "sha256": "3e48876b07fb73f88c489e3804c9676872b6c1972650447de5820f1966b5811e"
110
+ },
111
  {
112
  "path": "compatibility/cache-tests/validation.json",
113
  "bytes": 1984,
USAGE.md CHANGED
@@ -134,3 +134,11 @@ complete customized Swift BF16 source identified in `QUANTIZATION_MANIFEST.json`
134
  with its 18 shards verified before conversion, the pinned patched converter, and
135
  affine / 5-bit / group size 64. Do not fill missing weights with base Qwen or another
136
  quantization. Existing quantized weights are unchanged by these packaging repairs.
 
 
 
 
 
 
 
 
 
134
  with its 18 shards verified before conversion, the pinned patched converter, and
135
  affine / 5-bit / group size 64. Do not fill missing weights with base Qwen or another
136
  quantization. Existing quantized weights are unchanged by these packaging repairs.
137
+
138
+ ## Full-model Mac follow-up
139
+
140
+ Full-model follow-up validation is now recorded in
141
+ [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md): both complete checkpoints passed
142
+ an 86k-token synthetic text conversation and two cached follow-ups on a 48 GiB
143
+ M4 Pro using the patched server. GUI integration and other memory/context sizes
144
+ remain outside that test.
compatibility/cache-tests/full-mac-4bit-results.json ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "status": "PASS_FULL_MODEL_SYNTHETIC_LONG_CONTEXT_FOLLOWUPS",
3
+ "physical_memory_bytes": 51539607552,
4
+ "offline": true,
5
+ "model_weights_changed": false,
6
+ "memory_limit_overrides": false,
7
+ "device": "Device(gpu, 0)",
8
+ "device_info": {
9
+ "device_name": "Apple M4 Pro",
10
+ "max_recommended_working_set_size": 40200896512,
11
+ "memory_size": 51539607552,
12
+ "architecture": "applegpu_g16s",
13
+ "max_buffer_length": 30150672384,
14
+ "resource_limit": 499000
15
+ },
16
+ "packages": {
17
+ "mlx": "0.32.2",
18
+ "mlx-lm": "0.32.0",
19
+ "transformers": "5.14.1",
20
+ "huggingface_hub": "1.31.0"
21
+ },
22
+ "prompt_origin": "Synthetic repeated records; no third-party conversation",
23
+ "initial_prompt_tokens": 86004,
24
+ "requests": [
25
+ {
26
+ "turn": 1,
27
+ "elapsed_seconds": 893.6356901669999,
28
+ "usage": {
29
+ "prompt_tokens": 86004,
30
+ "completion_tokens": 2,
31
+ "total_tokens": 86006,
32
+ "prompt_tokens_details": {
33
+ "cached_tokens": 0
34
+ }
35
+ },
36
+ "generation": "READY",
37
+ "finish_reason": "stop",
38
+ "retained_bytes": 5945229312,
39
+ "retained_limit": 9674659152,
40
+ "retained_sequences": 2,
41
+ "peak_mlx_bytes": 33428188365
42
+ },
43
+ {
44
+ "turn": 2,
45
+ "elapsed_seconds": 1.4563554589999512,
46
+ "usage": {
47
+ "prompt_tokens": 86024,
48
+ "completion_tokens": 2,
49
+ "total_tokens": 86026,
50
+ "prompt_tokens_details": {
51
+ "cached_tokens": 86000
52
+ }
53
+ },
54
+ "generation": "READY",
55
+ "finish_reason": "stop",
56
+ "retained_bytes": 5946540032,
57
+ "retained_limit": 9674659152,
58
+ "retained_sequences": 2,
59
+ "peak_mlx_bytes": 33885106168
60
+ },
61
+ {
62
+ "turn": 3,
63
+ "elapsed_seconds": 1.4185313340001358,
64
+ "usage": {
65
+ "prompt_tokens": 86044,
66
+ "completion_tokens": 2,
67
+ "total_tokens": 86046,
68
+ "prompt_tokens_details": {
69
+ "cached_tokens": 86020
70
+ }
71
+ },
72
+ "generation": "READY",
73
+ "finish_reason": "stop",
74
+ "retained_bytes": 5947850752,
75
+ "retained_limit": 9674659152,
76
+ "retained_sequences": 2,
77
+ "peak_mlx_bytes": 33902489592
78
+ }
79
+ ],
80
+ "warm_reuse_observed": true,
81
+ "warm_reuse_fraction": [
82
+ 0.9997210080907654,
83
+ 0.9997210729394264
84
+ ],
85
+ "monitor_summary": {
86
+ "samples": 90,
87
+ "peak_active_mlx_bytes": 24817128310,
88
+ "peak_allocator_cache_bytes": 5144874723,
89
+ "swap_observations": [
90
+ "total = 0.00M used = 0.00M free = 0.00M (encrypted)"
91
+ ],
92
+ "worker_healthy_in_all_samples": true,
93
+ "retained_limit_observed_in_all_samples": true
94
+ },
95
+ "model_repository": "ukisai/Swift-1.5-4bit-MLX",
96
+ "tested_revision": "c4765c752caee282928f62685722e11000ec64cf"
97
+ }
compatibility/cache-tests/full-mac-5bit-results.json ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "status": "PASS_FULL_MODEL_SYNTHETIC_LONG_CONTEXT_FOLLOWUPS",
3
+ "physical_memory_bytes": 51539607552,
4
+ "offline": true,
5
+ "model_weights_changed": false,
6
+ "memory_limit_overrides": false,
7
+ "device": "Device(gpu, 0)",
8
+ "device_info": {
9
+ "device_name": "Apple M4 Pro",
10
+ "max_recommended_working_set_size": 40200896512,
11
+ "memory_size": 51539607552,
12
+ "architecture": "applegpu_g16s",
13
+ "max_buffer_length": 30150672384,
14
+ "resource_limit": 499000
15
+ },
16
+ "packages": {
17
+ "mlx": "0.32.2",
18
+ "mlx-lm": "0.32.0",
19
+ "transformers": "5.14.1",
20
+ "huggingface_hub": "1.31.0"
21
+ },
22
+ "prompt_origin": "Synthetic repeated records; no third-party conversation",
23
+ "initial_prompt_tokens": 86004,
24
+ "requests": [
25
+ {
26
+ "turn": 1,
27
+ "elapsed_seconds": 906.998991333,
28
+ "usage": {
29
+ "prompt_tokens": 86004,
30
+ "completion_tokens": 2,
31
+ "total_tokens": 86006,
32
+ "prompt_tokens_details": {
33
+ "cached_tokens": 0
34
+ }
35
+ },
36
+ "generation": "READY",
37
+ "finish_reason": "stop",
38
+ "retained_bytes": 5945229312,
39
+ "retained_limit": 7946990032,
40
+ "retained_sequences": 2,
41
+ "peak_mlx_bytes": 36879525149
42
+ },
43
+ {
44
+ "turn": 2,
45
+ "elapsed_seconds": 1.4516940830003477,
46
+ "usage": {
47
+ "prompt_tokens": 86024,
48
+ "completion_tokens": 2,
49
+ "total_tokens": 86026,
50
+ "prompt_tokens_details": {
51
+ "cached_tokens": 86000
52
+ }
53
+ },
54
+ "generation": "READY",
55
+ "finish_reason": "stop",
56
+ "retained_bytes": 5946540032,
57
+ "retained_limit": 7946990032,
58
+ "retained_sequences": 2,
59
+ "peak_mlx_bytes": 37340425976
60
+ },
61
+ {
62
+ "turn": 3,
63
+ "elapsed_seconds": 1.4471444169998904,
64
+ "usage": {
65
+ "prompt_tokens": 86044,
66
+ "completion_tokens": 2,
67
+ "total_tokens": 86046,
68
+ "prompt_tokens_details": {
69
+ "cached_tokens": 86020
70
+ }
71
+ },
72
+ "generation": "READY",
73
+ "finish_reason": "stop",
74
+ "retained_bytes": 5947850752,
75
+ "retained_limit": 7946990032,
76
+ "retained_sequences": 2,
77
+ "peak_mlx_bytes": 37358006008
78
+ }
79
+ ],
80
+ "warm_reuse_observed": true,
81
+ "warm_reuse_fraction": [
82
+ 0.9997210080907654,
83
+ 0.9997210729394264
84
+ ],
85
+ "monitor_summary": {
86
+ "samples": 91,
87
+ "peak_active_mlx_bytes": 28351155232,
88
+ "peak_allocator_cache_bytes": 5353876378,
89
+ "swap_observations": [
90
+ "total = 0.00M used = 0.00M free = 0.00M (encrypted)"
91
+ ],
92
+ "worker_healthy_in_all_samples": true,
93
+ "retained_limit_observed_in_all_samples": true
94
+ },
95
+ "model_repository": "ukisai/Swift-1.5-5bit-MLX",
96
+ "tested_revision": "5359607ded7c88e3c7379a0adfb879eaf483cd98"
97
+ }
compatibility/cache-tests/validate_full_mac_cache.py ADDED
@@ -0,0 +1,166 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Validate the existing full checkpoint on an Apple Silicon Mac with >=48 GiB."""
2
+
3
+ import argparse
4
+ import importlib.util
5
+ import importlib.metadata
6
+ import json
7
+ import logging
8
+ import os
9
+ from pathlib import Path
10
+ import platform
11
+ import subprocess
12
+ import sys
13
+ import threading
14
+ import time
15
+ import urllib.request
16
+
17
+
18
+ def main():
19
+ parser = argparse.ArgumentParser(description=__doc__)
20
+ parser.add_argument("--snapshot", type=Path, required=True)
21
+ parser.add_argument("--output", type=Path, required=True)
22
+ parser.add_argument("--target-prompt-tokens", type=int, default=86000)
23
+ args = parser.parse_args()
24
+ if platform.system() != "Darwin" or platform.machine() != "arm64":
25
+ raise SystemExit("This validation requires a native Apple Silicon Mac.")
26
+ physical_bytes = int(subprocess.check_output(["sysctl", "-n", "hw.memsize"]))
27
+ if physical_bytes < 48 * 1024**3:
28
+ raise SystemExit("Refusing a full checkpoint load: at least 48 GiB is required for this test.")
29
+ if args.target_prompt_tokens < 4096:
30
+ raise SystemExit("Use at least 4096 tokens for the full-model follow-up test.")
31
+ args.snapshot = args.snapshot.resolve(strict=True)
32
+ args.output.mkdir(parents=True, exist_ok=False)
33
+ os.environ["HF_HUB_OFFLINE"] = "1"
34
+ os.environ["TRANSFORMERS_OFFLINE"] = "1"
35
+ logging.basicConfig(level=logging.INFO)
36
+
37
+ # Check files and the installed patch before loading model weights.
38
+ for command in (
39
+ [sys.executable, str(args.snapshot / "check_download.py"),
40
+ str(args.snapshot), "--hash", "--runtime"],
41
+ [sys.executable, str(args.snapshot / "compatibility/cache-tests/verify_server_patch.py")],
42
+ ):
43
+ subprocess.run(command, check=True)
44
+
45
+ import mlx.core as mx
46
+ from transformers import AutoTokenizer
47
+
48
+ assert mx.metal.is_available(), "Metal is required; a CPU run cannot validate this crash"
49
+ assert mx.default_device() != mx.cpu, "The validation must execute on the Metal GPU"
50
+ assert importlib.metadata.version("mlx") == "0.32.2"
51
+ assert importlib.metadata.version("mlx-lm") == "0.32.0"
52
+
53
+ harness_path = args.snapshot / "compatibility/cache-tests/test_server_cache_http.py"
54
+ spec = importlib.util.spec_from_file_location("cache_harness", harness_path)
55
+ harness = importlib.util.module_from_spec(spec)
56
+ spec.loader.exec_module(harness)
57
+ tokenizer = AutoTokenizer.from_pretrained(args.snapshot, local_files_only=True)
58
+
59
+ def prompt_length(messages):
60
+ tokens = tokenizer.apply_chat_template(
61
+ messages, tokenize=True, return_dict=False,
62
+ add_generation_prompt=True, enable_thinking=False,
63
+ )
64
+ assert isinstance(tokens, list) and tokens and isinstance(tokens[0], int)
65
+ return len(tokens)
66
+
67
+ def make_messages(count):
68
+ return [
69
+ {"role": "system", "content": "Read the records. Reply briefly to the final question."},
70
+ {"role": "user", "content": ("Record: alpha beta gamma delta epsilon.\n" * count)},
71
+ {"role": "assistant", "content": "I have read the records."},
72
+ {"role": "user", "content": "Reply with the single word READY."},
73
+ ]
74
+
75
+ lo, hi = 1, args.target_prompt_tokens
76
+ while lo < hi:
77
+ mid = (lo + hi) // 2
78
+ if prompt_length(make_messages(mid)) < args.target_prompt_tokens:
79
+ lo = mid + 1
80
+ else:
81
+ hi = mid
82
+ messages = make_messages(lo)
83
+ initial_tokens = prompt_length(messages)
84
+ assert args.target_prompt_tokens <= initial_tokens < args.target_prompt_tokens + 128
85
+ print(f"Validated synthetic prompt length: {initial_tokens} tokens", flush=True)
86
+ report = {"status": "RUNNING", "physical_memory_bytes": physical_bytes,
87
+ "offline": True, "model_weights_changed": False,
88
+ "memory_limit_overrides": False,
89
+ "device": str(mx.default_device()), "device_info": mx.device_info(),
90
+ "packages": {name: importlib.metadata.version(name)
91
+ for name in ("mlx", "mlx-lm", "transformers", "huggingface_hub")},
92
+ "prompt_origin": "Synthetic repeated records; no third-party conversation",
93
+ "initial_prompt_tokens": initial_tokens, "requests": []}
94
+ output_file = args.output / "results.json"
95
+ stopped = threading.Event()
96
+
97
+ def save():
98
+ output_file.write_text(json.dumps(report, indent=2) + "\n")
99
+
100
+ def monitor(generator):
101
+ with (args.output / "memory.jsonl").open("w") as stream:
102
+ while not stopped.is_set():
103
+ cache = generator.prompt_cache
104
+ row = {"time": time.time(), "active_mlx_bytes": mx.get_active_memory(),
105
+ "allocator_cache_bytes": mx.get_cache_memory(),
106
+ "peak_mlx_bytes": mx.get_peak_memory(),
107
+ "retained_bytes": cache.nbytes, "retained_limit": cache.max_bytes,
108
+ "retained_sequences": len(cache),
109
+ "worker_available": generator.generation_available(),
110
+ "swap": subprocess.check_output(["sysctl", "-n", "vm.swapusage"], text=True).strip()}
111
+ stream.write(json.dumps(row) + "\n")
112
+ stream.flush()
113
+ stopped.wait(10)
114
+
115
+ save()
116
+ try:
117
+ with harness.running(args.snapshot, []) as (url, generator, cli):
118
+ observer = threading.Thread(target=monitor, args=(generator,), daemon=True)
119
+ observer.start()
120
+ try:
121
+ for turn in range(3):
122
+ body = {"model": "default_model", "messages": messages, "max_tokens": 16,
123
+ "temperature": 0, "chat_template_kwargs": {"enable_thinking": False}}
124
+ req = urllib.request.Request(url + "/v1/chat/completions",
125
+ data=json.dumps(body).encode(), headers={"Content-Type": "application/json"})
126
+ started = time.monotonic()
127
+ with urllib.request.urlopen(req, timeout=3600) as response:
128
+ result = json.load(response)
129
+ assert response.status == 200
130
+ choice = result["choices"][0]
131
+ content = choice["message"].get("content")
132
+ assert isinstance(content, str) and content.strip(), result
133
+ assert generator.generation_available()
134
+ cache = generator.prompt_cache
135
+ row = {"turn": turn + 1, "elapsed_seconds": time.monotonic() - started,
136
+ "usage": result["usage"], "generation": content,
137
+ "finish_reason": choice["finish_reason"],
138
+ "retained_bytes": cache.nbytes, "retained_limit": cache.max_bytes,
139
+ "retained_sequences": len(cache), "peak_mlx_bytes": mx.get_peak_memory()}
140
+ report["requests"].append(row)
141
+ save()
142
+ print(json.dumps(row), flush=True)
143
+ messages.extend([{"role": "assistant", "content": content},
144
+ {"role": "user", "content": "Reply with READY again."}])
145
+ hits = [row["usage"]["prompt_tokens_details"]["cached_tokens"]
146
+ for row in report["requests"][1:]]
147
+ report["warm_reuse_observed"] = any(hits)
148
+ report["warm_reuse_fraction"] = [
149
+ row["usage"]["prompt_tokens_details"]["cached_tokens"] / row["usage"]["prompt_tokens"]
150
+ for row in report["requests"][1:]]
151
+ assert all(fraction >= 0.5 for fraction in report["warm_reuse_fraction"]), \
152
+ "Requests completed, but the cache did not reuse most of the long history"
153
+ report["status"] = "PASS_FULL_MODEL_SYNTHETIC_LONG_CONTEXT_FOLLOWUPS"
154
+ save()
155
+ finally:
156
+ stopped.set()
157
+ observer.join(timeout=15)
158
+ except BaseException as error:
159
+ report["status"] = "FAIL"
160
+ report["error"] = f"{type(error).__name__}: {error}"
161
+ save()
162
+ raise
163
+
164
+
165
+ if __name__ == "__main__":
166
+ main()