ukisai commited on
Commit
8744db9
·
verified ·
1 Parent(s): 3e6f20e

Add server settings for limited memory

Browse files

Document existing public MLX-LM cache and concurrency options and their synthetic multi-turn checks. No third-party log details are included. Model files and runtime code remain unchanged; full-model capacity is not established by these tests.

Files changed (3) hide show
  1. TROUBLESHOOTING.md +31 -1
  2. UPLOAD_MANIFEST.json +7 -6
  3. USAGE.md +30 -0
TROUBLESHOOTING.md CHANGED
@@ -43,6 +43,36 @@ separate runtime. GUI integration has not been validated for these full checkpoi
43
  Do not remove vision/MTP parameters or change the config to force a stock text
44
  loader to accept the file. All 2,379 saved tensor entries belong to this release.
45
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
46
  ## 3. If a server reports `404 generation thread died`
47
 
48
  In the pinned official MLX-LM server, this message means its generation worker
@@ -75,5 +105,5 @@ against a dead worker will not recover it; restart only after addressing the cau
75
  a tool definition. These are architecture/server checks, not full Swift generation
76
  or a test of the reported third-party plugin.
77
 
78
- The reporter's original traceback is still needed to identify their specific failure.
79
  These checks do not establish full-model Mac execution or GUI compatibility.
 
 
43
  Do not remove vision/MTP parameters or change the config to force a stock text
44
  loader to accept the file. All 2,379 saved tensor entries belong to this release.
45
 
46
+ ## Server settings for limited memory
47
+
48
+ The pinned MLX-LM server retains prompt caches for reuse by default. For a setup
49
+ that favors lower memory use over reuse speed, stop the existing server and start
50
+ it from the same patched Python environment with:
51
+
52
+ ```bash
53
+ mlx_lm.server --model ukisai/Swift-1.5-4bit-MLX \
54
+ --host 127.0.0.1 --port 8080 \
55
+ --prompt-cache-size 0 \
56
+ --prompt-concurrency 1 --decode-concurrency 1 \
57
+ --prefill-step-size 512
58
+ ```
59
+
60
+ This disables retained prompt caches, limits concurrency to one, and processes
61
+ prefill in smaller chunks. The model weights and their quantization are unchanged.
62
+ The active request still needs its own cache and temporary memory; these options
63
+ do not set a total process-memory cap or guarantee that any context length fits.
64
+
65
+ Disabling reuse can slow later turns because their history must be processed
66
+ again. It does not remove history that the client includes in the next request.
67
+ Start with a fresh short conversation and increase history while monitoring RAM.
68
+ The model's configured context limit is distinct from the memory needed to run it.
69
+
70
+ Validation scope: three successive HTTP requests, including a streaming request,
71
+ passed on each small synthetic 4-bit and 5-bit architecture with these settings.
72
+ Retained-cache accounting stayed at zero. These tests did not load the complete
73
+ Swift checkpoint or establish full-model long-context capacity on any Mac.
74
+ No macOS or MLX memory-limit overrides were used in those tests.
75
+
76
  ## 3. If a server reports `404 generation thread died`
77
 
78
  In the pinned official MLX-LM server, this message means its generation worker
 
105
  a tool definition. These are architecture/server checks, not full Swift generation
106
  or a test of the reported third-party plugin.
107
 
 
108
  These checks do not establish full-model Mac execution or GUI compatibility.
109
+ Diagnosing an individual crash requires its original exception and runtime configuration.
UPLOAD_MANIFEST.json CHANGED
@@ -36,13 +36,13 @@
36
  },
37
  {
38
  "path": "TROUBLESHOOTING.md",
39
- "bytes": 4360,
40
- "sha256": "6460abf4aa4d2de46f07ce8dfa095476467eb466c6250c727ba0d6f2c77399eb"
41
  },
42
  {
43
  "path": "USAGE.md",
44
- "bytes": 5924,
45
- "sha256": "70e4b505d7acb75d33774e19b4b6202e5e45f0dcaf0f5973d6ea73f5d618a941"
46
  },
47
  {
48
  "path": "chat_template.jinja",
@@ -280,8 +280,9 @@
280
  "sha256": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003"
281
  }
282
  ],
283
- "total_bytes_excluding_manifest": 15859840103,
284
  "manifest_self_excluded": true,
285
  "packaging_update": "2026-09-24: diagnostic fixtures archived; model weights unchanged",
286
- "diagnostic_update": "2026-09-25: clarify required runtime; add download/runtime checker; weights unchanged"
 
287
  }
 
36
  },
37
  {
38
  "path": "TROUBLESHOOTING.md",
39
+ "bytes": 5845,
40
+ "sha256": "461ca48d8ec617a226f5c1d726c41bf5dfba2121fe8e7c4ad389d39fb1ccec4c"
41
  },
42
  {
43
  "path": "USAGE.md",
44
+ "bytes": 7405,
45
+ "sha256": "51f0a3c01149644dfc17db8b08c4b02c4de485e871a13c53e361165718711618"
46
  },
47
  {
48
  "path": "chat_template.jinja",
 
280
  "sha256": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003"
281
  }
282
  ],
283
+ "total_bytes_excluding_manifest": 15859843069,
284
  "manifest_self_excluded": true,
285
  "packaging_update": "2026-09-24: diagnostic fixtures archived; model weights unchanged",
286
+ "diagnostic_update": "2026-09-25: clarify required runtime; add download/runtime checker; weights unchanged",
287
+ "cache_settings_update": "2026-09-25: document public server options and synthetic multi-turn tests; model unchanged"
288
  }
USAGE.md CHANGED
@@ -100,6 +100,36 @@ encoder (`model.visual`). Their weights are retained and the release validation
100
  records their component execution. Speculative generation and integrated image/video
101
  chat are not implemented. Unsupported multimodal generation calls raise an error.
102
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
103
  Reproduce conversion only from the complete original Swift BF16 export identified
104
  in `QUANTIZATION_MANIFEST.json`, after verifying its 18 shards and original assets:
105
 
 
100
  records their component execution. Speculative generation and integrated image/video
101
  chat are not implemented. Unsupported multimodal generation calls raise an error.
102
 
103
+ ## Server settings for limited memory
104
+
105
+ The pinned MLX-LM server retains prompt caches for reuse by default. For a setup
106
+ that favors lower memory use over reuse speed, stop the existing server and start
107
+ it from the same patched Python environment with:
108
+
109
+ ```bash
110
+ mlx_lm.server --model ukisai/Swift-1.5-4bit-MLX \
111
+ --host 127.0.0.1 --port 8080 \
112
+ --prompt-cache-size 0 \
113
+ --prompt-concurrency 1 --decode-concurrency 1 \
114
+ --prefill-step-size 512
115
+ ```
116
+
117
+ This disables retained prompt caches, limits concurrency to one, and processes
118
+ prefill in smaller chunks. The model weights and their quantization are unchanged.
119
+ The active request still needs its own cache and temporary memory; these options
120
+ do not set a total process-memory cap or guarantee that any context length fits.
121
+
122
+ Disabling reuse can slow later turns because their history must be processed
123
+ again. It does not remove history that the client includes in the next request.
124
+ Start with a fresh short conversation and increase history while monitoring RAM.
125
+ The model's configured context limit is distinct from the memory needed to run it.
126
+
127
+ Validation scope: three successive HTTP requests, including a streaming request,
128
+ passed on each small synthetic 4-bit and 5-bit architecture with these settings.
129
+ Retained-cache accounting stayed at zero. These tests did not load the complete
130
+ Swift checkpoint or establish full-model long-context capacity on any Mac.
131
+ No macOS or MLX memory-limit overrides were used in those tests.
132
+
133
  Reproduce conversion only from the complete original Swift BF16 export identified
134
  in `QUANTIZATION_MANIFEST.json`, after verifying its 18 shards and original assets:
135