ukisai commited on
Commit
ff55db2
·
verified ·
1 Parent(s): 4a7ea58

Add server settings for limited memory

Browse files

Document existing public MLX-LM cache and concurrency options and their synthetic multi-turn checks. No third-party log details are included. Model files and runtime code remain unchanged; full-model capacity is not established by these tests.

Files changed (3) hide show
  1. TROUBLESHOOTING.md +31 -1
  2. UPLOAD_MANIFEST.json +7 -7
  3. USAGE.md +30 -0
TROUBLESHOOTING.md CHANGED
@@ -43,6 +43,36 @@ separate runtime. GUI integration has not been validated for these full checkpoi
43
  Do not remove vision/MTP parameters or change the config to force a stock text
44
  loader to accept the file. All 2,379 saved tensor entries belong to this release.
45
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
46
  ## 3. If a server reports `404 generation thread died`
47
 
48
  In the pinned official MLX-LM server, this message means its generation worker
@@ -75,5 +105,5 @@ against a dead worker will not recover it; restart only after addressing the cau
75
  a tool definition. These are architecture/server checks, not full Swift generation
76
  or a test of the reported third-party plugin.
77
 
78
- The reporter's original traceback is still needed to identify their specific failure.
79
  These checks do not establish full-model Mac execution or GUI compatibility.
 
 
43
  Do not remove vision/MTP parameters or change the config to force a stock text
44
  loader to accept the file. All 2,379 saved tensor entries belong to this release.
45
 
46
+ ## Server settings for limited memory
47
+
48
+ The pinned MLX-LM server retains prompt caches for reuse by default. For a setup
49
+ that favors lower memory use over reuse speed, stop the existing server and start
50
+ it from the same patched Python environment with:
51
+
52
+ ```bash
53
+ mlx_lm.server --model ukisai/Swift-1.5-5bit-MLX \
54
+ --host 127.0.0.1 --port 8080 \
55
+ --prompt-cache-size 0 \
56
+ --prompt-concurrency 1 --decode-concurrency 1 \
57
+ --prefill-step-size 512
58
+ ```
59
+
60
+ This disables retained prompt caches, limits concurrency to one, and processes
61
+ prefill in smaller chunks. The model weights and their quantization are unchanged.
62
+ The active request still needs its own cache and temporary memory; these options
63
+ do not set a total process-memory cap or guarantee that any context length fits.
64
+
65
+ Disabling reuse can slow later turns because their history must be processed
66
+ again. It does not remove history that the client includes in the next request.
67
+ Start with a fresh short conversation and increase history while monitoring RAM.
68
+ The model's configured context limit is distinct from the memory needed to run it.
69
+
70
+ Validation scope: three successive HTTP requests, including a streaming request,
71
+ passed on each small synthetic 4-bit and 5-bit architecture with these settings.
72
+ Retained-cache accounting stayed at zero. These tests did not load the complete
73
+ Swift checkpoint or establish full-model long-context capacity on any Mac.
74
+ No macOS or MLX memory-limit overrides were used in those tests.
75
+
76
  ## 3. If a server reports `404 generation thread died`
77
 
78
  In the pinned official MLX-LM server, this message means its generation worker
 
105
  a tool definition. These are architecture/server checks, not full Swift generation
106
  or a test of the reported third-party plugin.
107
 
 
108
  These checks do not establish full-model Mac execution or GUI compatibility.
109
+ Diagnosing an individual crash requires its original exception and runtime configuration.
UPLOAD_MANIFEST.json CHANGED
@@ -38,15 +38,15 @@
38
  },
39
  {
40
  "path": "TROUBLESHOOTING.md",
41
- "bytes": 4380,
42
- "sha256": "32b67c70616257a0e26e91bbc6b1b880391872f98c4292aece9488940cda9d00",
43
- "git_blob_sha1": "2968859f3f32fc644fc20c2fb9e9644e10535e43"
44
  },
45
  {
46
  "path": "USAGE.md",
47
- "bytes": 5431,
48
- "sha256": "be75f4abc1868a29866013d9fd6490644f6c06cf3effe6e2704f3e6657affddb",
49
- "git_blob_sha1": "db137784896c403a7ee030bb56f335503227baf6"
50
  },
51
  {
52
  "path": "chat_template.jinja",
@@ -217,6 +217,6 @@
217
  }
218
  ],
219
  "file_count": 37,
220
- "total_bytes": 19310877524,
221
  "self_excluded": true
222
  }
 
38
  },
39
  {
40
  "path": "TROUBLESHOOTING.md",
41
+ "bytes": 5865,
42
+ "sha256": "d94c7c741398c55b8f9abaa1f1542d7046dc6fab1df7e9ea4b13b59af3945dee",
43
+ "git_blob_sha1": "3af3c505aff07e4e527e1d2d01b7b8f4cdb5288f"
44
  },
45
  {
46
  "path": "USAGE.md",
47
+ "bytes": 6912,
48
+ "sha256": "5c401d9f19fdd22c7bdc495bfdc925bc4ef6f9f829827a22918e26e0ba67d0e4",
49
+ "git_blob_sha1": "8532eaa9e4c661635164a08e3fb00c597191b1a0"
50
  },
51
  {
52
  "path": "chat_template.jinja",
 
217
  }
218
  ],
219
  "file_count": 37,
220
+ "total_bytes": 19310880490,
221
  "self_excluded": true
222
  }
USAGE.md CHANGED
@@ -102,6 +102,36 @@ The explicit `model.mtp_logits` step and `model.visual` encoder have historical
102
  component evidence. Integrated image/video chat and speculative generation are
103
  not implemented; unsupported multimodal generation must not be reported as working.
104
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
105
  ## Conversion provenance
106
 
107
  Conversion is not part of installation. If separately authorized, use only the
 
102
  component evidence. Integrated image/video chat and speculative generation are
103
  not implemented; unsupported multimodal generation must not be reported as working.
104
 
105
+ ## Server settings for limited memory
106
+
107
+ The pinned MLX-LM server retains prompt caches for reuse by default. For a setup
108
+ that favors lower memory use over reuse speed, stop the existing server and start
109
+ it from the same patched Python environment with:
110
+
111
+ ```bash
112
+ mlx_lm.server --model ukisai/Swift-1.5-5bit-MLX \
113
+ --host 127.0.0.1 --port 8080 \
114
+ --prompt-cache-size 0 \
115
+ --prompt-concurrency 1 --decode-concurrency 1 \
116
+ --prefill-step-size 512
117
+ ```
118
+
119
+ This disables retained prompt caches, limits concurrency to one, and processes
120
+ prefill in smaller chunks. The model weights and their quantization are unchanged.
121
+ The active request still needs its own cache and temporary memory; these options
122
+ do not set a total process-memory cap or guarantee that any context length fits.
123
+
124
+ Disabling reuse can slow later turns because their history must be processed
125
+ again. It does not remove history that the client includes in the next request.
126
+ Start with a fresh short conversation and increase history while monitoring RAM.
127
+ The model's configured context limit is distinct from the memory needed to run it.
128
+
129
+ Validation scope: three successive HTTP requests, including a streaming request,
130
+ passed on each small synthetic 4-bit and 5-bit architecture with these settings.
131
+ Retained-cache accounting stayed at zero. These tests did not load the complete
132
+ Swift checkpoint or establish full-model long-context capacity on any Mac.
133
+ No macOS or MLX memory-limit overrides were used in those tests.
134
+
135
  ## Conversion provenance
136
 
137
  Conversion is not part of installation. If separately authorized, use only the