Instructions to use ukisai/Swift-1.5-5bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ukisai/Swift-1.5-5bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ukisai/Swift-1.5-5bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ukisai/Swift-1.5-5bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-1.5-5bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ukisai/Swift-1.5-5bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ukisai/Swift-1.5-5bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-1.5-5bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ukisai/Swift-1.5-5bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-1.5-5bit-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-1.5-5bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-1.5-5bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Bound retained server cache while keeping prompt reuse enabled
Browse filesThe pinned server constructed its retained-cache LRU without the configured byte limit. Entries saved during prefill or completion could therefore exceed `--prompt-cache-bytes` after the earlier request-admission trim. This patch enforces the retention limit on each insertion and keeps prompt reuse enabled with bounded defaults.
The Swift server profile defaults to two retained entries, one prompt and one decode stream, and 512-token prefill steps. A Metal memory estimate sets the default retained-cache budget; an explicit byte limit overrides it. The budget excludes active-request and total process memory, so this is not a promise that every context length fits.
Validation completed on clean copies of the pinned source for both architecture patch stacks:
- PASS: 44 unit/upstream server tests per patch stack.
- PASS: 36 offline HTTP requests per format, including warm cache reuse, explicit byte-cap enforcement, streaming, sequential generation and disabled retention.
- PASS: clean patch application, runtime source identity, unchanged model implementation, reverse-apply check, Black 25.1.0, isort 6.0.0 and Ruff 0.16.6.
These are reproducible local Apple Silicon checks, not hosted CI statuses. Results and the offline regression harness are included under `compatibility/cache-tests/`. They use tiny random model fixtures; full Swift long-context capacity remains unverified by this change.
Only runtime patch, tests, instructions and file manifest change. Checkpoint weights, quantization, architecture code and tokenizer assets are unchanged. Existing installations must apply the patch and restart; fetching model files alone does not update an installed server. `SERVER_CACHE_UPDATE.md` includes installation, offline startup, validation and rollback instructions.
This PR is for review and has not been merged into main.
- SERVER_CACHE_UPDATE.md +93 -0
- TROUBLESHOOTING.md +22 -29
- UPLOAD_MANIFEST.json +47 -7
- USAGE.md +21 -26
- compatibility/cache-tests/offline-4bit-results.json +222 -0
- compatibility/cache-tests/offline-5bit-results.json +222 -0
- compatibility/cache-tests/synthetic-config.json +71 -0
- compatibility/cache-tests/test_server_cache_http.py +271 -0
- compatibility/cache-tests/validation.json +55 -0
- compatibility/cache-tests/verify_server_patch.py +11 -0
- compatibility/swift15-server-cache.patch +253 -0
|
@@ -0,0 +1,93 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Server cache update
|
| 2 |
+
|
| 3 |
+
This update changes the pinned MLX-LM server, not the checkpoint. It is a Swift
|
| 4 |
+
runtime patch based on official MLX-LM `c69d1288440a0dc4e6401fc417098b07598dccd5`, not an upstream release.
|
| 5 |
+
The existing architecture patch remains required.
|
| 6 |
+
|
| 7 |
+
## Apply to an existing installation
|
| 8 |
+
|
| 9 |
+
Stop the server. Activate the same Python environment used for serving. From the
|
| 10 |
+
directory containing `swift5-mlx-lm` and `Swift-1.5-5bit-MLX`, after downloading this patch:
|
| 11 |
+
|
| 12 |
+
```bash
|
| 13 |
+
git -C swift5-mlx-lm rev-parse HEAD
|
| 14 |
+
git -C swift5-mlx-lm apply --check ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
|
| 15 |
+
git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
|
| 16 |
+
python -m pip install --no-deps -e ./swift5-mlx-lm
|
| 17 |
+
python Swift-1.5-5bit-MLX/compatibility/cache-tests/verify_server_patch.py
|
| 18 |
+
```
|
| 19 |
+
|
| 20 |
+
The revision must equal the pinned revision above. Stop if the patch check fails;
|
| 21 |
+
do not force it onto different server code or apply it twice. For a fresh install,
|
| 22 |
+
follow [USAGE.md](USAGE.md), which applies the architecture patches first.
|
| 23 |
+
The server verifier checks the active Python installation without loading weights.
|
| 24 |
+
|
| 25 |
+
Once dependencies and the complete checkpoint are already local, startup works
|
| 26 |
+
without Hub access:
|
| 27 |
+
|
| 28 |
+
```bash
|
| 29 |
+
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python -m mlx_lm.server \
|
| 30 |
+
--model ./Swift-1.5-5bit-MLX --host 127.0.0.1 --port 8080
|
| 31 |
+
```
|
| 32 |
+
|
| 33 |
+
Remove an earlier `--prompt-cache-size 0` override to enable reuse. Explicit
|
| 34 |
+
arguments override the new defaults. Installing/downloading prerequisites still
|
| 35 |
+
requires internet; offline operation requires those prerequisites already present.
|
| 36 |
+
|
| 37 |
+
## Behavior and limits
|
| 38 |
+
|
| 39 |
+
Defaults: two retained cache entries, prompt concurrency 1, decode concurrency 1,
|
| 40 |
+
and prefill steps of 512 tokens. The byte cap is now passed to the LRU cache, so
|
| 41 |
+
every insertion enforces it. Previously the CLI byte limit only triggered a trim
|
| 42 |
+
when adding a batched request; later retained entries could exceed that setting.
|
| 43 |
+
|
| 44 |
+
The automatic Metal budget uses the reported recommended working-set size `R`,
|
| 45 |
+
the loaded model parameter bytes `W`, and workspace reserve `max(4 GiB, R/8)`.
|
| 46 |
+
It budgets half the remaining space for retained entries:
|
| 47 |
+
`max(0, (R - W - reserve) / 2)`. This is a conservative estimate, not a guarantee
|
| 48 |
+
that an arbitrary active context fits. It does not change macOS or MLX limits.
|
| 49 |
+
When device metadata is unavailable, the entry limit still applies; supply an
|
| 50 |
+
explicit byte limit if needed. `--prompt-cache-bytes 2GB`, for example, sets a
|
| 51 |
+
retention budget of 2,000,000,000 bytes; it is not a universal recommended value.
|
| 52 |
+
|
| 53 |
+
The active request, temporary tensors, allocator and other apps use additional
|
| 54 |
+
memory. Cache byte accounting is logical and may include shared buffers. Entries
|
| 55 |
+
that cannot fit are evicted whole; the client's input history is not truncated.
|
| 56 |
+
Reuse can therefore vary with context size. This patch does not add KV
|
| 57 |
+
quantization, change attention or alter any checkpoint tensors.
|
| 58 |
+
|
| 59 |
+
## Reproduce the checks
|
| 60 |
+
|
| 61 |
+
Use the patched environment with Python 3.12, MLX 0.32.2 and MLX-LM 0.32.0.
|
| 62 |
+
`validation.json` and the two `offline-*-results.json` files record our local
|
| 63 |
+
results. No hosted CI status is claimed. The tests use random tiny models and the
|
| 64 |
+
local tokenizer; they never open the released Swift weight shards.
|
| 65 |
+
|
| 66 |
+
```bash
|
| 67 |
+
python -m pip install 'pytest==9.1.1' 'requests==2.32.5'
|
| 68 |
+
python -m pytest -q swift5-mlx-lm/tests/test_server_cache_budget.py swift5-mlx-lm/tests/test_server.py
|
| 69 |
+
python Swift-1.5-5bit-MLX/compatibility/cache-tests/test_server_cache_http.py \
|
| 70 |
+
--assets ./Swift-1.5-5bit-MLX \
|
| 71 |
+
--config Swift-1.5-5bit-MLX/compatibility/cache-tests/synthetic-config.json \
|
| 72 |
+
--bits 5 --output cache-http-results.json
|
| 73 |
+
```
|
| 74 |
+
|
| 75 |
+
The upstream server suite may fetch its small upstream test model if not already
|
| 76 |
+
cached. The separate HTTP harness sets offline mode and needs only local assets.
|
| 77 |
+
It covers 12 requests each with automatic budgeting, an explicit byte limit, and
|
| 78 |
+
disabled retention; streaming and seeded sequential execution are included.
|
| 79 |
+
Synthetic output is deliberately fixed for transport/cache checks, not quality.
|
| 80 |
+
Full Swift long-context execution on a larger Mac has not been independently
|
| 81 |
+
retested by this change. GUI runtimes need their own integration.
|
| 82 |
+
|
| 83 |
+
## Roll back this runtime patch
|
| 84 |
+
|
| 85 |
+
Stop the server, then reverse only this patch and restart with explicit settings:
|
| 86 |
+
|
| 87 |
+
```bash
|
| 88 |
+
git -C swift5-mlx-lm apply --reverse --check ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
|
| 89 |
+
git -C swift5-mlx-lm apply --reverse ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
|
| 90 |
+
```
|
| 91 |
+
|
| 92 |
+
This leaves the architecture patch and model files intact. An editable install
|
| 93 |
+
uses the source change after restart; reinstall if a non-editable copy was used.
|
|
@@ -43,35 +43,28 @@ separate runtime. GUI integration has not been validated for these full checkpoi
|
|
| 43 |
Do not remove vision/MTP parameters or change the config to force a stock text
|
| 44 |
loader to accept the file. All 2,379 saved tensor entries belong to this release.
|
| 45 |
|
| 46 |
-
## Server
|
| 47 |
-
|
| 48 |
-
The
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
```
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
The model's configured context limit is distinct from the memory needed to run it.
|
| 69 |
-
|
| 70 |
-
Validation scope: three successive HTTP requests, including a streaming request,
|
| 71 |
-
passed on each small synthetic 4-bit and 5-bit architecture with these settings.
|
| 72 |
-
Retained-cache accounting stayed at zero. These tests did not load the complete
|
| 73 |
-
Swift checkpoint or establish full-model long-context capacity on any Mac.
|
| 74 |
-
No macOS or MLX memory-limit overrides were used in those tests.
|
| 75 |
|
| 76 |
## 3. If a server reports `404 generation thread died`
|
| 77 |
|
|
|
|
| 43 |
Do not remove vision/MTP parameters or change the config to force a stock text
|
| 44 |
loader to accept the file. All 2,379 saved tensor entries belong to this release.
|
| 45 |
|
| 46 |
+
## Server cache update
|
| 47 |
+
|
| 48 |
+
The supplied server patch keeps prompt reuse enabled with at most two retained
|
| 49 |
+
entries, enforces a retained-cache byte budget at insertion, and defaults to one
|
| 50 |
+
prompt and one decode stream with 512-token prefill steps. On Metal the default
|
| 51 |
+
budget estimates headroom after model weights and workspace; an explicit
|
| 52 |
+
`--prompt-cache-bytes` overrides it. `--prompt-cache-size 2` means two entries,
|
| 53 |
+
not two GB.
|
| 54 |
+
|
| 55 |
+
See [SERVER_CACHE_UPDATE.md](SERVER_CACHE_UPDATE.md) for installation, offline
|
| 56 |
+
startup, validation and rollback. Existing installations must apply this runtime
|
| 57 |
+
update and restart. No model-weight download or re-quantization is needed.
|
| 58 |
+
|
| 59 |
+
The retained-cache limit does not cap active-request or total process memory.
|
| 60 |
+
An entry larger than the budget is evicted, so reuse depends on what fits. Use
|
| 61 |
+
`--prompt-cache-size 0` only when intentionally disabling reuse. The model's
|
| 62 |
+
configured context length is not a promise that it fits every Mac.
|
| 63 |
+
|
| 64 |
+
Validation: 44 unit/upstream server tests and 72 offline HTTP requests passed
|
| 65 |
+
across tiny 4-bit and 5-bit fixtures. These checks cover retention, warm reuse,
|
| 66 |
+
streaming and sequential generation. They do not establish full-model
|
| 67 |
+
long-context capacity or GUI integration.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 68 |
|
| 69 |
## 3. If a server reports `404 generation thread died`
|
| 70 |
|
|
@@ -36,17 +36,20 @@
|
|
| 36 |
"sha256": "e0ce3327646a3dde90ee001a3bb819c8cb24e550a83d28340671f2542e7ad55b",
|
| 37 |
"git_blob_sha1": "008e4ab150bc188e372cad12fda89317afda850e"
|
| 38 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
{
|
| 40 |
"path": "TROUBLESHOOTING.md",
|
| 41 |
-
"bytes":
|
| 42 |
-
"sha256": "
|
| 43 |
-
"git_blob_sha1": "3af3c505aff07e4e527e1d2d01b7b8f4cdb5288f"
|
| 44 |
},
|
| 45 |
{
|
| 46 |
"path": "USAGE.md",
|
| 47 |
-
"bytes":
|
| 48 |
-
"sha256": "
|
| 49 |
-
"git_blob_sha1": "8532eaa9e4c661635164a08e3fb00c597191b1a0"
|
| 50 |
},
|
| 51 |
{
|
| 52 |
"path": "chat_template.jinja",
|
|
@@ -66,6 +69,36 @@
|
|
| 66 |
"sha256": "ccfab7ccb2ea306f71531c8ca77bb55507606cd90768b1e32b8b52ab5b48cf01",
|
| 67 |
"git_blob_sha1": "98ff47b9ef9d4ac9f1a4bde3db13dd27456c5ea1"
|
| 68 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 69 |
{
|
| 70 |
"path": "compatibility/conversion-result.json",
|
| 71 |
"bytes": 694,
|
|
@@ -126,6 +159,11 @@
|
|
| 126 |
"sha256": "f6f1d0bdafa45863bfbf93dac0398c481c993ea04fdf38b9bae98c643f89eaec",
|
| 127 |
"git_blob_sha1": "b739acc2e3b4cb82f073b7280255bcb898bb2a33"
|
| 128 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 129 |
{
|
| 130 |
"path": "config.json",
|
| 131 |
"bytes": 4577,
|
|
@@ -218,5 +256,7 @@
|
|
| 218 |
],
|
| 219 |
"file_count": 37,
|
| 220 |
"total_bytes": 19310880490,
|
| 221 |
-
"self_excluded": true
|
|
|
|
|
|
|
| 222 |
}
|
|
|
|
| 36 |
"sha256": "e0ce3327646a3dde90ee001a3bb819c8cb24e550a83d28340671f2542e7ad55b",
|
| 37 |
"git_blob_sha1": "008e4ab150bc188e372cad12fda89317afda850e"
|
| 38 |
},
|
| 39 |
+
{
|
| 40 |
+
"path": "SERVER_CACHE_UPDATE.md",
|
| 41 |
+
"bytes": 4765,
|
| 42 |
+
"sha256": "31ae8115dcf0ccba9e91dd03f9eab30dc56a6a752fa26955c05555cc655bdcd7"
|
| 43 |
+
},
|
| 44 |
{
|
| 45 |
"path": "TROUBLESHOOTING.md",
|
| 46 |
+
"bytes": 5623,
|
| 47 |
+
"sha256": "8d631dcdacfd6f98d0ac37c4d2532f1a8139ab2e336c8ae27bfe03c251444327"
|
|
|
|
| 48 |
},
|
| 49 |
{
|
| 50 |
"path": "USAGE.md",
|
| 51 |
+
"bytes": 6858,
|
| 52 |
+
"sha256": "946ba223dd2a7cf1aceb069cf4fd82f4bf609366abacba61017f08e2bf9a29b4"
|
|
|
|
| 53 |
},
|
| 54 |
{
|
| 55 |
"path": "chat_template.jinja",
|
|
|
|
| 69 |
"sha256": "ccfab7ccb2ea306f71531c8ca77bb55507606cd90768b1e32b8b52ab5b48cf01",
|
| 70 |
"git_blob_sha1": "98ff47b9ef9d4ac9f1a4bde3db13dd27456c5ea1"
|
| 71 |
},
|
| 72 |
+
{
|
| 73 |
+
"path": "compatibility/cache-tests/offline-4bit-results.json",
|
| 74 |
+
"bytes": 5014,
|
| 75 |
+
"sha256": "ddcfa5f200aa57d1cdadc6225299329b592e858bea2502e9bf07d313f4cf149a"
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"path": "compatibility/cache-tests/offline-5bit-results.json",
|
| 79 |
+
"bytes": 5014,
|
| 80 |
+
"sha256": "f50ca1c21eafc43623f1461999e33e99b9e684a990d926aa77fbe59cb919fe25"
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"path": "compatibility/cache-tests/synthetic-config.json",
|
| 84 |
+
"bytes": 1814,
|
| 85 |
+
"sha256": "852fa793054753245559ffeae66d4367b6009f3ebdb7ca611747ec7564d785c7"
|
| 86 |
+
},
|
| 87 |
+
{
|
| 88 |
+
"path": "compatibility/cache-tests/test_server_cache_http.py",
|
| 89 |
+
"bytes": 9689,
|
| 90 |
+
"sha256": "e7aff50b441145338575a50c7abc5fb6d40fdac3d11bbf6433d909ecfe22483b"
|
| 91 |
+
},
|
| 92 |
+
{
|
| 93 |
+
"path": "compatibility/cache-tests/validation.json",
|
| 94 |
+
"bytes": 1984,
|
| 95 |
+
"sha256": "288450e5a3b34d4c45ce16489de749f49c7bd03cd2d28d4940d4bd94953022a5"
|
| 96 |
+
},
|
| 97 |
+
{
|
| 98 |
+
"path": "compatibility/cache-tests/verify_server_patch.py",
|
| 99 |
+
"bytes": 504,
|
| 100 |
+
"sha256": "4c3acf33b4b9ffe0131c45ad0584785767a81cf455f9f60d9ec814e4b7eaf61a"
|
| 101 |
+
},
|
| 102 |
{
|
| 103 |
"path": "compatibility/conversion-result.json",
|
| 104 |
"bytes": 694,
|
|
|
|
| 159 |
"sha256": "f6f1d0bdafa45863bfbf93dac0398c481c993ea04fdf38b9bae98c643f89eaec",
|
| 160 |
"git_blob_sha1": "b739acc2e3b4cb82f073b7280255bcb898bb2a33"
|
| 161 |
},
|
| 162 |
+
{
|
| 163 |
+
"path": "compatibility/swift15-server-cache.patch",
|
| 164 |
+
"bytes": 9356,
|
| 165 |
+
"sha256": "a7fbc0f0524ee7864d9f41a98a2e35acf0e63f53e369ee5b9a87ca12926a9eab"
|
| 166 |
+
},
|
| 167 |
{
|
| 168 |
"path": "config.json",
|
| 169 |
"bytes": 4577,
|
|
|
|
| 256 |
],
|
| 257 |
"file_count": 37,
|
| 258 |
"total_bytes": 19310880490,
|
| 259 |
+
"self_excluded": true,
|
| 260 |
+
"total_bytes_excluding_manifest": 19310918334,
|
| 261 |
+
"server_cache_update": "2026-09-25: bounded cache reuse runtime patch and recorded local tests; weights unchanged"
|
| 262 |
}
|
|
@@ -38,6 +38,8 @@ git -C swift5-mlx-lm apply --check ../Swift-1.5-5bit-MLX/compatibility/swift15-m
|
|
| 38 |
git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/swift15-mlx-lm.patch
|
| 39 |
git -C swift5-mlx-lm apply --check ../Swift-1.5-5bit-MLX/compatibility/enable-5bit.patch
|
| 40 |
git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/enable-5bit.patch
|
|
|
|
|
|
|
| 41 |
```
|
| 42 |
|
| 43 |
Apple Silicon:
|
|
@@ -102,35 +104,28 @@ The explicit `model.mtp_logits` step and `model.visual` encoder have historical
|
|
| 102 |
component evidence. Integrated image/video chat and speculative generation are
|
| 103 |
not implemented; unsupported multimodal generation must not be reported as working.
|
| 104 |
|
| 105 |
-
## Server
|
| 106 |
|
| 107 |
-
The
|
| 108 |
-
|
| 109 |
-
|
|
|
|
|
|
|
|
|
|
| 110 |
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
|
| 115 |
-
|
| 116 |
-
|
| 117 |
-
``
|
|
|
|
| 118 |
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
Disabling reuse can slow later turns because their history must be processed
|
| 125 |
-
again. It does not remove history that the client includes in the next request.
|
| 126 |
-
Start with a fresh short conversation and increase history while monitoring RAM.
|
| 127 |
-
The model's configured context limit is distinct from the memory needed to run it.
|
| 128 |
-
|
| 129 |
-
Validation scope: three successive HTTP requests, including a streaming request,
|
| 130 |
-
passed on each small synthetic 4-bit and 5-bit architecture with these settings.
|
| 131 |
-
Retained-cache accounting stayed at zero. These tests did not load the complete
|
| 132 |
-
Swift checkpoint or establish full-model long-context capacity on any Mac.
|
| 133 |
-
No macOS or MLX memory-limit overrides were used in those tests.
|
| 134 |
|
| 135 |
## Conversion provenance
|
| 136 |
|
|
|
|
| 38 |
git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/swift15-mlx-lm.patch
|
| 39 |
git -C swift5-mlx-lm apply --check ../Swift-1.5-5bit-MLX/compatibility/enable-5bit.patch
|
| 40 |
git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/enable-5bit.patch
|
| 41 |
+
git -C swift5-mlx-lm apply --check ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
|
| 42 |
+
git -C swift5-mlx-lm apply ../Swift-1.5-5bit-MLX/compatibility/swift15-server-cache.patch
|
| 43 |
```
|
| 44 |
|
| 45 |
Apple Silicon:
|
|
|
|
| 104 |
component evidence. Integrated image/video chat and speculative generation are
|
| 105 |
not implemented; unsupported multimodal generation must not be reported as working.
|
| 106 |
|
| 107 |
+
## Server cache update
|
| 108 |
|
| 109 |
+
The supplied server patch keeps prompt reuse enabled with at most two retained
|
| 110 |
+
entries, enforces a retained-cache byte budget at insertion, and defaults to one
|
| 111 |
+
prompt and one decode stream with 512-token prefill steps. On Metal the default
|
| 112 |
+
budget estimates headroom after model weights and workspace; an explicit
|
| 113 |
+
`--prompt-cache-bytes` overrides it. `--prompt-cache-size 2` means two entries,
|
| 114 |
+
not two GB.
|
| 115 |
|
| 116 |
+
See [SERVER_CACHE_UPDATE.md](SERVER_CACHE_UPDATE.md) for installation, offline
|
| 117 |
+
startup, validation and rollback. Existing installations must apply this runtime
|
| 118 |
+
update and restart. No model-weight download or re-quantization is needed.
|
| 119 |
+
|
| 120 |
+
The retained-cache limit does not cap active-request or total process memory.
|
| 121 |
+
An entry larger than the budget is evicted, so reuse depends on what fits. Use
|
| 122 |
+
`--prompt-cache-size 0` only when intentionally disabling reuse. The model's
|
| 123 |
+
configured context length is not a promise that it fits every Mac.
|
| 124 |
|
| 125 |
+
Validation: 44 unit/upstream server tests and 72 offline HTTP requests passed
|
| 126 |
+
across tiny 4-bit and 5-bit fixtures. These checks cover retention, warm reuse,
|
| 127 |
+
streaming and sequential generation. They do not establish full-model
|
| 128 |
+
long-context capacity or GUI integration.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 129 |
|
| 130 |
## Conversion provenance
|
| 131 |
|
|
@@ -0,0 +1,222 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"scope": "Tiny random weights with real local tokenizer assets. No released model weights loaded or changed; no full-model capacity claim.",
|
| 3 |
+
"bits": 4,
|
| 4 |
+
"offline_hub": true,
|
| 5 |
+
"memory_limit_overrides": false,
|
| 6 |
+
"machine": "arm64",
|
| 7 |
+
"python": "3.12.13",
|
| 8 |
+
"packages": {
|
| 9 |
+
"mlx": "0.32.2",
|
| 10 |
+
"mlx-lm": "0.32.0",
|
| 11 |
+
"transformers": "5.14.1",
|
| 12 |
+
"huggingface_hub": "1.31.0"
|
| 13 |
+
},
|
| 14 |
+
"profiles": [
|
| 15 |
+
{
|
| 16 |
+
"name": "defaults",
|
| 17 |
+
"requests": [
|
| 18 |
+
{
|
| 19 |
+
"http_status": 200,
|
| 20 |
+
"prompt_tokens": 225,
|
| 21 |
+
"cached_tokens": 0
|
| 22 |
+
},
|
| 23 |
+
{
|
| 24 |
+
"http_status": 200,
|
| 25 |
+
"prompt_tokens": 246,
|
| 26 |
+
"cached_tokens": 221
|
| 27 |
+
},
|
| 28 |
+
{
|
| 29 |
+
"http_status": 200,
|
| 30 |
+
"prompt_tokens": 267,
|
| 31 |
+
"cached_tokens": 242
|
| 32 |
+
},
|
| 33 |
+
{
|
| 34 |
+
"http_status": 200,
|
| 35 |
+
"prompt_tokens": 288,
|
| 36 |
+
"cached_tokens": 263
|
| 37 |
+
},
|
| 38 |
+
{
|
| 39 |
+
"http_status": 200,
|
| 40 |
+
"prompt_tokens": 309,
|
| 41 |
+
"cached_tokens": 284
|
| 42 |
+
},
|
| 43 |
+
{
|
| 44 |
+
"http_status": 200,
|
| 45 |
+
"prompt_tokens": 330,
|
| 46 |
+
"cached_tokens": 305
|
| 47 |
+
},
|
| 48 |
+
{
|
| 49 |
+
"http_status": 200,
|
| 50 |
+
"prompt_tokens": 351,
|
| 51 |
+
"cached_tokens": 326
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"http_status": 200,
|
| 55 |
+
"prompt_tokens": 372,
|
| 56 |
+
"cached_tokens": 347
|
| 57 |
+
},
|
| 58 |
+
{
|
| 59 |
+
"http_status": 200,
|
| 60 |
+
"prompt_tokens": 393,
|
| 61 |
+
"cached_tokens": 368
|
| 62 |
+
},
|
| 63 |
+
{
|
| 64 |
+
"http_status": 200,
|
| 65 |
+
"prompt_tokens": 414,
|
| 66 |
+
"cached_tokens": 389
|
| 67 |
+
},
|
| 68 |
+
{
|
| 69 |
+
"http_status": 200,
|
| 70 |
+
"stream_complete": true
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"http_status": 200,
|
| 74 |
+
"prompt_tokens": 457,
|
| 75 |
+
"cached_tokens": 431
|
| 76 |
+
}
|
| 77 |
+
],
|
| 78 |
+
"byte_limit": 4190923688,
|
| 79 |
+
"peak_retained_bytes": 246528,
|
| 80 |
+
"retained_sequences": 2,
|
| 81 |
+
"insertions": 35
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "explicit_bytes",
|
| 85 |
+
"requests": [
|
| 86 |
+
{
|
| 87 |
+
"http_status": 200,
|
| 88 |
+
"prompt_tokens": 225,
|
| 89 |
+
"cached_tokens": 0
|
| 90 |
+
},
|
| 91 |
+
{
|
| 92 |
+
"http_status": 200,
|
| 93 |
+
"prompt_tokens": 246,
|
| 94 |
+
"cached_tokens": 112
|
| 95 |
+
},
|
| 96 |
+
{
|
| 97 |
+
"http_status": 200,
|
| 98 |
+
"prompt_tokens": 267,
|
| 99 |
+
"cached_tokens": 0
|
| 100 |
+
},
|
| 101 |
+
{
|
| 102 |
+
"http_status": 200,
|
| 103 |
+
"prompt_tokens": 288,
|
| 104 |
+
"cached_tokens": 112
|
| 105 |
+
},
|
| 106 |
+
{
|
| 107 |
+
"http_status": 200,
|
| 108 |
+
"prompt_tokens": 309,
|
| 109 |
+
"cached_tokens": 0
|
| 110 |
+
},
|
| 111 |
+
{
|
| 112 |
+
"http_status": 200,
|
| 113 |
+
"prompt_tokens": 330,
|
| 114 |
+
"cached_tokens": 112
|
| 115 |
+
},
|
| 116 |
+
{
|
| 117 |
+
"http_status": 200,
|
| 118 |
+
"prompt_tokens": 351,
|
| 119 |
+
"cached_tokens": 0
|
| 120 |
+
},
|
| 121 |
+
{
|
| 122 |
+
"http_status": 200,
|
| 123 |
+
"prompt_tokens": 372,
|
| 124 |
+
"cached_tokens": 112
|
| 125 |
+
},
|
| 126 |
+
{
|
| 127 |
+
"http_status": 200,
|
| 128 |
+
"prompt_tokens": 393,
|
| 129 |
+
"cached_tokens": 0
|
| 130 |
+
},
|
| 131 |
+
{
|
| 132 |
+
"http_status": 200,
|
| 133 |
+
"prompt_tokens": 414,
|
| 134 |
+
"cached_tokens": 112
|
| 135 |
+
},
|
| 136 |
+
{
|
| 137 |
+
"http_status": 200,
|
| 138 |
+
"stream_complete": true
|
| 139 |
+
},
|
| 140 |
+
{
|
| 141 |
+
"http_status": 200,
|
| 142 |
+
"prompt_tokens": 457,
|
| 143 |
+
"cached_tokens": 112
|
| 144 |
+
}
|
| 145 |
+
],
|
| 146 |
+
"byte_limit": 100000,
|
| 147 |
+
"peak_retained_bytes": 82432,
|
| 148 |
+
"retained_sequences": 1,
|
| 149 |
+
"insertions": 40
|
| 150 |
+
},
|
| 151 |
+
{
|
| 152 |
+
"name": "no_retention",
|
| 153 |
+
"requests": [
|
| 154 |
+
{
|
| 155 |
+
"http_status": 200,
|
| 156 |
+
"prompt_tokens": 225,
|
| 157 |
+
"cached_tokens": 0
|
| 158 |
+
},
|
| 159 |
+
{
|
| 160 |
+
"http_status": 200,
|
| 161 |
+
"prompt_tokens": 246,
|
| 162 |
+
"cached_tokens": 0
|
| 163 |
+
},
|
| 164 |
+
{
|
| 165 |
+
"http_status": 200,
|
| 166 |
+
"prompt_tokens": 267,
|
| 167 |
+
"cached_tokens": 0
|
| 168 |
+
},
|
| 169 |
+
{
|
| 170 |
+
"http_status": 200,
|
| 171 |
+
"prompt_tokens": 288,
|
| 172 |
+
"cached_tokens": 0
|
| 173 |
+
},
|
| 174 |
+
{
|
| 175 |
+
"http_status": 200,
|
| 176 |
+
"prompt_tokens": 309,
|
| 177 |
+
"cached_tokens": 0
|
| 178 |
+
},
|
| 179 |
+
{
|
| 180 |
+
"http_status": 200,
|
| 181 |
+
"prompt_tokens": 330,
|
| 182 |
+
"cached_tokens": 0
|
| 183 |
+
},
|
| 184 |
+
{
|
| 185 |
+
"http_status": 200,
|
| 186 |
+
"prompt_tokens": 351,
|
| 187 |
+
"cached_tokens": 0
|
| 188 |
+
},
|
| 189 |
+
{
|
| 190 |
+
"http_status": 200,
|
| 191 |
+
"prompt_tokens": 372,
|
| 192 |
+
"cached_tokens": 0
|
| 193 |
+
},
|
| 194 |
+
{
|
| 195 |
+
"http_status": 200,
|
| 196 |
+
"prompt_tokens": 393,
|
| 197 |
+
"cached_tokens": 0
|
| 198 |
+
},
|
| 199 |
+
{
|
| 200 |
+
"http_status": 200,
|
| 201 |
+
"prompt_tokens": 414,
|
| 202 |
+
"cached_tokens": 0
|
| 203 |
+
},
|
| 204 |
+
{
|
| 205 |
+
"http_status": 200,
|
| 206 |
+
"stream_complete": true
|
| 207 |
+
},
|
| 208 |
+
{
|
| 209 |
+
"http_status": 200,
|
| 210 |
+
"prompt_tokens": 457,
|
| 211 |
+
"cached_tokens": 0
|
| 212 |
+
}
|
| 213 |
+
],
|
| 214 |
+
"byte_limit": 4190923688,
|
| 215 |
+
"peak_retained_bytes": 0,
|
| 216 |
+
"retained_sequences": 0,
|
| 217 |
+
"insertions": 45
|
| 218 |
+
}
|
| 219 |
+
],
|
| 220 |
+
"peak_mlx_bytes": 338846824,
|
| 221 |
+
"status": "PASS"
|
| 222 |
+
}
|
|
@@ -0,0 +1,222 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"scope": "Tiny random weights with real local tokenizer assets. No released model weights loaded or changed; no full-model capacity claim.",
|
| 3 |
+
"bits": 5,
|
| 4 |
+
"offline_hub": true,
|
| 5 |
+
"memory_limit_overrides": false,
|
| 6 |
+
"machine": "arm64",
|
| 7 |
+
"python": "3.12.13",
|
| 8 |
+
"packages": {
|
| 9 |
+
"mlx": "0.32.2",
|
| 10 |
+
"mlx-lm": "0.32.0",
|
| 11 |
+
"transformers": "5.14.1",
|
| 12 |
+
"huggingface_hub": "1.31.0"
|
| 13 |
+
},
|
| 14 |
+
"profiles": [
|
| 15 |
+
{
|
| 16 |
+
"name": "defaults",
|
| 17 |
+
"requests": [
|
| 18 |
+
{
|
| 19 |
+
"http_status": 200,
|
| 20 |
+
"prompt_tokens": 225,
|
| 21 |
+
"cached_tokens": 0
|
| 22 |
+
},
|
| 23 |
+
{
|
| 24 |
+
"http_status": 200,
|
| 25 |
+
"prompt_tokens": 246,
|
| 26 |
+
"cached_tokens": 221
|
| 27 |
+
},
|
| 28 |
+
{
|
| 29 |
+
"http_status": 200,
|
| 30 |
+
"prompt_tokens": 267,
|
| 31 |
+
"cached_tokens": 242
|
| 32 |
+
},
|
| 33 |
+
{
|
| 34 |
+
"http_status": 200,
|
| 35 |
+
"prompt_tokens": 288,
|
| 36 |
+
"cached_tokens": 263
|
| 37 |
+
},
|
| 38 |
+
{
|
| 39 |
+
"http_status": 200,
|
| 40 |
+
"prompt_tokens": 309,
|
| 41 |
+
"cached_tokens": 284
|
| 42 |
+
},
|
| 43 |
+
{
|
| 44 |
+
"http_status": 200,
|
| 45 |
+
"prompt_tokens": 330,
|
| 46 |
+
"cached_tokens": 305
|
| 47 |
+
},
|
| 48 |
+
{
|
| 49 |
+
"http_status": 200,
|
| 50 |
+
"prompt_tokens": 351,
|
| 51 |
+
"cached_tokens": 326
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"http_status": 200,
|
| 55 |
+
"prompt_tokens": 372,
|
| 56 |
+
"cached_tokens": 347
|
| 57 |
+
},
|
| 58 |
+
{
|
| 59 |
+
"http_status": 200,
|
| 60 |
+
"prompt_tokens": 393,
|
| 61 |
+
"cached_tokens": 368
|
| 62 |
+
},
|
| 63 |
+
{
|
| 64 |
+
"http_status": 200,
|
| 65 |
+
"prompt_tokens": 414,
|
| 66 |
+
"cached_tokens": 389
|
| 67 |
+
},
|
| 68 |
+
{
|
| 69 |
+
"http_status": 200,
|
| 70 |
+
"stream_complete": true
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"http_status": 200,
|
| 74 |
+
"prompt_tokens": 457,
|
| 75 |
+
"cached_tokens": 431
|
| 76 |
+
}
|
| 77 |
+
],
|
| 78 |
+
"byte_limit": 4186895080,
|
| 79 |
+
"peak_retained_bytes": 246528,
|
| 80 |
+
"retained_sequences": 2,
|
| 81 |
+
"insertions": 35
|
| 82 |
+
},
|
| 83 |
+
{
|
| 84 |
+
"name": "explicit_bytes",
|
| 85 |
+
"requests": [
|
| 86 |
+
{
|
| 87 |
+
"http_status": 200,
|
| 88 |
+
"prompt_tokens": 225,
|
| 89 |
+
"cached_tokens": 0
|
| 90 |
+
},
|
| 91 |
+
{
|
| 92 |
+
"http_status": 200,
|
| 93 |
+
"prompt_tokens": 246,
|
| 94 |
+
"cached_tokens": 112
|
| 95 |
+
},
|
| 96 |
+
{
|
| 97 |
+
"http_status": 200,
|
| 98 |
+
"prompt_tokens": 267,
|
| 99 |
+
"cached_tokens": 0
|
| 100 |
+
},
|
| 101 |
+
{
|
| 102 |
+
"http_status": 200,
|
| 103 |
+
"prompt_tokens": 288,
|
| 104 |
+
"cached_tokens": 112
|
| 105 |
+
},
|
| 106 |
+
{
|
| 107 |
+
"http_status": 200,
|
| 108 |
+
"prompt_tokens": 309,
|
| 109 |
+
"cached_tokens": 0
|
| 110 |
+
},
|
| 111 |
+
{
|
| 112 |
+
"http_status": 200,
|
| 113 |
+
"prompt_tokens": 330,
|
| 114 |
+
"cached_tokens": 112
|
| 115 |
+
},
|
| 116 |
+
{
|
| 117 |
+
"http_status": 200,
|
| 118 |
+
"prompt_tokens": 351,
|
| 119 |
+
"cached_tokens": 0
|
| 120 |
+
},
|
| 121 |
+
{
|
| 122 |
+
"http_status": 200,
|
| 123 |
+
"prompt_tokens": 372,
|
| 124 |
+
"cached_tokens": 112
|
| 125 |
+
},
|
| 126 |
+
{
|
| 127 |
+
"http_status": 200,
|
| 128 |
+
"prompt_tokens": 393,
|
| 129 |
+
"cached_tokens": 0
|
| 130 |
+
},
|
| 131 |
+
{
|
| 132 |
+
"http_status": 200,
|
| 133 |
+
"prompt_tokens": 414,
|
| 134 |
+
"cached_tokens": 112
|
| 135 |
+
},
|
| 136 |
+
{
|
| 137 |
+
"http_status": 200,
|
| 138 |
+
"stream_complete": true
|
| 139 |
+
},
|
| 140 |
+
{
|
| 141 |
+
"http_status": 200,
|
| 142 |
+
"prompt_tokens": 457,
|
| 143 |
+
"cached_tokens": 112
|
| 144 |
+
}
|
| 145 |
+
],
|
| 146 |
+
"byte_limit": 100000,
|
| 147 |
+
"peak_retained_bytes": 82432,
|
| 148 |
+
"retained_sequences": 1,
|
| 149 |
+
"insertions": 40
|
| 150 |
+
},
|
| 151 |
+
{
|
| 152 |
+
"name": "no_retention",
|
| 153 |
+
"requests": [
|
| 154 |
+
{
|
| 155 |
+
"http_status": 200,
|
| 156 |
+
"prompt_tokens": 225,
|
| 157 |
+
"cached_tokens": 0
|
| 158 |
+
},
|
| 159 |
+
{
|
| 160 |
+
"http_status": 200,
|
| 161 |
+
"prompt_tokens": 246,
|
| 162 |
+
"cached_tokens": 0
|
| 163 |
+
},
|
| 164 |
+
{
|
| 165 |
+
"http_status": 200,
|
| 166 |
+
"prompt_tokens": 267,
|
| 167 |
+
"cached_tokens": 0
|
| 168 |
+
},
|
| 169 |
+
{
|
| 170 |
+
"http_status": 200,
|
| 171 |
+
"prompt_tokens": 288,
|
| 172 |
+
"cached_tokens": 0
|
| 173 |
+
},
|
| 174 |
+
{
|
| 175 |
+
"http_status": 200,
|
| 176 |
+
"prompt_tokens": 309,
|
| 177 |
+
"cached_tokens": 0
|
| 178 |
+
},
|
| 179 |
+
{
|
| 180 |
+
"http_status": 200,
|
| 181 |
+
"prompt_tokens": 330,
|
| 182 |
+
"cached_tokens": 0
|
| 183 |
+
},
|
| 184 |
+
{
|
| 185 |
+
"http_status": 200,
|
| 186 |
+
"prompt_tokens": 351,
|
| 187 |
+
"cached_tokens": 0
|
| 188 |
+
},
|
| 189 |
+
{
|
| 190 |
+
"http_status": 200,
|
| 191 |
+
"prompt_tokens": 372,
|
| 192 |
+
"cached_tokens": 0
|
| 193 |
+
},
|
| 194 |
+
{
|
| 195 |
+
"http_status": 200,
|
| 196 |
+
"prompt_tokens": 393,
|
| 197 |
+
"cached_tokens": 0
|
| 198 |
+
},
|
| 199 |
+
{
|
| 200 |
+
"http_status": 200,
|
| 201 |
+
"prompt_tokens": 414,
|
| 202 |
+
"cached_tokens": 0
|
| 203 |
+
},
|
| 204 |
+
{
|
| 205 |
+
"http_status": 200,
|
| 206 |
+
"stream_complete": true
|
| 207 |
+
},
|
| 208 |
+
{
|
| 209 |
+
"http_status": 200,
|
| 210 |
+
"prompt_tokens": 457,
|
| 211 |
+
"cached_tokens": 0
|
| 212 |
+
}
|
| 213 |
+
],
|
| 214 |
+
"byte_limit": 4186895080,
|
| 215 |
+
"peak_retained_bytes": 0,
|
| 216 |
+
"retained_sequences": 0,
|
| 217 |
+
"insertions": 45
|
| 218 |
+
}
|
| 219 |
+
],
|
| 220 |
+
"peak_mlx_bytes": 430070044,
|
| 221 |
+
"status": "PASS"
|
| 222 |
+
}
|
|
@@ -0,0 +1,71 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model_type": "qwen3_5",
|
| 3 |
+
"architectures": [
|
| 4 |
+
"Qwen3_5ForConditionalGeneration"
|
| 5 |
+
],
|
| 6 |
+
"language_model_only": false,
|
| 7 |
+
"tie_word_embeddings": false,
|
| 8 |
+
"image_token_id": 248056,
|
| 9 |
+
"video_token_id": 248057,
|
| 10 |
+
"vision_start_token_id": 248053,
|
| 11 |
+
"text_config": {
|
| 12 |
+
"model_type": "qwen3_5_text",
|
| 13 |
+
"hidden_size": 128,
|
| 14 |
+
"intermediate_size": 256,
|
| 15 |
+
"num_hidden_layers": 4,
|
| 16 |
+
"num_attention_heads": 4,
|
| 17 |
+
"num_key_value_heads": 2,
|
| 18 |
+
"head_dim": 32,
|
| 19 |
+
"vocab_size": 248320,
|
| 20 |
+
"full_attention_interval": 4,
|
| 21 |
+
"layer_types": [
|
| 22 |
+
"linear_attention",
|
| 23 |
+
"linear_attention",
|
| 24 |
+
"linear_attention",
|
| 25 |
+
"full_attention"
|
| 26 |
+
],
|
| 27 |
+
"linear_num_key_heads": 2,
|
| 28 |
+
"linear_num_value_heads": 4,
|
| 29 |
+
"linear_key_head_dim": 32,
|
| 30 |
+
"linear_value_head_dim": 32,
|
| 31 |
+
"linear_conv_kernel_dim": 4,
|
| 32 |
+
"hidden_act": "silu",
|
| 33 |
+
"attn_output_gate": true,
|
| 34 |
+
"output_gate_type": "swish",
|
| 35 |
+
"mamba_ssm_dtype": "float32",
|
| 36 |
+
"rms_norm_eps": 1e-06,
|
| 37 |
+
"max_position_embeddings": 1024,
|
| 38 |
+
"tie_word_embeddings": false,
|
| 39 |
+
"attention_bias": false,
|
| 40 |
+
"attention_dropout": 0.0,
|
| 41 |
+
"mtp_num_hidden_layers": 1,
|
| 42 |
+
"mtp_use_dedicated_embeddings": false,
|
| 43 |
+
"rope_parameters": {
|
| 44 |
+
"rope_type": "default",
|
| 45 |
+
"rope_theta": 10000000,
|
| 46 |
+
"partial_rotary_factor": 0.5,
|
| 47 |
+
"mrope_interleaved": true,
|
| 48 |
+
"mrope_section": [
|
| 49 |
+
3,
|
| 50 |
+
3,
|
| 51 |
+
2
|
| 52 |
+
]
|
| 53 |
+
}
|
| 54 |
+
},
|
| 55 |
+
"vision_config": {
|
| 56 |
+
"model_type": "qwen3_5",
|
| 57 |
+
"depth": 2,
|
| 58 |
+
"hidden_size": 32,
|
| 59 |
+
"intermediate_size": 48,
|
| 60 |
+
"num_heads": 4,
|
| 61 |
+
"out_hidden_size": 128,
|
| 62 |
+
"num_position_embeddings": 16,
|
| 63 |
+
"patch_size": 2,
|
| 64 |
+
"temporal_patch_size": 2,
|
| 65 |
+
"spatial_merge_size": 2,
|
| 66 |
+
"in_channels": 3,
|
| 67 |
+
"hidden_act": "gelu_pytorch_tanh",
|
| 68 |
+
"deepstack_visual_indexes": []
|
| 69 |
+
},
|
| 70 |
+
"vision_end_token_id": 248054
|
| 71 |
+
}
|
|
@@ -0,0 +1,271 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Copyright © 2026 Swift contributors.
|
| 2 |
+
|
| 3 |
+
"""Offline HTTP regression checks on tiny random models, never Swift weights."""
|
| 4 |
+
|
| 5 |
+
import argparse
|
| 6 |
+
import contextlib
|
| 7 |
+
import copy
|
| 8 |
+
import functools
|
| 9 |
+
import importlib.metadata
|
| 10 |
+
import json
|
| 11 |
+
import logging
|
| 12 |
+
import os
|
| 13 |
+
import platform
|
| 14 |
+
import shutil
|
| 15 |
+
import sys
|
| 16 |
+
import tempfile
|
| 17 |
+
import threading
|
| 18 |
+
import urllib.request
|
| 19 |
+
from http.server import ThreadingHTTPServer
|
| 20 |
+
from pathlib import Path
|
| 21 |
+
from unittest.mock import patch
|
| 22 |
+
|
| 23 |
+
os.environ["HF_HUB_OFFLINE"] = "1"
|
| 24 |
+
os.environ["TRANSFORMERS_OFFLINE"] = "1"
|
| 25 |
+
|
| 26 |
+
import mlx.core as mx
|
| 27 |
+
import mlx.nn as nn
|
| 28 |
+
from mlx.utils import tree_flatten
|
| 29 |
+
|
| 30 |
+
from mlx_lm import server
|
| 31 |
+
from mlx_lm.models.cache import LRUPromptCache
|
| 32 |
+
from mlx_lm.models.qwen3_5_full import Model, ModelArgs
|
| 33 |
+
|
| 34 |
+
|
| 35 |
+
class RecordingCache(LRUPromptCache):
|
| 36 |
+
def __init__(self, *args, **kwargs):
|
| 37 |
+
super().__init__(*args, **kwargs)
|
| 38 |
+
self.peak_bytes = 0
|
| 39 |
+
self.insertions = 0
|
| 40 |
+
|
| 41 |
+
def insert_cache(self, *args, **kwargs):
|
| 42 |
+
super().insert_cache(*args, **kwargs)
|
| 43 |
+
self.insertions += 1
|
| 44 |
+
self.peak_bytes = max(self.peak_bytes, self.nbytes)
|
| 45 |
+
assert self.nbytes <= self.max_bytes
|
| 46 |
+
assert len(self) <= self.max_size
|
| 47 |
+
|
| 48 |
+
|
| 49 |
+
def make_fixture(folder, assets, config_path, bits):
|
| 50 |
+
config = json.loads(config_path.read_text())
|
| 51 |
+
source = json.loads((assets / "config.json").read_text())
|
| 52 |
+
config["text_config"]["vocab_size"] = source["text_config"]["vocab_size"]
|
| 53 |
+
config.pop("quantization", None)
|
| 54 |
+
mx.random.seed(25)
|
| 55 |
+
model = Model(ModelArgs.from_dict(copy.deepcopy(config)))
|
| 56 |
+
model.language_model.lm_head.weight = mx.zeros_like(
|
| 57 |
+
model.language_model.lm_head.weight
|
| 58 |
+
)
|
| 59 |
+
model.apply(
|
| 60 |
+
lambda value: (
|
| 61 |
+
value.astype(mx.bfloat16)
|
| 62 |
+
if mx.issubdtype(value.dtype, mx.floating)
|
| 63 |
+
else value
|
| 64 |
+
)
|
| 65 |
+
)
|
| 66 |
+
nn.quantize(
|
| 67 |
+
model,
|
| 68 |
+
bits=bits,
|
| 69 |
+
group_size=64,
|
| 70 |
+
mode="affine",
|
| 71 |
+
class_predicate=lambda path, layer: hasattr(layer, "to_quantized")
|
| 72 |
+
and layer.weight.shape[-1] % 64 == 0,
|
| 73 |
+
)
|
| 74 |
+
mx.eval(model.parameters())
|
| 75 |
+
weights = dict(tree_flatten(model.parameters()))
|
| 76 |
+
mx.save_safetensors(
|
| 77 |
+
str(folder / "model.safetensors"),
|
| 78 |
+
weights,
|
| 79 |
+
metadata={"purpose": "synthetic regression only"},
|
| 80 |
+
)
|
| 81 |
+
index = {
|
| 82 |
+
"metadata": {"total_size": sum(w.nbytes for w in weights.values())},
|
| 83 |
+
"weight_map": {name: "model.safetensors" for name in weights},
|
| 84 |
+
}
|
| 85 |
+
(folder / "model.safetensors.index.json").write_text(json.dumps(index))
|
| 86 |
+
config["quantization"] = {"bits": bits, "group_size": 64, "mode": "affine"}
|
| 87 |
+
(folder / "config.json").write_text(json.dumps(config))
|
| 88 |
+
for name in (
|
| 89 |
+
"tokenizer.json",
|
| 90 |
+
"tokenizer_config.json",
|
| 91 |
+
"vocab.json",
|
| 92 |
+
"merges.txt",
|
| 93 |
+
"chat_template.jinja",
|
| 94 |
+
"generation_config.json",
|
| 95 |
+
):
|
| 96 |
+
shutil.copyfile(assets / name, folder / name)
|
| 97 |
+
del model, weights
|
| 98 |
+
mx.clear_cache()
|
| 99 |
+
|
| 100 |
+
|
| 101 |
+
@contextlib.contextmanager
|
| 102 |
+
def running(folder, options):
|
| 103 |
+
captured = []
|
| 104 |
+
argv = ["mlx_lm.server", "--model", str(folder), "--port", "0", *options]
|
| 105 |
+
with patch.object(sys, "argv", argv), patch.object(
|
| 106 |
+
server, "run", side_effect=lambda h, p, m: captured.append(m)
|
| 107 |
+
), patch.object(server, "maybe_set_recommended_wired_limit", return_value=None):
|
| 108 |
+
server.main()
|
| 109 |
+
provider = captured[0]
|
| 110 |
+
instances = []
|
| 111 |
+
|
| 112 |
+
def start_http(host, port, generator):
|
| 113 |
+
handler = functools.partial(
|
| 114 |
+
server.APIHandler, generator, system_fingerprint="synthetic-offline-test"
|
| 115 |
+
)
|
| 116 |
+
httpd = ThreadingHTTPServer((host, port), handler)
|
| 117 |
+
thread = threading.Thread(target=httpd.serve_forever, daemon=True)
|
| 118 |
+
instances.append((httpd, thread, generator))
|
| 119 |
+
thread.start()
|
| 120 |
+
|
| 121 |
+
with patch.object(server, "LRUPromptCache", RecordingCache), patch.object(
|
| 122 |
+
server, "_run_http_server", start_http
|
| 123 |
+
):
|
| 124 |
+
server.run("127.0.0.1", 0, provider)
|
| 125 |
+
httpd, thread, generator = instances[0]
|
| 126 |
+
try:
|
| 127 |
+
yield f"http://127.0.0.1:{httpd.server_port}", generator, provider.cli_args
|
| 128 |
+
finally:
|
| 129 |
+
httpd.shutdown()
|
| 130 |
+
httpd.server_close()
|
| 131 |
+
thread.join(timeout=5)
|
| 132 |
+
generator.stop_and_join()
|
| 133 |
+
|
| 134 |
+
|
| 135 |
+
def request(url, messages, stream=False, seeded=False):
|
| 136 |
+
data = {
|
| 137 |
+
"model": "default_model",
|
| 138 |
+
"messages": messages,
|
| 139 |
+
"max_tokens": 4,
|
| 140 |
+
"temperature": 0,
|
| 141 |
+
"chat_template_kwargs": {"enable_thinking": False},
|
| 142 |
+
"stream": stream,
|
| 143 |
+
}
|
| 144 |
+
if seeded:
|
| 145 |
+
data["seed"] = 25
|
| 146 |
+
req = urllib.request.Request(
|
| 147 |
+
url + "/v1/chat/completions",
|
| 148 |
+
data=json.dumps(data).encode(),
|
| 149 |
+
headers={"Content-Type": "application/json"},
|
| 150 |
+
)
|
| 151 |
+
with urllib.request.urlopen(req, timeout=60) as response:
|
| 152 |
+
body = response.read().decode()
|
| 153 |
+
assert response.status == 200
|
| 154 |
+
if stream:
|
| 155 |
+
assert "data: [DONE]" in body
|
| 156 |
+
chunks = [
|
| 157 |
+
json.loads(line[6:])
|
| 158 |
+
for line in body.splitlines()
|
| 159 |
+
if line.startswith("data: ") and line != "data: [DONE]"
|
| 160 |
+
]
|
| 161 |
+
text = "".join(
|
| 162 |
+
choice.get("delta", {}).get("content", "") or ""
|
| 163 |
+
for chunk in chunks
|
| 164 |
+
for choice in chunk.get("choices", [])
|
| 165 |
+
)
|
| 166 |
+
assert text == "!!!!", text
|
| 167 |
+
return {"http_status": 200, "stream_complete": True}
|
| 168 |
+
result = json.loads(body)
|
| 169 |
+
assert result["choices"][0]["message"]["content"] == "!!!!"
|
| 170 |
+
usage = result["usage"]
|
| 171 |
+
return {
|
| 172 |
+
"http_status": 200,
|
| 173 |
+
"prompt_tokens": usage["prompt_tokens"],
|
| 174 |
+
"cached_tokens": usage["prompt_tokens_details"]["cached_tokens"],
|
| 175 |
+
}
|
| 176 |
+
|
| 177 |
+
|
| 178 |
+
def main():
|
| 179 |
+
parser = argparse.ArgumentParser(description=__doc__)
|
| 180 |
+
parser.add_argument(
|
| 181 |
+
"--assets",
|
| 182 |
+
type=Path,
|
| 183 |
+
required=True,
|
| 184 |
+
help="Local Swift snapshot; only tokenizer and config assets are read",
|
| 185 |
+
)
|
| 186 |
+
parser.add_argument("--config", type=Path, required=True)
|
| 187 |
+
parser.add_argument("--bits", type=int, choices=(4, 5), required=True)
|
| 188 |
+
parser.add_argument("--output", type=Path, required=True)
|
| 189 |
+
args = parser.parse_args()
|
| 190 |
+
logging.basicConfig(level=logging.WARNING)
|
| 191 |
+
report = {
|
| 192 |
+
"scope": "Tiny random weights with real local tokenizer assets. "
|
| 193 |
+
"No released model weights loaded or changed; no full-model capacity claim.",
|
| 194 |
+
"bits": args.bits,
|
| 195 |
+
"offline_hub": True,
|
| 196 |
+
"memory_limit_overrides": False,
|
| 197 |
+
"machine": platform.machine(),
|
| 198 |
+
"python": platform.python_version(),
|
| 199 |
+
"packages": {
|
| 200 |
+
name: importlib.metadata.version(name)
|
| 201 |
+
for name in ("mlx", "mlx-lm", "transformers", "huggingface_hub")
|
| 202 |
+
},
|
| 203 |
+
"profiles": [],
|
| 204 |
+
}
|
| 205 |
+
profiles = [
|
| 206 |
+
("defaults", []),
|
| 207 |
+
("explicit_bytes", ["--prompt-cache-bytes", "100000"]),
|
| 208 |
+
("no_retention", ["--prompt-cache-size", "0"]),
|
| 209 |
+
]
|
| 210 |
+
with tempfile.TemporaryDirectory(prefix="swift-cache-test-") as temp:
|
| 211 |
+
folder = Path(temp)
|
| 212 |
+
make_fixture(folder, args.assets, args.config, args.bits)
|
| 213 |
+
for name, options in profiles:
|
| 214 |
+
with running(folder, options) as (url, generator, cli):
|
| 215 |
+
assert cli.prompt_concurrency == cli.decode_concurrency == 1
|
| 216 |
+
assert cli.prefill_step_size == 512
|
| 217 |
+
messages = [
|
| 218 |
+
{
|
| 219 |
+
"role": "system",
|
| 220 |
+
"content": "Help with coding. "
|
| 221 |
+
+ "Read the session carefully. " * 20,
|
| 222 |
+
},
|
| 223 |
+
{
|
| 224 |
+
"role": "user",
|
| 225 |
+
"content": "Remember this context. " + "sample words " * 40,
|
| 226 |
+
},
|
| 227 |
+
{"role": "assistant", "content": "Previous response."},
|
| 228 |
+
{"role": "user", "content": "Continue briefly."},
|
| 229 |
+
]
|
| 230 |
+
row = {"name": name, "requests": []}
|
| 231 |
+
for index in range(12):
|
| 232 |
+
result = request(
|
| 233 |
+
url, messages, stream=index == 10, seeded=index == 11
|
| 234 |
+
)
|
| 235 |
+
assert generator.generation_available()
|
| 236 |
+
row["requests"].append(result)
|
| 237 |
+
messages.extend(
|
| 238 |
+
[
|
| 239 |
+
{"role": "assistant", "content": "!!!!"},
|
| 240 |
+
{
|
| 241 |
+
"role": "user",
|
| 242 |
+
"content": f"Another short answer {index}.",
|
| 243 |
+
},
|
| 244 |
+
]
|
| 245 |
+
)
|
| 246 |
+
cache = generator.prompt_cache
|
| 247 |
+
assert cache.insertions > 0
|
| 248 |
+
row.update(
|
| 249 |
+
byte_limit=cache.max_bytes,
|
| 250 |
+
peak_retained_bytes=cache.peak_bytes,
|
| 251 |
+
retained_sequences=len(cache),
|
| 252 |
+
insertions=cache.insertions,
|
| 253 |
+
)
|
| 254 |
+
nonstream = [r for r in row["requests"] if "cached_tokens" in r]
|
| 255 |
+
if name == "no_retention":
|
| 256 |
+
assert cache.peak_bytes == 0
|
| 257 |
+
assert all(r["cached_tokens"] == 0 for r in nonstream)
|
| 258 |
+
else:
|
| 259 |
+
assert any(r["cached_tokens"] > 0 for r in nonstream[1:])
|
| 260 |
+
assert 0 < cache.max_bytes < 1 << 63
|
| 261 |
+
if name == "explicit_bytes":
|
| 262 |
+
assert cache.max_bytes == 100000
|
| 263 |
+
report["profiles"].append(row)
|
| 264 |
+
report["peak_mlx_bytes"] = mx.get_peak_memory()
|
| 265 |
+
report["status"] = "PASS"
|
| 266 |
+
args.output.write_text(json.dumps(report, indent=2) + "\n")
|
| 267 |
+
print(json.dumps(report, indent=2))
|
| 268 |
+
|
| 269 |
+
|
| 270 |
+
if __name__ == "__main__":
|
| 271 |
+
main()
|
|
@@ -0,0 +1,55 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"status": "PASS_LOCAL_REGRESSION_CHECKS",
|
| 3 |
+
"base_commit": "c69d1288440a0dc4e6401fc417098b07598dccd5",
|
| 4 |
+
"mlx_version": "0.32.2",
|
| 5 |
+
"mlx_lm_version": "0.32.0",
|
| 6 |
+
"server_sha256_before": "57392576de114af80ede723976b9dd156dc234f856b9943b8b44b8b78bc27f79",
|
| 7 |
+
"server_sha256_after": "791f8c1eb3431c4d24b4f9084f67b665c38e3d547ee4a307e52c5b2b6f12aba0",
|
| 8 |
+
"patch_sha256": "a7fbc0f0524ee7864d9f41a98a2e35acf0e63f53e369ee5b9a87ca12926a9eab",
|
| 9 |
+
"unit_and_upstream_server_tests": {
|
| 10 |
+
"passed": 44,
|
| 11 |
+
"failed": 0
|
| 12 |
+
},
|
| 13 |
+
"offline_http_tests": {
|
| 14 |
+
"passed_requests": 72,
|
| 15 |
+
"failed_requests": 0
|
| 16 |
+
},
|
| 17 |
+
"scope": "Local Apple Silicon tests using small fixtures. Full Swift long-context capacity is not established.",
|
| 18 |
+
"hosted_ci": "NOT_RUN; these are recorded local checks, not a hosted CI status",
|
| 19 |
+
"weights_changed": false,
|
| 20 |
+
"quantization_changed": false,
|
| 21 |
+
"fresh_patch_replay": {
|
| 22 |
+
"base_commit": "c69d1288440a0dc4e6401fc417098b07598dccd5",
|
| 23 |
+
"status": "PASS",
|
| 24 |
+
"releases": [
|
| 25 |
+
{
|
| 26 |
+
"bits": 4,
|
| 27 |
+
"clean_patch_apply": "PASS",
|
| 28 |
+
"rollback_check": "PASS",
|
| 29 |
+
"runtime_identity": "PASS",
|
| 30 |
+
"model_implementation_unchanged": "PASS",
|
| 31 |
+
"unit_and_upstream_tests_passed": 44,
|
| 32 |
+
"offline_http_requests_passed": 36,
|
| 33 |
+
"black": "25.1.0 PASS",
|
| 34 |
+
"isort": "6.0.0 PASS",
|
| 35 |
+
"ruff": "0.16.6 PASS",
|
| 36 |
+
"patch_sha256": "a7fbc0f0524ee7864d9f41a98a2e35acf0e63f53e369ee5b9a87ca12926a9eab",
|
| 37 |
+
"status": "PASS"
|
| 38 |
+
},
|
| 39 |
+
{
|
| 40 |
+
"bits": 5,
|
| 41 |
+
"clean_patch_apply": "PASS",
|
| 42 |
+
"rollback_check": "PASS",
|
| 43 |
+
"runtime_identity": "PASS",
|
| 44 |
+
"model_implementation_unchanged": "PASS",
|
| 45 |
+
"unit_and_upstream_tests_passed": 44,
|
| 46 |
+
"offline_http_requests_passed": 36,
|
| 47 |
+
"black": "25.1.0 PASS",
|
| 48 |
+
"isort": "6.0.0 PASS",
|
| 49 |
+
"ruff": "0.16.6 PASS",
|
| 50 |
+
"patch_sha256": "a7fbc0f0524ee7864d9f41a98a2e35acf0e63f53e369ee5b9a87ca12926a9eab",
|
| 51 |
+
"status": "PASS"
|
| 52 |
+
}
|
| 53 |
+
]
|
| 54 |
+
}
|
| 55 |
+
}
|
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Verify the installed server patch without loading weights."""
|
| 2 |
+
import hashlib
|
| 3 |
+
import json
|
| 4 |
+
from pathlib import Path
|
| 5 |
+
from mlx_lm import server
|
| 6 |
+
|
| 7 |
+
manifest = json.loads((Path(__file__).parent / "validation.json").read_text())
|
| 8 |
+
actual = hashlib.sha256(Path(server.__file__).read_bytes()).hexdigest()
|
| 9 |
+
if actual != manifest["server_sha256_after"]:
|
| 10 |
+
raise SystemExit("FAIL: this Python environment does not use the checked server patch")
|
| 11 |
+
print("PASS: installed server source matches the checked cache patch")
|
|
@@ -0,0 +1,253 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
diff --git a/mlx_lm/server.py b/mlx_lm/server.py
|
| 2 |
+
index 9462d6e..36f7787 100644
|
| 3 |
+
--- a/mlx_lm/server.py
|
| 4 |
+
+++ b/mlx_lm/server.py
|
| 5 |
+
@@ -29,6 +29,7 @@ from typing import (
|
| 6 |
+
|
| 7 |
+
import mlx.core as mx
|
| 8 |
+
from huggingface_hub import scan_cache_dir
|
| 9 |
+
+from mlx.utils import tree_flatten
|
| 10 |
+
|
| 11 |
+
from ._version import __version__
|
| 12 |
+
from .generate import (
|
| 13 |
+
@@ -428,6 +429,26 @@ def _format_top_logprobs(logprobs, top_n, tokenizer) -> Tuple[Dict[str, Any]]:
|
| 14 |
+
)
|
| 15 |
+
|
| 16 |
+
|
| 17 |
+
+def _prompt_cache_byte_limit(cli_args, model):
|
| 18 |
+
+ limit = getattr(cli_args, "prompt_cache_bytes", None)
|
| 19 |
+
+ if limit is not None:
|
| 20 |
+
+ if limit < 0:
|
| 21 |
+
+ raise ValueError("--prompt-cache-bytes must be non-negative")
|
| 22 |
+
+ return limit
|
| 23 |
+
+
|
| 24 |
+
+ # Reserve workspace and half the remaining capacity for the active request.
|
| 25 |
+
+ try:
|
| 26 |
+
+ recommended = mx.device_info().get("max_recommended_working_set_size")
|
| 27 |
+
+ except (RuntimeError, ValueError):
|
| 28 |
+
+ recommended = None
|
| 29 |
+
+ if not recommended:
|
| 30 |
+
+ return 1 << 63
|
| 31 |
+
+ parameters = {id(p): p for _, p in tree_flatten(model.parameters())}
|
| 32 |
+
+ model_bytes = sum(p.nbytes for p in parameters.values())
|
| 33 |
+
+ reserve = max(4 * 1024**3, recommended // 8)
|
| 34 |
+
+ return max(0, (recommended - model_bytes - reserve) // 2)
|
| 35 |
+
+
|
| 36 |
+
+
|
| 37 |
+
class ResponseGenerator:
|
| 38 |
+
def __init__(self, model_provider: ModelProvider, prompt_cache: LRUPromptCache):
|
| 39 |
+
self.model_provider = model_provider
|
| 40 |
+
@@ -443,6 +464,13 @@ class ResponseGenerator:
|
| 41 |
+
self._generation_thread = Thread(target=self._run_generate)
|
| 42 |
+
self._generation_thread.start()
|
| 43 |
+
|
| 44 |
+
+ def _configure_prompt_cache(self, model):
|
| 45 |
+
+ limit = _prompt_cache_byte_limit(self.cli_args, model)
|
| 46 |
+
+ if self.prompt_cache.max_bytes != limit:
|
| 47 |
+
+ self.prompt_cache.max_bytes = limit
|
| 48 |
+
+ self.prompt_cache.trim_to(n_bytes=limit)
|
| 49 |
+
+ logging.info("Retained prompt-cache limit: %.2f GB", limit / 1e9)
|
| 50 |
+
+
|
| 51 |
+
def _run_generate(self):
|
| 52 |
+
try:
|
| 53 |
+
self._generate()
|
| 54 |
+
@@ -772,6 +800,7 @@ class ResponseGenerator:
|
| 55 |
+
model, tokenizer = self.model_provider.load(
|
| 56 |
+
args.model.model, args.model.adapter, args.model.draft
|
| 57 |
+
)
|
| 58 |
+
+ self._configure_prompt_cache(model)
|
| 59 |
+
except Exception as e:
|
| 60 |
+
rqueue.put(e)
|
| 61 |
+
continue
|
| 62 |
+
@@ -1762,7 +1791,13 @@ def run(
|
| 63 |
+
handler_class=APIHandler,
|
| 64 |
+
):
|
| 65 |
+
group = mx.distributed.init()
|
| 66 |
+
- prompt_cache = LRUPromptCache(model_provider.cli_args.prompt_cache_size)
|
| 67 |
+
+ cache_bytes = model_provider.cli_args.prompt_cache_bytes
|
| 68 |
+
+ if cache_bytes is not None and cache_bytes < 0:
|
| 69 |
+
+ raise ValueError("--prompt-cache-bytes must be non-negative")
|
| 70 |
+
+ prompt_cache = LRUPromptCache(
|
| 71 |
+
+ model_provider.cli_args.prompt_cache_size,
|
| 72 |
+
+ max_bytes=cache_bytes if cache_bytes is not None else 1 << 63,
|
| 73 |
+
+ )
|
| 74 |
+
response_generator = ResponseGenerator(model_provider, prompt_cache)
|
| 75 |
+
if group.rank() == 0:
|
| 76 |
+
_run_http_server(host, port, response_generator)
|
| 77 |
+
@@ -1875,31 +1910,32 @@ def main():
|
| 78 |
+
parser.add_argument(
|
| 79 |
+
"--decode-concurrency",
|
| 80 |
+
type=int,
|
| 81 |
+
- default=32,
|
| 82 |
+
+ default=1,
|
| 83 |
+
help="When a request is batchable then decode that many requests in parallel",
|
| 84 |
+
)
|
| 85 |
+
parser.add_argument(
|
| 86 |
+
"--prompt-concurrency",
|
| 87 |
+
type=int,
|
| 88 |
+
- default=8,
|
| 89 |
+
+ default=1,
|
| 90 |
+
help="When a request is batchable then process that many prompts in parallel",
|
| 91 |
+
)
|
| 92 |
+
parser.add_argument(
|
| 93 |
+
"--prefill-step-size",
|
| 94 |
+
type=int,
|
| 95 |
+
- default=2048,
|
| 96 |
+
- help="Step size for prefill processing (default: 2048)",
|
| 97 |
+
+ default=512,
|
| 98 |
+
+ help="Step size for prefill processing (default: 512)",
|
| 99 |
+
)
|
| 100 |
+
parser.add_argument(
|
| 101 |
+
"--prompt-cache-size",
|
| 102 |
+
type=int,
|
| 103 |
+
- default=10,
|
| 104 |
+
- help="Maximum number of distinct KV caches to hold in the prompt cache",
|
| 105 |
+
+ default=2,
|
| 106 |
+
+ help="Maximum retained prompt-cache entries (default: 2)",
|
| 107 |
+
)
|
| 108 |
+
parser.add_argument(
|
| 109 |
+
"--prompt-cache-bytes",
|
| 110 |
+
type=_parse_size,
|
| 111 |
+
- help="Maximum size in bytes of the KV caches",
|
| 112 |
+
+ help="Maximum retained prompt-cache bytes. Default: automatic on Metal. "
|
| 113 |
+
+ "This does not cap active-request or total process memory.",
|
| 114 |
+
)
|
| 115 |
+
parser.add_argument(
|
| 116 |
+
"--kv-bits",
|
| 117 |
+
diff --git a/tests/test_server_cache_budget.py b/tests/test_server_cache_budget.py
|
| 118 |
+
new file mode 100644
|
| 119 |
+
--- /dev/null
|
| 120 |
+
+++ b/tests/test_server_cache_budget.py
|
| 121 |
+
@@ -0,0 +1,132 @@
|
| 122 |
+
+# Copyright © 2026 Swift contributors.
|
| 123 |
+
+
|
| 124 |
+
+import sys
|
| 125 |
+
+from types import SimpleNamespace
|
| 126 |
+
+from unittest.mock import Mock
|
| 127 |
+
+
|
| 128 |
+
+import pytest
|
| 129 |
+
+
|
| 130 |
+
+from mlx_lm import server
|
| 131 |
+
+from mlx_lm.models.cache import LRUPromptCache
|
| 132 |
+
+
|
| 133 |
+
+
|
| 134 |
+
+class CacheState:
|
| 135 |
+
+ def __init__(self, nbytes):
|
| 136 |
+
+ self.nbytes = nbytes
|
| 137 |
+
+
|
| 138 |
+
+ def is_trimmable(self):
|
| 139 |
+
+ return False
|
| 140 |
+
+
|
| 141 |
+
+
|
| 142 |
+
+@pytest.mark.parametrize("limit", [0, 100])
|
| 143 |
+
+def test_server_factory_enforces_configured_bytes(monkeypatch, limit):
|
| 144 |
+
+ caches = []
|
| 145 |
+
+ provider = SimpleNamespace(
|
| 146 |
+
+ cli_args=SimpleNamespace(prompt_cache_size=10, prompt_cache_bytes=limit)
|
| 147 |
+
+ )
|
| 148 |
+
+ monkeypatch.setattr(
|
| 149 |
+
+ server, "ResponseGenerator", lambda provider, cache: caches.append(cache)
|
| 150 |
+
+ )
|
| 151 |
+
+ monkeypatch.setattr(server, "_run_http_server", lambda *args: None)
|
| 152 |
+
+ server.run("127.0.0.1", 0, provider)
|
| 153 |
+
+ cache = caches[0]
|
| 154 |
+
+ for i in range(5):
|
| 155 |
+
+ cache.insert_cache("model", [i, 1], [CacheState(80)])
|
| 156 |
+
+ assert cache.nbytes <= limit
|
| 157 |
+
+ if limit:
|
| 158 |
+
+ reused, remaining = cache.fetch_nearest_cache("model", [4, 1, 9])
|
| 159 |
+
+ assert reused is not None
|
| 160 |
+
+ assert remaining == [9]
|
| 161 |
+
+
|
| 162 |
+
+
|
| 163 |
+
+def test_negative_limit_rejected_before_worker_start(monkeypatch):
|
| 164 |
+
+ worker = Mock()
|
| 165 |
+
+ monkeypatch.setattr(server, "ResponseGenerator", worker)
|
| 166 |
+
+ provider = SimpleNamespace(
|
| 167 |
+
+ cli_args=SimpleNamespace(prompt_cache_size=2, prompt_cache_bytes=-1)
|
| 168 |
+
+ )
|
| 169 |
+
+ with pytest.raises(ValueError, match="non-negative"):
|
| 170 |
+
+ server.run("127.0.0.1", 0, provider)
|
| 171 |
+
+ worker.assert_not_called()
|
| 172 |
+
+
|
| 173 |
+
+
|
| 174 |
+
+def test_explicit_limit_does_not_probe_model(monkeypatch):
|
| 175 |
+
+ probe = Mock(side_effect=AssertionError("device probe was not needed"))
|
| 176 |
+
+ monkeypatch.setattr(server.mx, "device_info", probe)
|
| 177 |
+
+ model = Mock()
|
| 178 |
+
+ for limit in (0, 1024, 8 * 1024**3):
|
| 179 |
+
+ assert (
|
| 180 |
+
+ server._prompt_cache_byte_limit(
|
| 181 |
+
+ SimpleNamespace(prompt_cache_bytes=limit), model
|
| 182 |
+
+ )
|
| 183 |
+
+ == limit
|
| 184 |
+
+ )
|
| 185 |
+
+ model.parameters.assert_not_called()
|
| 186 |
+
+
|
| 187 |
+
+
|
| 188 |
+
+def test_auto_limit_reserves_room_for_active_request(monkeypatch):
|
| 189 |
+
+ gib = 1024**3
|
| 190 |
+
+ monkeypatch.setattr(
|
| 191 |
+
+ server.mx,
|
| 192 |
+
+ "device_info",
|
| 193 |
+
+ lambda: {"max_recommended_working_set_size": 36 * gib},
|
| 194 |
+
+ )
|
| 195 |
+
+ model = SimpleNamespace(parameters=lambda: {"weight": CacheState(15 * gib)})
|
| 196 |
+
+ args = SimpleNamespace(prompt_cache_bytes=None)
|
| 197 |
+
+ limit = server._prompt_cache_byte_limit(args, model)
|
| 198 |
+
+ assert 6 * gib <= limit < (36 - 15) * gib // 2
|
| 199 |
+
+ larger = SimpleNamespace(parameters=lambda: {"weight": CacheState(30 * gib)})
|
| 200 |
+
+ assert server._prompt_cache_byte_limit(args, larger) < limit
|
| 201 |
+
+ full = SimpleNamespace(parameters=lambda: {"weight": CacheState(36 * gib)})
|
| 202 |
+
+ assert server._prompt_cache_byte_limit(args, full) == 0
|
| 203 |
+
+
|
| 204 |
+
+
|
| 205 |
+
+def test_model_swap_trims_retained_states(monkeypatch):
|
| 206 |
+
+ obj = server.ResponseGenerator.__new__(server.ResponseGenerator)
|
| 207 |
+
+ obj.model_provider = SimpleNamespace(
|
| 208 |
+
+ cli_args=SimpleNamespace(prompt_cache_bytes=100)
|
| 209 |
+
+ )
|
| 210 |
+
+ obj.prompt_cache = LRUPromptCache(max_size=10)
|
| 211 |
+
+ obj.prompt_cache.insert_cache("old", [1], [CacheState(200)])
|
| 212 |
+
+ obj._configure_prompt_cache(Mock())
|
| 213 |
+
+ assert obj.prompt_cache.nbytes == 0
|
| 214 |
+
+ assert obj.prompt_cache.max_bytes == 100
|
| 215 |
+
+
|
| 216 |
+
+
|
| 217 |
+
+@pytest.mark.parametrize("device_info", [{}, {"max_recommended_working_set_size": 0}])
|
| 218 |
+
+def test_auto_limit_without_metal_metadata(monkeypatch, device_info):
|
| 219 |
+
+ monkeypatch.setattr(server.mx, "device_info", lambda: device_info)
|
| 220 |
+
+ model = Mock()
|
| 221 |
+
+ assert (
|
| 222 |
+
+ server._prompt_cache_byte_limit(SimpleNamespace(prompt_cache_bytes=None), model)
|
| 223 |
+
+ == 1 << 63
|
| 224 |
+
+ )
|
| 225 |
+
+ model.parameters.assert_not_called()
|
| 226 |
+
+
|
| 227 |
+
+
|
| 228 |
+
+def test_auto_limit_counts_tied_parameters_once(monkeypatch):
|
| 229 |
+
+ gib = 1024**3
|
| 230 |
+
+ monkeypatch.setattr(
|
| 231 |
+
+ server.mx,
|
| 232 |
+
+ "device_info",
|
| 233 |
+
+ lambda: {"max_recommended_working_set_size": 32 * gib},
|
| 234 |
+
+ )
|
| 235 |
+
+ weight = CacheState(8 * gib)
|
| 236 |
+
+ model = SimpleNamespace(parameters=lambda: {"embed": weight, "head": weight})
|
| 237 |
+
+ assert (
|
| 238 |
+
+ server._prompt_cache_byte_limit(SimpleNamespace(prompt_cache_bytes=None), model)
|
| 239 |
+
+ == 10 * gib
|
| 240 |
+
+ )
|
| 241 |
+
+
|
| 242 |
+
+
|
| 243 |
+
+def test_defaults_keep_reuse_enabled_and_limit_concurrency(monkeypatch):
|
| 244 |
+
+ captured = []
|
| 245 |
+
+ monkeypatch.setattr(sys, "argv", ["mlx_lm.server"])
|
| 246 |
+
+ monkeypatch.setattr(server, "maybe_set_recommended_wired_limit", lambda: None)
|
| 247 |
+
+ monkeypatch.setattr(server, "run", lambda h, p, m: captured.append(m.cli_args))
|
| 248 |
+
+ server.main()
|
| 249 |
+
+ args = captured[0]
|
| 250 |
+
+ assert args.prompt_cache_size == 2
|
| 251 |
+
+ assert args.prompt_cache_bytes is None
|
| 252 |
+
+ assert args.prompt_concurrency == args.decode_concurrency == 1
|
| 253 |
+
+ assert args.prefill_step_size == 512
|