Bound retained server cache while keeping prompt reuse enabled

#1
by ukisai - opened

The pinned server constructed its retained-cache LRU without the configured byte limit, so entries saved during prefill or completion could exceed --prompt-cache-bytes. The patch enforces the limit on insertion and keeps prompt reuse enabled with bounded defaults: two retained entries, one prompt and decode stream, and 512-token prefill steps. Metal memory estimates set the default retained-cache budget; an explicit byte cap overrides it.

Both complete, existing Swift 4-bit and 5-bit checkpoints now passed independent native Metal validation on an AWS M4 Pro with 48 GiB RAM. Each processed our own synthetic 86k-token conversation, generated a reply, and completed two follow-ups with most history reused from cache. Hub offline mode was enabled. The full report and reproduction instructions include measured results and limitations.

Checks passed:

  • Full checkpoint SHA-256 verification, runtime identity, three successful HTTP requests per format, retained-cache budget enforcement and substantial warm reuse.
  • 44 unit/upstream server tests per clean architecture patch stack.
  • 36 tiny-fixture offline HTTP requests per format, including streaming, sequential generation, disabled retention and explicit byte limits.
  • Clean patch application and rollback, Black 25.1.0, isort 6.0.0 and Ruff 0.16.6.

These are independently executed checks, not hosted CI badges. The tested workload does not establish arbitrary context lengths, smaller Mac memory sizes, GUI/plugin integration, image/video chat or speculative MTP support. The retention cap excludes active-request and total process memory.

Weights, quantization, model semantics and tokenizer assets are unchanged. Existing installations must apply the runtime patch and restart. This PR is merged into public main at 8aff72b145212e62c15146e41dcbef35ea5fa9ff. The published tree and unchanged checkpoint hashes were verified after merging.

Retrieve only the runtime update

With the complete model and dependencies already local, fetch these three small files:

hf download ukisai/Swift-1.5-5bit-MLX \
  compatibility/swift15-server-cache.patch \
  compatibility/cache-tests/validation.json \
  compatibility/cache-tests/verify_server_patch.py \
  --revision 80896f29f71f3f6b8131f4322e9a70262e8e7b51 --local-dir Swift-1.5-5bit-MLX

Then follow the patch, restart and offline-startup instructions. The command does not download any weight shards.

ukisai changed pull request status to merged

Sign up or log in to comment