Instructions to use ukisai/Swift-1.5-4bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ukisai/Swift-1.5-4bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ukisai/Swift-1.5-4bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ukisai/Swift-1.5-4bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-1.5-4bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ukisai/Swift-1.5-4bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ukisai/Swift-1.5-4bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-1.5-4bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ukisai/Swift-1.5-4bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-1.5-4bit-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-1.5-4bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-1.5-4bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Clarify required runtime and add download diagnostics
Browse filesDocument the reproduced unpatched-loader failure, add a bounded-memory download and runtime checker, and explain how to obtain the original generation-worker error. Preserve all weights and architecture patches. Full-model Mac execution and the reported third-party failure remain unverified.
- README.md +13 -2
- TROUBLESHOOTING.md +79 -0
- UPLOAD_MANIFEST.json +17 -6
- USAGE.md +12 -0
- check_download.py +188 -0
README.md
CHANGED
|
@@ -34,6 +34,17 @@ converted with the official Apple MLX-LM converter using 4 bits and group size 6
|
|
| 34 |
Swift 1.5 is UkisAI's reasoning-efficient Qwen3.8-27B derivative, focused on stronger
|
| 35 |
long-horizon, agentic and coding performance while using fewer thinking tokens.
|
| 36 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
Swift 1.5 uses **58.5% fewer thinking tokens** than base Qwen3.8-27B while scoring **0.35% higher**, for a **9.18× speed-up** on several tasks.
|
| 38 |
|
| 39 |
## Download the complete model
|
|
@@ -60,8 +71,8 @@ model. Those test files are now archived under `compatibility/mac-check/fixtures
|
|
| 60 |
to keep them out of model-file discovery. Model weights are unchanged.
|
| 61 |
|
| 62 |
**Runtime requirement:** use the included MLX-LM patch and [loading instructions](USAGE.md).
|
| 63 |
-
|
| 64 |
-
|
| 65 |
validated; the weights require additional memory for the runtime, cache and macOS.
|
| 66 |
|
| 67 |
## Demo
|
|
|
|
| 34 |
Swift 1.5 is UkisAI's reasoning-efficient Qwen3.8-27B derivative, focused on stronger
|
| 35 |
long-horizon, agentic and coding performance while using fewer thinking tokens.
|
| 36 |
|
| 37 |
+
**Runtime compatibility:** this complete checkpoint requires the supplied architecture patch
|
| 38 |
+
and the pinned Python installation in [USAGE.md](USAGE.md). The tested unpatched
|
| 39 |
+
MLX-LM 0.32.0 loader rejects 501 saved vision entries. Downloading the model does
|
| 40 |
+
not install these patches into an app's inference engine. GUI compatibility and
|
| 41 |
+
full 27B Apple Silicon generation remain unverified.
|
| 42 |
+
|
| 43 |
+
The complete weights total **15.83 GB (14.74 GiB) across 3 required shards**, plus
|
| 44 |
+
config, index and tokenizer files. A single 5–6 GB file is not the complete model.
|
| 45 |
+
For download checks or `404 generation thread died`, use
|
| 46 |
+
[TROUBLESHOOTING.md](TROUBLESHOOTING.md).
|
| 47 |
+
|
| 48 |
Swift 1.5 uses **58.5% fewer thinking tokens** than base Qwen3.8-27B while scoring **0.35% higher**, for a **9.18× speed-up** on several tasks.
|
| 49 |
|
| 50 |
## Download the complete model
|
|
|
|
| 71 |
to keep them out of model-file discovery. Model weights are unchanged.
|
| 72 |
|
| 73 |
**Runtime requirement:** use the included MLX-LM patch and [loading instructions](USAGE.md).
|
| 74 |
+
The tested stock MLX-LM 0.32.0 loader cannot load this complete checkpoint;
|
| 75 |
+
GUI integrations must supply a compatible loader and have not been validated. Full-model generation on a 24 GB Mac has not been
|
| 76 |
validated; the weights require additional memory for the runtime, cache and macOS.
|
| 77 |
|
| 78 |
## Demo
|
TROUBLESHOOTING.md
ADDED
|
@@ -0,0 +1,79 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Check this download and its runtime
|
| 2 |
+
|
| 3 |
+
This release needs **all 3 shards (15.83 GB (14.74 GiB))**, its index/config/tokenizer
|
| 4 |
+
assets, and the supplied architecture patch. Downloading an HF repository does not install
|
| 5 |
+
the MLX-LM patches into Python or into a GUI app's separate runtime.
|
| 6 |
+
|
| 7 |
+
## 1. Verify the files without loading the model
|
| 8 |
+
|
| 9 |
+
Run from the directory containing `Swift-1.5-4bit-MLX`:
|
| 10 |
+
|
| 11 |
+
```bash
|
| 12 |
+
python Swift-1.5-4bit-MLX/check_download.py Swift-1.5-4bit-MLX --hash
|
| 13 |
+
```
|
| 14 |
+
|
| 15 |
+
This uses standard Python, reads the headers and stream-checks the original shard
|
| 16 |
+
SHA-256 hashes with bounded memory. It does not load the model or contact a server.
|
| 17 |
+
A successful result verifies the listed download checks, not inference.
|
| 18 |
+
|
| 19 |
+
If files are missing or incomplete, resume the full download using the pinned
|
| 20 |
+
snapshot instructions in [USAGE.md](USAGE.md), then verify again. Keep all shards
|
| 21 |
+
in one directory alongside the index, config and tokenizer. Do not load a single
|
| 22 |
+
shard or a diagnostic sample as a standalone model. Do not combine 4-bit and 5-bit
|
| 23 |
+
files. The complete 4-bit payload is 15.83 GB; the complete 5-bit payload is 19.28 GB.
|
| 24 |
+
A reported 5.7 GB download is insufficient for either complete checkpoint, although
|
| 25 |
+
that number alone does not tell us what an app has downloaded or is displaying.
|
| 26 |
+
|
| 27 |
+
## 2. Verify the runtime
|
| 28 |
+
|
| 29 |
+
Follow [USAGE.md](USAGE.md) to install the supplied architecture patch in the pinned Python
|
| 30 |
+
environment. Then, in that same environment:
|
| 31 |
+
|
| 32 |
+
```bash
|
| 33 |
+
python Swift-1.5-4bit-MLX/check_download.py Swift-1.5-4bit-MLX --runtime
|
| 34 |
+
```
|
| 35 |
+
|
| 36 |
+
The checker confirms that MLX-LM dispatches to `qwen3_5_full` and that its architecture
|
| 37 |
+
source matches the audited supplied patch. It records the MLX and MLX-LM versions.
|
| 38 |
+
It deliberately rejects other architecture code until checked separately.
|
| 39 |
+
For 5-bit, the base architecture patch alone is insufficient: `enable-5bit.patch`
|
| 40 |
+
must also be applied. A passing check in a terminal does not change an app's
|
| 41 |
+
separate runtime. GUI integration has not been validated for these full checkpoints.
|
| 42 |
+
|
| 43 |
+
Do not remove vision/MTP parameters or change the config to force a stock text
|
| 44 |
+
loader to accept the file. All 2,379 saved tensor entries belong to this release.
|
| 45 |
+
|
| 46 |
+
## 3. If a server reports `404 generation thread died`
|
| 47 |
+
|
| 48 |
+
In the pinned official MLX-LM server, this message means its generation worker
|
| 49 |
+
failed. The HTTP 404 is not enough to diagnose a missing web page, a broken model,
|
| 50 |
+
or a plugin problem. The original exception is in the server's terminal/log, before
|
| 51 |
+
the later generic response.
|
| 52 |
+
|
| 53 |
+
After the file and runtime checks pass, run the short direct Python generation
|
| 54 |
+
example in [USAGE.md](USAGE.md) in the same environment, independently of any plugin.
|
| 55 |
+
Use a Mac with sufficient RAM for the complete weights, runtime, cache and macOS.
|
| 56 |
+
Full-model generation on a 24 GB Mac has not been verified. Neither full release
|
| 57 |
+
has been validated on the available 16 GiB test Mac; do not raise memory limits.
|
| 58 |
+
|
| 59 |
+
If it fails, retain the original exception and traceback. A useful report contains
|
| 60 |
+
the app/command and version, Mac RAM, model revision, checker output, and that first
|
| 61 |
+
exception. Do not include access tokens or private prompts. Repeating a request
|
| 62 |
+
against a dead worker will not recover it; restart only after addressing the cause.
|
| 63 |
+
|
| 64 |
+
## What was checked on 2026-09-25
|
| 65 |
+
|
| 66 |
+
- Both public releases contain all indexed shards and 2,379 matching tensor headers.
|
| 67 |
+
- The unpatched official MLX-LM 0.32.0 model at the revision pinned in USAGE rejects
|
| 68 |
+
501 saved vision entries and removes 31 saved MTP entries during sanitization.
|
| 69 |
+
The required patched parameter trees accept all 2,379 entries with no shape mismatch.
|
| 70 |
+
These comparisons use real headers and unevaluated zero fixtures, not full payloads.
|
| 71 |
+
- An incomplete checkpoint reproduced the exact HTTP 404 / `generation thread died`
|
| 72 |
+
response. That demonstrates one possible cause, not the cause of any specific report.
|
| 73 |
+
- Both patched 4-bit and 5-bit runtimes passed small synthetic HTTP tests for
|
| 74 |
+
non-thinking chat, default thinking format, seeded generation, and streaming with
|
| 75 |
+
a tool definition. These are architecture/server checks, not full Swift generation
|
| 76 |
+
or a test of the reported third-party plugin.
|
| 77 |
+
|
| 78 |
+
The reporter's original traceback is still needed to identify their specific failure.
|
| 79 |
+
These checks do not establish full-model Mac execution or GUI compatibility.
|
UPLOAD_MANIFEST.json
CHANGED
|
@@ -31,19 +31,29 @@
|
|
| 31 |
},
|
| 32 |
{
|
| 33 |
"path": "README.md",
|
| 34 |
-
"bytes":
|
| 35 |
-
"sha256": "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
},
|
| 37 |
{
|
| 38 |
"path": "USAGE.md",
|
| 39 |
-
"bytes":
|
| 40 |
-
"sha256": "
|
| 41 |
},
|
| 42 |
{
|
| 43 |
"path": "chat_template.jinja",
|
| 44 |
"bytes": 8952,
|
| 45 |
"sha256": "c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041"
|
| 46 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 47 |
{
|
| 48 |
"path": "compatibility/MLX-LM-LICENSE",
|
| 49 |
"bytes": 1066,
|
|
@@ -270,7 +280,8 @@
|
|
| 270 |
"sha256": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003"
|
| 271 |
}
|
| 272 |
],
|
| 273 |
-
"total_bytes_excluding_manifest":
|
| 274 |
"manifest_self_excluded": true,
|
| 275 |
-
"packaging_update": "2026-09-24: diagnostic fixtures archived; model weights unchanged"
|
|
|
|
| 276 |
}
|
|
|
|
| 31 |
},
|
| 32 |
{
|
| 33 |
"path": "README.md",
|
| 34 |
+
"bytes": 10074,
|
| 35 |
+
"sha256": "fc7a25233fe1f32a53318716b7c7a5a00818603cd72285ffae990958f9e3e030"
|
| 36 |
+
},
|
| 37 |
+
{
|
| 38 |
+
"path": "TROUBLESHOOTING.md",
|
| 39 |
+
"bytes": 4360,
|
| 40 |
+
"sha256": "6460abf4aa4d2de46f07ce8dfa095476467eb466c6250c727ba0d6f2c77399eb"
|
| 41 |
},
|
| 42 |
{
|
| 43 |
"path": "USAGE.md",
|
| 44 |
+
"bytes": 5924,
|
| 45 |
+
"sha256": "70e4b505d7acb75d33774e19b4b6202e5e45f0dcaf0f5973d6ea73f5d618a941"
|
| 46 |
},
|
| 47 |
{
|
| 48 |
"path": "chat_template.jinja",
|
| 49 |
"bytes": 8952,
|
| 50 |
"sha256": "c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041"
|
| 51 |
},
|
| 52 |
+
{
|
| 53 |
+
"path": "check_download.py",
|
| 54 |
+
"bytes": 9658,
|
| 55 |
+
"sha256": "1e424a32bb015431df22b8f058aa18de116ae63e0440d60424792a1cf2fd4018"
|
| 56 |
+
},
|
| 57 |
{
|
| 58 |
"path": "compatibility/MLX-LM-LICENSE",
|
| 59 |
"bytes": 1066,
|
|
|
|
| 280 |
"sha256": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003"
|
| 281 |
}
|
| 282 |
],
|
| 283 |
+
"total_bytes_excluding_manifest": 15859840103,
|
| 284 |
"manifest_self_excluded": true,
|
| 285 |
+
"packaging_update": "2026-09-24: diagnostic fixtures archived; model weights unchanged",
|
| 286 |
+
"diagnostic_update": "2026-09-25: clarify required runtime; add download/runtime checker; weights unchanged"
|
| 287 |
}
|
USAGE.md
CHANGED
|
@@ -21,6 +21,7 @@ python -m pip install 'huggingface_hub==1.31.0'
|
|
| 21 |
SWIFT_MLX_REVISION="$(python -c 'import json, urllib.request; print(json.load(urllib.request.urlopen("https://huggingface.co/api/models/ukisai/Swift-1.5-4bit-MLX"))["sha"])')"
|
| 22 |
hf download ukisai/Swift-1.5-4bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-4bit-MLX
|
| 23 |
hf cache verify ukisai/Swift-1.5-4bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-4bit-MLX --fail-on-missing-files
|
|
|
|
| 24 |
git clone https://github.com/ml-explore/mlx-lm.git swift15-mlx-lm
|
| 25 |
git -C swift15-mlx-lm checkout --detach c69d1288440a0dc4e6401fc417098b07598dccd5
|
| 26 |
git -C swift15-mlx-lm apply --check ../Swift-1.5-4bit-MLX/compatibility/swift15-mlx-lm.patch
|
|
@@ -50,6 +51,17 @@ pip install 'mlx[cpu]==0.32.2' 'transformers==5.14.1' 'huggingface_hub==1.31.0'
|
|
| 50 |
pip install -e ./swift15-mlx-lm
|
| 51 |
```
|
| 52 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 53 |
Text generation:
|
| 54 |
|
| 55 |
```python
|
|
|
|
| 21 |
SWIFT_MLX_REVISION="$(python -c 'import json, urllib.request; print(json.load(urllib.request.urlopen("https://huggingface.co/api/models/ukisai/Swift-1.5-4bit-MLX"))["sha"])')"
|
| 22 |
hf download ukisai/Swift-1.5-4bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-4bit-MLX
|
| 23 |
hf cache verify ukisai/Swift-1.5-4bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-4bit-MLX --fail-on-missing-files
|
| 24 |
+
python Swift-1.5-4bit-MLX/check_download.py Swift-1.5-4bit-MLX
|
| 25 |
git clone https://github.com/ml-explore/mlx-lm.git swift15-mlx-lm
|
| 26 |
git -C swift15-mlx-lm checkout --detach c69d1288440a0dc4e6401fc417098b07598dccd5
|
| 27 |
git -C swift15-mlx-lm apply --check ../Swift-1.5-4bit-MLX/compatibility/swift15-mlx-lm.patch
|
|
|
|
| 51 |
pip install -e ./swift15-mlx-lm
|
| 52 |
```
|
| 53 |
|
| 54 |
+
Before generating, check the Python environment that will actually run the model:
|
| 55 |
+
|
| 56 |
+
```bash
|
| 57 |
+
python Swift-1.5-4bit-MLX/check_download.py Swift-1.5-4bit-MLX --runtime
|
| 58 |
+
```
|
| 59 |
+
|
| 60 |
+
This checks the complete download and selects the pinned patched architecture
|
| 61 |
+
without loading weights. It does not check a separate GUI app's environment,
|
| 62 |
+
available inference memory, or generated quality. See [TROUBLESHOOTING.md](TROUBLESHOOTING.md)
|
| 63 |
+
if it fails or a server reports `generation thread died`.
|
| 64 |
+
|
| 65 |
Text generation:
|
| 66 |
|
| 67 |
```python
|
check_download.py
ADDED
|
@@ -0,0 +1,188 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Check a Swift 1.5 4/5-bit download and optional pinned Python runtime.
|
| 3 |
+
|
| 4 |
+
Does not download files, load weights, generate text, or certify GUI compatibility.
|
| 5 |
+
"""
|
| 6 |
+
import argparse
|
| 7 |
+
import hashlib
|
| 8 |
+
import json
|
| 9 |
+
from pathlib import Path
|
| 10 |
+
import struct
|
| 11 |
+
import sys
|
| 12 |
+
|
| 13 |
+
RELEASES = {4: {
|
| 14 |
+
'model-00001-of-00003.safetensors': (5328325554, 'd91b70ca84dff315addeeb8a599dce2beccfec33c7483686ede2e8ff3cc459dc'),
|
| 15 |
+
'model-00002-of-00003.safetensors': (5354185158, 'eddcff1a6ef0971990f7cd01dba77ab9a8c20d59821bb4cd7e25234fe3809346'),
|
| 16 |
+
'model-00003-of-00003.safetensors': (5144253923, '644cc5322ff706c9608b42ea1ff83e0beb12b828d49b36c468a366f44593b8eb'),
|
| 17 |
+
}, 5: {
|
| 18 |
+
'model-00001-of-00004.safetensors': (5360417859, '9da8750595ec42d8c5d2ff0ae0b695e387cbcb3c0beac40f262e6d9a7196b59d'),
|
| 19 |
+
'model-00002-of-00004.safetensors': (5338181777, 'b592723bd4e9f01344335cd7d4b4f41a7c8cde98e2e074ad1d943cdb951a6835'),
|
| 20 |
+
'model-00003-of-00004.safetensors': (5355925620, '0027e7fe05310d5b2d960414221ed0872de9b511aa7e6b1c3a8f0e4647558412'),
|
| 21 |
+
'model-00004-of-00004.safetensors': (3227578201, 'afec577f16805b327471561d6841662eb4acf7d866ad5e761c5ef8cd3149d8f4'),
|
| 22 |
+
}}
|
| 23 |
+
EXPECTED_TENSORS = 2379
|
| 24 |
+
MODEL_HASHES = {
|
| 25 |
+
4: {'f559e4559946bfad4ce20b93906028449b76e007f42f36bb83d2c2c5f6cb95a9',
|
| 26 |
+
'ecda666d6d5f1059388a2ce4267d9ff8625398bb925d6c0840d0059ce17b676f'},
|
| 27 |
+
5: {'ecda666d6d5f1059388a2ce4267d9ff8625398bb925d6c0840d0059ce17b676f'},
|
| 28 |
+
}
|
| 29 |
+
REQUIRED = (
|
| 30 |
+
'config.json', 'model.safetensors.index.json', 'tokenizer.json',
|
| 31 |
+
'tokenizer_config.json', 'chat_template.jinja', 'generation_config.json',
|
| 32 |
+
'vocab.json', 'merges.txt', 'preprocessor_config.json', 'video_preprocessor_config.json',
|
| 33 |
+
)
|
| 34 |
+
|
| 35 |
+
def unique(pairs):
|
| 36 |
+
result = {}
|
| 37 |
+
for key, value in pairs:
|
| 38 |
+
if key in result:
|
| 39 |
+
raise ValueError('Duplicate JSON key')
|
| 40 |
+
result[key] = value
|
| 41 |
+
return result
|
| 42 |
+
|
| 43 |
+
def read_json(path):
|
| 44 |
+
return json.loads(path.read_text(encoding='utf-8'), object_pairs_hook=unique)
|
| 45 |
+
|
| 46 |
+
def runtime_check(config):
|
| 47 |
+
"""Check dispatch and the audited architecture source, without model creation."""
|
| 48 |
+
import inspect
|
| 49 |
+
from importlib.metadata import version
|
| 50 |
+
from mlx_lm.utils import _get_classes
|
| 51 |
+
cls, args = _get_classes(config)
|
| 52 |
+
args.from_dict(config)
|
| 53 |
+
row = {'mlx': version('mlx'), 'mlx_lm': version('mlx-lm'),
|
| 54 |
+
'selected_model_class': cls.__module__+'.'+cls.__name__}
|
| 55 |
+
if cls.__module__ != 'mlx_lm.models.qwen3_5_full':
|
| 56 |
+
raise ValueError('The required full-model MLX-LM patch is not active. Follow USAGE.md in this Python environment.')
|
| 57 |
+
row['model_source_sha256'] = hashlib.sha256(Path(inspect.getsourcefile(cls)).read_bytes()).hexdigest()
|
| 58 |
+
if row['model_source_sha256'] not in MODEL_HASHES[config['quantization']['bits']]:
|
| 59 |
+
raise ValueError('This is not the pinned patched model implementation. For 5-bit, both patches are required. Follow USAGE.md.')
|
| 60 |
+
return row
|
| 61 |
+
|
| 62 |
+
def check(folder, check_hashes=False, check_runtime=False):
|
| 63 |
+
folder = Path(folder)
|
| 64 |
+
errors = []
|
| 65 |
+
report = {'files':[], 'weights_loaded':False, 'generation_tested':False,
|
| 66 |
+
'scope':'File sizes, headers, index and requested checks only; no full-model or GUI test'}
|
| 67 |
+
try:
|
| 68 |
+
config = read_json(folder/'config.json')
|
| 69 |
+
bits = config['quantization']['bits']
|
| 70 |
+
if type(bits) is not int or bits not in RELEASES:
|
| 71 |
+
raise ValueError('Expected 4-bit or 5-bit Swift release')
|
| 72 |
+
expected = RELEASES[bits]
|
| 73 |
+
except (OSError,ValueError,KeyError,TypeError) as exc:
|
| 74 |
+
return dict(report,status='FAIL',errors=[f'Cannot identify the release: {exc}'])
|
| 75 |
+
report.update(bits=bits,expected_shards=len(expected),expected_weight_bytes=sum(v[0] for v in expected.values()))
|
| 76 |
+
for name in REQUIRED:
|
| 77 |
+
path = folder/name
|
| 78 |
+
if not path.is_file():
|
| 79 |
+
errors.append(f'Missing required asset: {name}')
|
| 80 |
+
elif path.stat().st_size == 0:
|
| 81 |
+
errors.append(f'Empty required asset: {name}')
|
| 82 |
+
elif name.endswith('.json'):
|
| 83 |
+
try:
|
| 84 |
+
read_json(path)
|
| 85 |
+
except (ValueError, OSError) as exc:
|
| 86 |
+
errors.append(f'Cannot parse {name}: {type(exc).__name__}')
|
| 87 |
+
seen = {}
|
| 88 |
+
sizes = {'U32':4, 'BF16':2, 'F16':2, 'F32':4}
|
| 89 |
+
for name,(size,sha256) in expected.items():
|
| 90 |
+
path = folder/name
|
| 91 |
+
actual = path.stat().st_size if path.is_file() else None
|
| 92 |
+
row = {'file':name,'expected_bytes':size,'actual_bytes':actual}
|
| 93 |
+
report['files'].append(row)
|
| 94 |
+
if actual != size:
|
| 95 |
+
errors.append(f'{name}: expected {size} bytes, found {actual if actual is not None else "missing"}')
|
| 96 |
+
continue
|
| 97 |
+
try:
|
| 98 |
+
with path.open('rb') as stream:
|
| 99 |
+
length = struct.unpack('<Q',stream.read(8))[0]
|
| 100 |
+
if not 0 < length <= 16*1024*1024:
|
| 101 |
+
raise ValueError('Invalid safetensors header size')
|
| 102 |
+
headers = json.loads(stream.read(length),object_pairs_hook=unique)
|
| 103 |
+
if not isinstance(headers,dict):
|
| 104 |
+
raise ValueError('Invalid safetensors header object')
|
| 105 |
+
headers.pop('__metadata__',None)
|
| 106 |
+
ranges = []
|
| 107 |
+
for tensor,entry in headers.items():
|
| 108 |
+
if tensor in seen:
|
| 109 |
+
raise ValueError('Duplicate tensor across shards')
|
| 110 |
+
start,end = entry['data_offsets']
|
| 111 |
+
shape = entry['shape']
|
| 112 |
+
if not isinstance(shape,list) or type(start) is not int or type(end) is not int:
|
| 113 |
+
raise ValueError('Invalid shape or offset types')
|
| 114 |
+
count = 1
|
| 115 |
+
for dim in shape:
|
| 116 |
+
if type(dim) is not int or dim < 0:
|
| 117 |
+
raise ValueError('Invalid tensor shape')
|
| 118 |
+
count *= dim
|
| 119 |
+
if entry['dtype'] not in sizes or end-start != count*sizes[entry['dtype']]:
|
| 120 |
+
raise ValueError('Invalid tensor size or dtype')
|
| 121 |
+
if not 0 <= start <= end <= actual-8-length:
|
| 122 |
+
raise ValueError('Invalid tensor data offsets')
|
| 123 |
+
seen[tensor] = name
|
| 124 |
+
ranges.append((start,end))
|
| 125 |
+
offset = 0
|
| 126 |
+
for start,end in sorted(ranges):
|
| 127 |
+
if start != offset:
|
| 128 |
+
raise ValueError('Overlapping tensor data or a payload gap')
|
| 129 |
+
offset = end
|
| 130 |
+
if offset != actual-8-length:
|
| 131 |
+
raise ValueError('Tensor payload does not fill shard')
|
| 132 |
+
row['header_check'] = 'PASS'
|
| 133 |
+
if check_hashes:
|
| 134 |
+
digest = hashlib.sha256()
|
| 135 |
+
with path.open('rb') as stream:
|
| 136 |
+
for block in iter(lambda:stream.read(8*1024*1024),b''):
|
| 137 |
+
digest.update(block)
|
| 138 |
+
row['sha256'] = digest.hexdigest()
|
| 139 |
+
row['sha256_check'] = 'PASS' if row['sha256'] == sha256 else 'FAIL'
|
| 140 |
+
if row['sha256'] != sha256:
|
| 141 |
+
errors.append(f'SHA256 mismatch: {name}')
|
| 142 |
+
except (ValueError,KeyError,TypeError,OSError,struct.error) as exc:
|
| 143 |
+
errors.append(f'{name}: {exc}')
|
| 144 |
+
extra = sorted(p.name for p in folder.glob('*.safetensors') if p.name not in expected)
|
| 145 |
+
if extra:
|
| 146 |
+
errors.append(f'Unexpected root weight files: {extra}')
|
| 147 |
+
try:
|
| 148 |
+
index = read_json(folder/'model.safetensors.index.json')
|
| 149 |
+
if not isinstance(index['weight_map'],dict):
|
| 150 |
+
raise ValueError('Invalid weight map')
|
| 151 |
+
if set(index['weight_map'].values()) != set(expected):
|
| 152 |
+
errors.append(f'Index must reference exactly the {len(expected)} expected shards')
|
| 153 |
+
if seen != index['weight_map']:
|
| 154 |
+
errors.append('Available tensor headers do not match the complete index')
|
| 155 |
+
if len(index['weight_map']) != EXPECTED_TENSORS:
|
| 156 |
+
errors.append(f'Expected {EXPECTED_TENSORS} saved tensor entries')
|
| 157 |
+
except (OSError,ValueError,KeyError,TypeError):
|
| 158 |
+
errors.append('Cannot validate the checkpoint index')
|
| 159 |
+
try:
|
| 160 |
+
if config.get('quantization') != {'bits':bits,'group_size':64,'mode':'affine'}:
|
| 161 |
+
errors.append(f'Expected affine / {bits}-bit / group size 64')
|
| 162 |
+
if config.get('language_model_only') is not False:
|
| 163 |
+
errors.append('Expected the complete Swift architecture configuration')
|
| 164 |
+
if check_runtime:
|
| 165 |
+
try:
|
| 166 |
+
report['runtime'] = runtime_check(config)
|
| 167 |
+
except Exception as exc:
|
| 168 |
+
errors.append(f'Runtime check failed: {type(exc).__name__}: {exc}')
|
| 169 |
+
except (OSError,ValueError,KeyError,TypeError):
|
| 170 |
+
errors.append('Cannot validate model configuration')
|
| 171 |
+
report['hash_check'] = ('PASS' if all(row.get('sha256_check')=='PASS' for row in report['files']) else 'FAIL_OR_INCOMPLETE') if check_hashes else 'not run; add --hash to read and SHA256-check all weight bytes'
|
| 172 |
+
report['runtime_check_requested'] = check_runtime
|
| 173 |
+
report['errors'] = errors
|
| 174 |
+
report['status'] = 'FAIL' if errors else 'PASS_DOWNLOAD_AND_REQUESTED_CHECKS_ONLY'
|
| 175 |
+
return report
|
| 176 |
+
|
| 177 |
+
def main():
|
| 178 |
+
parser = argparse.ArgumentParser(description=__doc__)
|
| 179 |
+
parser.add_argument('folder',type=Path)
|
| 180 |
+
parser.add_argument('--hash',action='store_true',help='Stream-read every shard and compare its original SHA256')
|
| 181 |
+
parser.add_argument('--runtime',action='store_true',help='Check the installed Python MLX-LM dispatcher; does not instantiate the model')
|
| 182 |
+
args = parser.parse_args()
|
| 183 |
+
result = check(args.folder,args.hash,args.runtime)
|
| 184 |
+
print(json.dumps(result,indent=2))
|
| 185 |
+
return 1 if result['errors'] else 0
|
| 186 |
+
|
| 187 |
+
if __name__ == '__main__':
|
| 188 |
+
sys.exit(main())
|