Instructions to use ukisai/Swift-1.5-5bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ukisai/Swift-1.5-5bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ukisai/Swift-1.5-5bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ukisai/Swift-1.5-5bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-1.5-5bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ukisai/Swift-1.5-5bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ukisai/Swift-1.5-5bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-1.5-5bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ukisai/Swift-1.5-5bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-1.5-5bit-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-1.5-5bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-5bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-1.5-5bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Clarify required runtime and add download diagnostics
Browse filesDocument the reproduced unpatched-loader failure, add a bounded-memory download and runtime checker, and explain how to obtain the original generation-worker error. Preserve all weights and architecture patches. Full-model Mac execution and the reported third-party failure remain unverified.
- README.md +24 -0
- TROUBLESHOOTING.md +79 -0
- UPLOAD_MANIFEST.json +20 -8
- USAGE.md +13 -2
- check_download.py +188 -0
README.md
CHANGED
|
@@ -32,6 +32,17 @@ reasoning-efficient Qwen3.8-27B derivative, focused on long-horizon, agentic and
|
|
| 32 |
coding tasks. This export preserves the text, vision and MTP parameter tree;
|
| 33 |
its supported generation interface is text-only with the included MLX-LM patches.
|
| 34 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
Swift 1.5 uses **58.5% fewer thinking tokens** than base Qwen3.8-27B while scoring **0.35% higher**, for a **9.18× speed-up** on several tasks.
|
| 36 |
|
| 37 |
> [!CAUTION]
|
|
@@ -40,6 +51,19 @@ Swift 1.5 uses **58.5% fewer thinking tokens** than base Qwen3.8-27B while scori
|
|
| 40 |
> Full independent Apple Silicon generation and quality evaluation remain **NOT_RUN**.
|
| 41 |
> The 19.28 GB tensor payload must not be forced onto a 16 GiB Mac.
|
| 42 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
## Demo
|
| 44 |
|
| 45 |
We gave base Qwen3.8-27B and Swift 1.5 27B the same prompt:
|
|
|
|
| 32 |
coding tasks. This export preserves the text, vision and MTP parameter tree;
|
| 33 |
its supported generation interface is text-only with the included MLX-LM patches.
|
| 34 |
|
| 35 |
+
**Runtime compatibility:** this complete checkpoint requires both supplied patches, in the order shown
|
| 36 |
+
and the pinned Python installation in [USAGE.md](USAGE.md). The tested unpatched
|
| 37 |
+
MLX-LM 0.32.0 loader rejects 501 saved vision entries. Downloading the model does
|
| 38 |
+
not install these patches into an app's inference engine. GUI compatibility and
|
| 39 |
+
full 27B Apple Silicon generation remain unverified.
|
| 40 |
+
|
| 41 |
+
The complete weights total **19.28 GB (17.96 GiB) across 4 required shards**, plus
|
| 42 |
+
config, index and tokenizer files. A single 5–6 GB file is not the complete model.
|
| 43 |
+
For download checks or `404 generation thread died`, use
|
| 44 |
+
[TROUBLESHOOTING.md](TROUBLESHOOTING.md).
|
| 45 |
+
|
| 46 |
Swift 1.5 uses **58.5% fewer thinking tokens** than base Qwen3.8-27B while scoring **0.35% higher**, for a **9.18× speed-up** on several tasks.
|
| 47 |
|
| 48 |
> [!CAUTION]
|
|
|
|
| 51 |
> Full independent Apple Silicon generation and quality evaluation remain **NOT_RUN**.
|
| 52 |
> The 19.28 GB tensor payload must not be forced onto a 16 GiB Mac.
|
| 53 |
|
| 54 |
+
## Download the complete model
|
| 55 |
+
|
| 56 |
+
This repository is public; authentication is optional. Use a complete download:
|
| 57 |
+
|
| 58 |
+
```bash
|
| 59 |
+
hf download ukisai/Swift-1.5-5bit-MLX --local-dir Swift-1.5-5bit-MLX
|
| 60 |
+
python Swift-1.5-5bit-MLX/check_download.py Swift-1.5-5bit-MLX --hash
|
| 61 |
+
```
|
| 62 |
+
|
| 63 |
+
All four `model-0000*-of-00004.safetensors` files are required. Follow [USAGE.md](USAGE.md)
|
| 64 |
+
to install both runtime patches before attempting generation. A download check
|
| 65 |
+
does not establish successful inference or sufficient RAM.
|
| 66 |
+
|
| 67 |
## Demo
|
| 68 |
|
| 69 |
We gave base Qwen3.8-27B and Swift 1.5 27B the same prompt:
|
TROUBLESHOOTING.md
ADDED
|
@@ -0,0 +1,79 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Check this download and its runtime
|
| 2 |
+
|
| 3 |
+
This release needs **all 4 shards (19.28 GB (17.96 GiB))**, its index/config/tokenizer
|
| 4 |
+
assets, and both supplied patches, in the order shown. Downloading an HF repository does not install
|
| 5 |
+
the MLX-LM patches into Python or into a GUI app's separate runtime.
|
| 6 |
+
|
| 7 |
+
## 1. Verify the files without loading the model
|
| 8 |
+
|
| 9 |
+
Run from the directory containing `Swift-1.5-5bit-MLX`:
|
| 10 |
+
|
| 11 |
+
```bash
|
| 12 |
+
python Swift-1.5-5bit-MLX/check_download.py Swift-1.5-5bit-MLX --hash
|
| 13 |
+
```
|
| 14 |
+
|
| 15 |
+
This uses standard Python, reads the headers and stream-checks the original shard
|
| 16 |
+
SHA-256 hashes with bounded memory. It does not load the model or contact a server.
|
| 17 |
+
A successful result verifies the listed download checks, not inference.
|
| 18 |
+
|
| 19 |
+
If files are missing or incomplete, resume the full download using the pinned
|
| 20 |
+
snapshot instructions in [USAGE.md](USAGE.md), then verify again. Keep all shards
|
| 21 |
+
in one directory alongside the index, config and tokenizer. Do not load a single
|
| 22 |
+
shard or a diagnostic sample as a standalone model. Do not combine 4-bit and 5-bit
|
| 23 |
+
files. The complete 4-bit payload is 15.83 GB; the complete 5-bit payload is 19.28 GB.
|
| 24 |
+
A reported 5.7 GB download is insufficient for either complete checkpoint, although
|
| 25 |
+
that number alone does not tell us what an app has downloaded or is displaying.
|
| 26 |
+
|
| 27 |
+
## 2. Verify the runtime
|
| 28 |
+
|
| 29 |
+
Follow [USAGE.md](USAGE.md) to install both supplied patches, in the order shown in the pinned Python
|
| 30 |
+
environment. Then, in that same environment:
|
| 31 |
+
|
| 32 |
+
```bash
|
| 33 |
+
python Swift-1.5-5bit-MLX/check_download.py Swift-1.5-5bit-MLX --runtime
|
| 34 |
+
```
|
| 35 |
+
|
| 36 |
+
The checker confirms that MLX-LM dispatches to `qwen3_5_full` and that its architecture
|
| 37 |
+
source matches the audited supplied patch. It records the MLX and MLX-LM versions.
|
| 38 |
+
It deliberately rejects other architecture code until checked separately.
|
| 39 |
+
For 5-bit, the base architecture patch alone is insufficient: `enable-5bit.patch`
|
| 40 |
+
must also be applied. A passing check in a terminal does not change an app's
|
| 41 |
+
separate runtime. GUI integration has not been validated for these full checkpoints.
|
| 42 |
+
|
| 43 |
+
Do not remove vision/MTP parameters or change the config to force a stock text
|
| 44 |
+
loader to accept the file. All 2,379 saved tensor entries belong to this release.
|
| 45 |
+
|
| 46 |
+
## 3. If a server reports `404 generation thread died`
|
| 47 |
+
|
| 48 |
+
In the pinned official MLX-LM server, this message means its generation worker
|
| 49 |
+
failed. The HTTP 404 is not enough to diagnose a missing web page, a broken model,
|
| 50 |
+
or a plugin problem. The original exception is in the server's terminal/log, before
|
| 51 |
+
the later generic response.
|
| 52 |
+
|
| 53 |
+
After the file and runtime checks pass, run the short direct Python generation
|
| 54 |
+
example in [USAGE.md](USAGE.md) in the same environment, independently of any plugin.
|
| 55 |
+
Use a Mac with sufficient RAM for the complete weights, runtime, cache and macOS.
|
| 56 |
+
Full-model generation on a 24 GB Mac has not been verified. Neither full release
|
| 57 |
+
has been validated on the available 16 GiB test Mac; do not raise memory limits.
|
| 58 |
+
|
| 59 |
+
If it fails, retain the original exception and traceback. A useful report contains
|
| 60 |
+
the app/command and version, Mac RAM, model revision, checker output, and that first
|
| 61 |
+
exception. Do not include access tokens or private prompts. Repeating a request
|
| 62 |
+
against a dead worker will not recover it; restart only after addressing the cause.
|
| 63 |
+
|
| 64 |
+
## What was checked on 2026-09-25
|
| 65 |
+
|
| 66 |
+
- Both public releases contain all indexed shards and 2,379 matching tensor headers.
|
| 67 |
+
- The unpatched official MLX-LM 0.32.0 model at the revision pinned in USAGE rejects
|
| 68 |
+
501 saved vision entries and removes 31 saved MTP entries during sanitization.
|
| 69 |
+
The required patched parameter trees accept all 2,379 entries with no shape mismatch.
|
| 70 |
+
These comparisons use real headers and unevaluated zero fixtures, not full payloads.
|
| 71 |
+
- An incomplete checkpoint reproduced the exact HTTP 404 / `generation thread died`
|
| 72 |
+
response. That demonstrates one possible cause, not the cause of any specific report.
|
| 73 |
+
- Both patched 4-bit and 5-bit runtimes passed small synthetic HTTP tests for
|
| 74 |
+
non-thinking chat, default thinking format, seeded generation, and streaming with
|
| 75 |
+
a tool definition. These are architecture/server checks, not full Swift generation
|
| 76 |
+
or a test of the reported third-party plugin.
|
| 77 |
+
|
| 78 |
+
The reporter's original traceback is still needed to identify their specific failure.
|
| 79 |
+
These checks do not establish full-model Mac execution or GUI compatibility.
|
UPLOAD_MANIFEST.json
CHANGED
|
@@ -32,15 +32,21 @@
|
|
| 32 |
},
|
| 33 |
{
|
| 34 |
"path": "README.md",
|
| 35 |
-
"bytes":
|
| 36 |
-
"sha256": "
|
| 37 |
-
"git_blob_sha1": "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
},
|
| 39 |
{
|
| 40 |
"path": "USAGE.md",
|
| 41 |
-
"bytes":
|
| 42 |
-
"sha256": "
|
| 43 |
-
"git_blob_sha1": "
|
| 44 |
},
|
| 45 |
{
|
| 46 |
"path": "chat_template.jinja",
|
|
@@ -48,6 +54,12 @@
|
|
| 48 |
"sha256": "c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041",
|
| 49 |
"git_blob_sha1": "c0c686f9c38d70d179fb7b5f5aa7530bc913dda3"
|
| 50 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 51 |
{
|
| 52 |
"path": "compatibility/LICENSE-MLX-LM-MIT",
|
| 53 |
"bytes": 1066,
|
|
@@ -204,7 +216,7 @@
|
|
| 204 |
"git_blob_sha1": "0aa0ce0658d60ac4a5d609f4eadb0e8e43514176"
|
| 205 |
}
|
| 206 |
],
|
| 207 |
-
"file_count":
|
| 208 |
-
"total_bytes":
|
| 209 |
"self_excluded": true
|
| 210 |
}
|
|
|
|
| 32 |
},
|
| 33 |
{
|
| 34 |
"path": "README.md",
|
| 35 |
+
"bytes": 8067,
|
| 36 |
+
"sha256": "e0ce3327646a3dde90ee001a3bb819c8cb24e550a83d28340671f2542e7ad55b",
|
| 37 |
+
"git_blob_sha1": "008e4ab150bc188e372cad12fda89317afda850e"
|
| 38 |
+
},
|
| 39 |
+
{
|
| 40 |
+
"path": "TROUBLESHOOTING.md",
|
| 41 |
+
"bytes": 4380,
|
| 42 |
+
"sha256": "32b67c70616257a0e26e91bbc6b1b880391872f98c4292aece9488940cda9d00",
|
| 43 |
+
"git_blob_sha1": "2968859f3f32fc644fc20c2fb9e9644e10535e43"
|
| 44 |
},
|
| 45 |
{
|
| 46 |
"path": "USAGE.md",
|
| 47 |
+
"bytes": 5431,
|
| 48 |
+
"sha256": "be75f4abc1868a29866013d9fd6490644f6c06cf3effe6e2704f3e6657affddb",
|
| 49 |
+
"git_blob_sha1": "db137784896c403a7ee030bb56f335503227baf6"
|
| 50 |
},
|
| 51 |
{
|
| 52 |
"path": "chat_template.jinja",
|
|
|
|
| 54 |
"sha256": "c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041",
|
| 55 |
"git_blob_sha1": "c0c686f9c38d70d179fb7b5f5aa7530bc913dda3"
|
| 56 |
},
|
| 57 |
+
{
|
| 58 |
+
"path": "check_download.py",
|
| 59 |
+
"bytes": 9658,
|
| 60 |
+
"sha256": "1e424a32bb015431df22b8f058aa18de116ae63e0440d60424792a1cf2fd4018",
|
| 61 |
+
"git_blob_sha1": "57c4bddd7813e480c7a738ea997e2bb4e3cc51ff"
|
| 62 |
+
},
|
| 63 |
{
|
| 64 |
"path": "compatibility/LICENSE-MLX-LM-MIT",
|
| 65 |
"bytes": 1066,
|
|
|
|
| 216 |
"git_blob_sha1": "0aa0ce0658d60ac4a5d609f4eadb0e8e43514176"
|
| 217 |
}
|
| 218 |
],
|
| 219 |
+
"file_count": 37,
|
| 220 |
+
"total_bytes": 19310877524,
|
| 221 |
"self_excluded": true
|
| 222 |
}
|
USAGE.md
CHANGED
|
@@ -7,8 +7,7 @@ or raise system memory limits to conceal insufficient hardware.
|
|
| 7 |
|
| 8 |
## Pin the snapshot and verify it
|
| 9 |
|
| 10 |
-
Use a new working directory.
|
| 11 |
-
interactively if not already signed in with access to this private repository.
|
| 12 |
The command below resolves current main once to a full commit and then uses
|
| 13 |
only that pinned snapshot. For a repeat run, reuse the recorded commit.
|
| 14 |
Do not use an incomplete historical upload or proceed after verification failure.
|
|
@@ -21,6 +20,7 @@ SWIFT_MLX_REVISION="$(python -c 'from huggingface_hub import HfApi; print(HfApi(
|
|
| 21 |
printf 'Pinned model revision: %s\n' "$SWIFT_MLX_REVISION"
|
| 22 |
hf download ukisai/Swift-1.5-5bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-5bit-MLX
|
| 23 |
hf cache verify ukisai/Swift-1.5-5bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-5bit-MLX --fail-on-missing-files
|
|
|
|
| 24 |
python Swift-1.5-5bit-MLX/verify_release.py Swift-1.5-5bit-MLX
|
| 25 |
```
|
| 26 |
|
|
@@ -60,6 +60,17 @@ audit, where 18 synthetic patch tests passed. These environments are not identic
|
|
| 60 |
and the tests did not load the full 27B model. The upstream code
|
| 61 |
[MIT notice](compatibility/LICENSE-MLX-LM-MIT) is included separately from weight licenses.
|
| 62 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 63 |
## Text generation
|
| 64 |
|
| 65 |
```python
|
|
|
|
| 7 |
|
| 8 |
## Pin the snapshot and verify it
|
| 9 |
|
| 10 |
+
Use a new working directory. This repository is public; authentication is optional.
|
|
|
|
| 11 |
The command below resolves current main once to a full commit and then uses
|
| 12 |
only that pinned snapshot. For a repeat run, reuse the recorded commit.
|
| 13 |
Do not use an incomplete historical upload or proceed after verification failure.
|
|
|
|
| 20 |
printf 'Pinned model revision: %s\n' "$SWIFT_MLX_REVISION"
|
| 21 |
hf download ukisai/Swift-1.5-5bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-5bit-MLX
|
| 22 |
hf cache verify ukisai/Swift-1.5-5bit-MLX --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-5bit-MLX --fail-on-missing-files
|
| 23 |
+
python Swift-1.5-5bit-MLX/check_download.py Swift-1.5-5bit-MLX
|
| 24 |
python Swift-1.5-5bit-MLX/verify_release.py Swift-1.5-5bit-MLX
|
| 25 |
```
|
| 26 |
|
|
|
|
| 60 |
and the tests did not load the full 27B model. The upstream code
|
| 61 |
[MIT notice](compatibility/LICENSE-MLX-LM-MIT) is included separately from weight licenses.
|
| 62 |
|
| 63 |
+
Before generating, check the Python environment that will actually run the model:
|
| 64 |
+
|
| 65 |
+
```bash
|
| 66 |
+
python Swift-1.5-5bit-MLX/check_download.py Swift-1.5-5bit-MLX --runtime
|
| 67 |
+
```
|
| 68 |
+
|
| 69 |
+
This checks the complete download and selects the pinned patched architecture
|
| 70 |
+
without loading weights. It does not check a separate GUI app's environment,
|
| 71 |
+
available inference memory, or generated quality. See [TROUBLESHOOTING.md](TROUBLESHOOTING.md)
|
| 72 |
+
if it fails or a server reports `generation thread died`.
|
| 73 |
+
|
| 74 |
## Text generation
|
| 75 |
|
| 76 |
```python
|
check_download.py
ADDED
|
@@ -0,0 +1,188 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Check a Swift 1.5 4/5-bit download and optional pinned Python runtime.
|
| 3 |
+
|
| 4 |
+
Does not download files, load weights, generate text, or certify GUI compatibility.
|
| 5 |
+
"""
|
| 6 |
+
import argparse
|
| 7 |
+
import hashlib
|
| 8 |
+
import json
|
| 9 |
+
from pathlib import Path
|
| 10 |
+
import struct
|
| 11 |
+
import sys
|
| 12 |
+
|
| 13 |
+
RELEASES = {4: {
|
| 14 |
+
'model-00001-of-00003.safetensors': (5328325554, 'd91b70ca84dff315addeeb8a599dce2beccfec33c7483686ede2e8ff3cc459dc'),
|
| 15 |
+
'model-00002-of-00003.safetensors': (5354185158, 'eddcff1a6ef0971990f7cd01dba77ab9a8c20d59821bb4cd7e25234fe3809346'),
|
| 16 |
+
'model-00003-of-00003.safetensors': (5144253923, '644cc5322ff706c9608b42ea1ff83e0beb12b828d49b36c468a366f44593b8eb'),
|
| 17 |
+
}, 5: {
|
| 18 |
+
'model-00001-of-00004.safetensors': (5360417859, '9da8750595ec42d8c5d2ff0ae0b695e387cbcb3c0beac40f262e6d9a7196b59d'),
|
| 19 |
+
'model-00002-of-00004.safetensors': (5338181777, 'b592723bd4e9f01344335cd7d4b4f41a7c8cde98e2e074ad1d943cdb951a6835'),
|
| 20 |
+
'model-00003-of-00004.safetensors': (5355925620, '0027e7fe05310d5b2d960414221ed0872de9b511aa7e6b1c3a8f0e4647558412'),
|
| 21 |
+
'model-00004-of-00004.safetensors': (3227578201, 'afec577f16805b327471561d6841662eb4acf7d866ad5e761c5ef8cd3149d8f4'),
|
| 22 |
+
}}
|
| 23 |
+
EXPECTED_TENSORS = 2379
|
| 24 |
+
MODEL_HASHES = {
|
| 25 |
+
4: {'f559e4559946bfad4ce20b93906028449b76e007f42f36bb83d2c2c5f6cb95a9',
|
| 26 |
+
'ecda666d6d5f1059388a2ce4267d9ff8625398bb925d6c0840d0059ce17b676f'},
|
| 27 |
+
5: {'ecda666d6d5f1059388a2ce4267d9ff8625398bb925d6c0840d0059ce17b676f'},
|
| 28 |
+
}
|
| 29 |
+
REQUIRED = (
|
| 30 |
+
'config.json', 'model.safetensors.index.json', 'tokenizer.json',
|
| 31 |
+
'tokenizer_config.json', 'chat_template.jinja', 'generation_config.json',
|
| 32 |
+
'vocab.json', 'merges.txt', 'preprocessor_config.json', 'video_preprocessor_config.json',
|
| 33 |
+
)
|
| 34 |
+
|
| 35 |
+
def unique(pairs):
|
| 36 |
+
result = {}
|
| 37 |
+
for key, value in pairs:
|
| 38 |
+
if key in result:
|
| 39 |
+
raise ValueError('Duplicate JSON key')
|
| 40 |
+
result[key] = value
|
| 41 |
+
return result
|
| 42 |
+
|
| 43 |
+
def read_json(path):
|
| 44 |
+
return json.loads(path.read_text(encoding='utf-8'), object_pairs_hook=unique)
|
| 45 |
+
|
| 46 |
+
def runtime_check(config):
|
| 47 |
+
"""Check dispatch and the audited architecture source, without model creation."""
|
| 48 |
+
import inspect
|
| 49 |
+
from importlib.metadata import version
|
| 50 |
+
from mlx_lm.utils import _get_classes
|
| 51 |
+
cls, args = _get_classes(config)
|
| 52 |
+
args.from_dict(config)
|
| 53 |
+
row = {'mlx': version('mlx'), 'mlx_lm': version('mlx-lm'),
|
| 54 |
+
'selected_model_class': cls.__module__+'.'+cls.__name__}
|
| 55 |
+
if cls.__module__ != 'mlx_lm.models.qwen3_5_full':
|
| 56 |
+
raise ValueError('The required full-model MLX-LM patch is not active. Follow USAGE.md in this Python environment.')
|
| 57 |
+
row['model_source_sha256'] = hashlib.sha256(Path(inspect.getsourcefile(cls)).read_bytes()).hexdigest()
|
| 58 |
+
if row['model_source_sha256'] not in MODEL_HASHES[config['quantization']['bits']]:
|
| 59 |
+
raise ValueError('This is not the pinned patched model implementation. For 5-bit, both patches are required. Follow USAGE.md.')
|
| 60 |
+
return row
|
| 61 |
+
|
| 62 |
+
def check(folder, check_hashes=False, check_runtime=False):
|
| 63 |
+
folder = Path(folder)
|
| 64 |
+
errors = []
|
| 65 |
+
report = {'files':[], 'weights_loaded':False, 'generation_tested':False,
|
| 66 |
+
'scope':'File sizes, headers, index and requested checks only; no full-model or GUI test'}
|
| 67 |
+
try:
|
| 68 |
+
config = read_json(folder/'config.json')
|
| 69 |
+
bits = config['quantization']['bits']
|
| 70 |
+
if type(bits) is not int or bits not in RELEASES:
|
| 71 |
+
raise ValueError('Expected 4-bit or 5-bit Swift release')
|
| 72 |
+
expected = RELEASES[bits]
|
| 73 |
+
except (OSError,ValueError,KeyError,TypeError) as exc:
|
| 74 |
+
return dict(report,status='FAIL',errors=[f'Cannot identify the release: {exc}'])
|
| 75 |
+
report.update(bits=bits,expected_shards=len(expected),expected_weight_bytes=sum(v[0] for v in expected.values()))
|
| 76 |
+
for name in REQUIRED:
|
| 77 |
+
path = folder/name
|
| 78 |
+
if not path.is_file():
|
| 79 |
+
errors.append(f'Missing required asset: {name}')
|
| 80 |
+
elif path.stat().st_size == 0:
|
| 81 |
+
errors.append(f'Empty required asset: {name}')
|
| 82 |
+
elif name.endswith('.json'):
|
| 83 |
+
try:
|
| 84 |
+
read_json(path)
|
| 85 |
+
except (ValueError, OSError) as exc:
|
| 86 |
+
errors.append(f'Cannot parse {name}: {type(exc).__name__}')
|
| 87 |
+
seen = {}
|
| 88 |
+
sizes = {'U32':4, 'BF16':2, 'F16':2, 'F32':4}
|
| 89 |
+
for name,(size,sha256) in expected.items():
|
| 90 |
+
path = folder/name
|
| 91 |
+
actual = path.stat().st_size if path.is_file() else None
|
| 92 |
+
row = {'file':name,'expected_bytes':size,'actual_bytes':actual}
|
| 93 |
+
report['files'].append(row)
|
| 94 |
+
if actual != size:
|
| 95 |
+
errors.append(f'{name}: expected {size} bytes, found {actual if actual is not None else "missing"}')
|
| 96 |
+
continue
|
| 97 |
+
try:
|
| 98 |
+
with path.open('rb') as stream:
|
| 99 |
+
length = struct.unpack('<Q',stream.read(8))[0]
|
| 100 |
+
if not 0 < length <= 16*1024*1024:
|
| 101 |
+
raise ValueError('Invalid safetensors header size')
|
| 102 |
+
headers = json.loads(stream.read(length),object_pairs_hook=unique)
|
| 103 |
+
if not isinstance(headers,dict):
|
| 104 |
+
raise ValueError('Invalid safetensors header object')
|
| 105 |
+
headers.pop('__metadata__',None)
|
| 106 |
+
ranges = []
|
| 107 |
+
for tensor,entry in headers.items():
|
| 108 |
+
if tensor in seen:
|
| 109 |
+
raise ValueError('Duplicate tensor across shards')
|
| 110 |
+
start,end = entry['data_offsets']
|
| 111 |
+
shape = entry['shape']
|
| 112 |
+
if not isinstance(shape,list) or type(start) is not int or type(end) is not int:
|
| 113 |
+
raise ValueError('Invalid shape or offset types')
|
| 114 |
+
count = 1
|
| 115 |
+
for dim in shape:
|
| 116 |
+
if type(dim) is not int or dim < 0:
|
| 117 |
+
raise ValueError('Invalid tensor shape')
|
| 118 |
+
count *= dim
|
| 119 |
+
if entry['dtype'] not in sizes or end-start != count*sizes[entry['dtype']]:
|
| 120 |
+
raise ValueError('Invalid tensor size or dtype')
|
| 121 |
+
if not 0 <= start <= end <= actual-8-length:
|
| 122 |
+
raise ValueError('Invalid tensor data offsets')
|
| 123 |
+
seen[tensor] = name
|
| 124 |
+
ranges.append((start,end))
|
| 125 |
+
offset = 0
|
| 126 |
+
for start,end in sorted(ranges):
|
| 127 |
+
if start != offset:
|
| 128 |
+
raise ValueError('Overlapping tensor data or a payload gap')
|
| 129 |
+
offset = end
|
| 130 |
+
if offset != actual-8-length:
|
| 131 |
+
raise ValueError('Tensor payload does not fill shard')
|
| 132 |
+
row['header_check'] = 'PASS'
|
| 133 |
+
if check_hashes:
|
| 134 |
+
digest = hashlib.sha256()
|
| 135 |
+
with path.open('rb') as stream:
|
| 136 |
+
for block in iter(lambda:stream.read(8*1024*1024),b''):
|
| 137 |
+
digest.update(block)
|
| 138 |
+
row['sha256'] = digest.hexdigest()
|
| 139 |
+
row['sha256_check'] = 'PASS' if row['sha256'] == sha256 else 'FAIL'
|
| 140 |
+
if row['sha256'] != sha256:
|
| 141 |
+
errors.append(f'SHA256 mismatch: {name}')
|
| 142 |
+
except (ValueError,KeyError,TypeError,OSError,struct.error) as exc:
|
| 143 |
+
errors.append(f'{name}: {exc}')
|
| 144 |
+
extra = sorted(p.name for p in folder.glob('*.safetensors') if p.name not in expected)
|
| 145 |
+
if extra:
|
| 146 |
+
errors.append(f'Unexpected root weight files: {extra}')
|
| 147 |
+
try:
|
| 148 |
+
index = read_json(folder/'model.safetensors.index.json')
|
| 149 |
+
if not isinstance(index['weight_map'],dict):
|
| 150 |
+
raise ValueError('Invalid weight map')
|
| 151 |
+
if set(index['weight_map'].values()) != set(expected):
|
| 152 |
+
errors.append(f'Index must reference exactly the {len(expected)} expected shards')
|
| 153 |
+
if seen != index['weight_map']:
|
| 154 |
+
errors.append('Available tensor headers do not match the complete index')
|
| 155 |
+
if len(index['weight_map']) != EXPECTED_TENSORS:
|
| 156 |
+
errors.append(f'Expected {EXPECTED_TENSORS} saved tensor entries')
|
| 157 |
+
except (OSError,ValueError,KeyError,TypeError):
|
| 158 |
+
errors.append('Cannot validate the checkpoint index')
|
| 159 |
+
try:
|
| 160 |
+
if config.get('quantization') != {'bits':bits,'group_size':64,'mode':'affine'}:
|
| 161 |
+
errors.append(f'Expected affine / {bits}-bit / group size 64')
|
| 162 |
+
if config.get('language_model_only') is not False:
|
| 163 |
+
errors.append('Expected the complete Swift architecture configuration')
|
| 164 |
+
if check_runtime:
|
| 165 |
+
try:
|
| 166 |
+
report['runtime'] = runtime_check(config)
|
| 167 |
+
except Exception as exc:
|
| 168 |
+
errors.append(f'Runtime check failed: {type(exc).__name__}: {exc}')
|
| 169 |
+
except (OSError,ValueError,KeyError,TypeError):
|
| 170 |
+
errors.append('Cannot validate model configuration')
|
| 171 |
+
report['hash_check'] = ('PASS' if all(row.get('sha256_check')=='PASS' for row in report['files']) else 'FAIL_OR_INCOMPLETE') if check_hashes else 'not run; add --hash to read and SHA256-check all weight bytes'
|
| 172 |
+
report['runtime_check_requested'] = check_runtime
|
| 173 |
+
report['errors'] = errors
|
| 174 |
+
report['status'] = 'FAIL' if errors else 'PASS_DOWNLOAD_AND_REQUESTED_CHECKS_ONLY'
|
| 175 |
+
return report
|
| 176 |
+
|
| 177 |
+
def main():
|
| 178 |
+
parser = argparse.ArgumentParser(description=__doc__)
|
| 179 |
+
parser.add_argument('folder',type=Path)
|
| 180 |
+
parser.add_argument('--hash',action='store_true',help='Stream-read every shard and compare its original SHA256')
|
| 181 |
+
parser.add_argument('--runtime',action='store_true',help='Check the installed Python MLX-LM dispatcher; does not instantiate the model')
|
| 182 |
+
args = parser.parse_args()
|
| 183 |
+
result = check(args.folder,args.hash,args.runtime)
|
| 184 |
+
print(json.dumps(result,indent=2))
|
| 185 |
+
return 1 if result['errors'] else 0
|
| 186 |
+
|
| 187 |
+
if __name__ == '__main__':
|
| 188 |
+
sys.exit(main())
|