Instructions to use ukisai/Swift-1.5-4bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ukisai/Swift-1.5-4bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ukisai/Swift-1.5-4bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ukisai/Swift-1.5-4bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-1.5-4bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ukisai/Swift-1.5-4bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ukisai/Swift-1.5-4bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-1.5-4bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ukisai/Swift-1.5-4bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-1.5-4bit-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-1.5-4bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-1.5-4bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download QUANTIZATION_MANIFEST.json from ukisai/Swift-1.5-4bit-MLX: direct link, hf CLI and curl.
- Browser
- Download file 8.23 kB
-
https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/resolve/deba7dee1987ecc2f4417c53fb49230af035e6d9/QUANTIZATION_MANIFEST.json
- Command line
-
hf download hf://ukisai/Swift-1.5-4bit-MLX@deba7dee1987ecc2f4417c53fb49230af035e6d9/QUANTIZATION_MANIFEST.json
-
curl -L -o QUANTIZATION_MANIFEST.json https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/resolve/deba7dee1987ecc2f4417c53fb49230af035e6d9/QUANTIZATION_MANIFEST.json
8.23 kB
| { | |
| "status": "VALIDATED", | |
| "created_at": "2026-09-21T18:16:33.223816+00:00", | |
| "source_repository": "ukisai/Swift-1.5-Qwen3.8-27b", | |
| "source_revision": "00ccd14e006897d28cb0ed5bf26390e60d274251", | |
| "source_manifest_sha256": "0a00065b88ab003281853a7fb9bd5ce0086bc3781b36136d8c39da19933923ae", | |
| "source_manifest_note": "SHA-256 of the original pre-publication-redaction export manifest.", | |
| "source_weight_shards": 18, | |
| "source_weight_bytes": 55563006776, | |
| "source_file_bytes": 55586102681, | |
| "source_recovery": "The pinned Hub revision was incomplete at conversion time. All 18 original shards and runtime assets were supplied by the verified project BF16 export and checked against the original export manifest; no weights were reconstructed or borrowed.", | |
| "quantization": { | |
| "mode": "affine", | |
| "bits": 4, | |
| "group_size": 64 | |
| }, | |
| "official_mlx_lm_repository": "https://github.com/ml-explore/mlx-lm", | |
| "official_mlx_lm_base_commit": "c69d1288440a0dc4e6401fc417098b07598dccd5", | |
| "architecture_patch_sha256": "f6f1d0bdafa45863bfbf93dac0398c481c993ea04fdf38b9bae98c643f89eaec", | |
| "packages": { | |
| "mlx": "0.32.2", | |
| "mlx-cpu": "0.32.2", | |
| "mlx-lm": "0.32.0", | |
| "transformers": "5.14.1", | |
| "huggingface_hub": "1.31.0", | |
| "torch": "2.11.0+cpu", | |
| "torchvision": "0.26.0+cpu", | |
| "safetensors": "0.8.0" | |
| }, | |
| "python": "3.12.3", | |
| "platform": "Linux x86_64", | |
| "validation_runtime": { | |
| "device": "CPU" | |
| }, | |
| "conversion_command": [ | |
| "mlx_lm.convert", | |
| "--hf-path", | |
| "<SOURCE_MODEL_DIR>", | |
| "--mlx-path", | |
| "<OUTPUT_MODEL_DIR>", | |
| "--quantize", | |
| "--q-mode", | |
| "affine", | |
| "--q-bits", | |
| "4", | |
| "--q-group-size", | |
| "64" | |
| ], | |
| "conversion_elapsed_seconds": 120.018799242, | |
| "output_weight_shards": 3, | |
| "output_weight_bytes": 15826764635, | |
| "source_tensors": 1199, | |
| "mapped_source_tensors": 1199, | |
| "saved_tensors": 2379, | |
| "categories": { | |
| "text": 851, | |
| "MTP": 15, | |
| "vision": 333 | |
| }, | |
| "quantized_weight_tensors": 590, | |
| "original_bf16_tensors": 609, | |
| "ignored_tensors": 0, | |
| "unexplained_tensors": 0, | |
| "validation": { | |
| "status": "PASS", | |
| "recorded_at": "2026-09-21T18:14:21.887850+00:00", | |
| "source_repo": "ukisai/Swift-1.5-Qwen3.8-27b", | |
| "source_revision": "00ccd14e006897d28cb0ed5bf26390e60d274251", | |
| "source_shards": 18, | |
| "source_shard_bytes": 55563006776, | |
| "source_tensors": 1199, | |
| "mapped_source_tensors": 1199, | |
| "saved_tensors": 2379, | |
| "categories": { | |
| "text": 851, | |
| "MTP": 15, | |
| "vision": 333 | |
| }, | |
| "ignored_tensors": 0, | |
| "unexplained_tensors": 0, | |
| "exact_unquantized_tensors": 609, | |
| "quantization": { | |
| "group_size": 64, | |
| "bits": 4, | |
| "mode": "affine" | |
| }, | |
| "all_floating_tensors_finite": true, | |
| "tokenizer": "Qwen2Tokenizer", | |
| "processor": "Qwen3VLProcessor", | |
| "assets_sha256": { | |
| "generation_config.json": "e70c136c1b78ddc1fb0905bac8e733a4dc448d4f852a5dd75143fffc70be550e", | |
| "preprocessor_config.json": "27225450ac9c6529872ee1924fcb0962ff5634834f817040f444118116f4e516", | |
| "video_preprocessor_config.json": "7768af27c1fafa9cc9011c1dc20067e03f8915e03b63504550e11d5066986d13", | |
| "tokenizer.json": "0997f410c57a1f4e53b09e4be8f4a172d90edd9564368fb0847030937229b9f3", | |
| "tokenizer_config.json": "b11349aafa7cdc6a320767cf7ceb29ed82f7eda5d65e8e0819e76f0ce947bf27", | |
| "vocab.json": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003", | |
| "merges.txt": "a9d356d7bdf1ef4949e3e748e95b8e10ad9d4e2e838eddc38a0a7b6b94d1db8d", | |
| "chat_template.jinja": "c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041" | |
| }, | |
| "chat_templates": [ | |
| { | |
| "options": { | |
| "enable_thinking": false | |
| }, | |
| "rendered": "<|im_start|>user\nSay hello.<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n" | |
| }, | |
| { | |
| "options": { | |
| "reasoning_effort": "low" | |
| }, | |
| "rendered": "<|im_start|>system\nReasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.<|im_end|>\n<|im_start|>user\nSay hello.<|im_end|>\n<|im_start|>assistant\n<think>\n" | |
| }, | |
| { | |
| "options": { | |
| "reasoning_effort": "xhigh" | |
| }, | |
| "rendered": "<|im_start|>system\nReasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.<|im_end|>\n<|im_start|>user\nSay hello.<|im_end|>\n<|im_start|>assistant\n<think>\n" | |
| } | |
| ], | |
| "load_seconds": 3.351265648000208, | |
| "load_memory_bytes": 15826466152, | |
| "process_peak_rss_bytes": 21142360064, | |
| "inference_floating_dtype": "float32, CPU runtime only; stored floating tensors remain BF16", | |
| "native_bf16_cpu_inference": "Aborted after reproducing incorrect accumulation in the official Linux BF16 quantized matmul. See cpu-quantized-matmul-diagnostic.json.", | |
| "generation": { | |
| "prompt": "<|im_start|>user\nReply with exactly: Hello from Swift.<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n", | |
| "text": "Hello from Swift.", | |
| "token_ids": [ | |
| 9419, | |
| 494, | |
| 22929, | |
| 13, | |
| 248046 | |
| ], | |
| "tokens": 5, | |
| "tokens_per_second": 0.0797672292904007, | |
| "prompt_tokens_per_second": 0.06443886297275572, | |
| "elapsed_seconds": 374.00736365800003, | |
| "finish_reason": "stop" | |
| }, | |
| "mtp": { | |
| "status": "PASS", | |
| "shape": [ | |
| 1, | |
| 1, | |
| 248320 | |
| ], | |
| "path": "Explicit MTP step with real text hidden states and shared LM head; speculative generation is not integrated" | |
| }, | |
| "vision": { | |
| "status": "PASS", | |
| "shape": [ | |
| 64, | |
| 5120 | |
| ], | |
| "grid": [ | |
| [ | |
| 1, | |
| 16, | |
| 16 | |
| ] | |
| ], | |
| "path": "Vision encoder only; image/video insertion and multimodal text generation are not implemented" | |
| }, | |
| "total_validation_seconds": 467.15810263900016 | |
| }, | |
| "private_repository": "ukisai/Swift-1.5-4bit-MLX", | |
| "apple_silicon_verification": { | |
| "status": "PASS_MAC_NATIVE_MLX_FORMAT_AND_REAL_METAL_SAMPLES", | |
| "platform": "macOS-26.6-arm64-arm-64bit", | |
| "machine": "arm64", | |
| "mlx": "0.32.2", | |
| "mlx_lm": "0.32.0", | |
| "device": "Device(gpu, 0)", | |
| "quantization": { | |
| "group_size": 64, | |
| "bits": 4, | |
| "mode": "affine" | |
| }, | |
| "all_saved_tensor_headers_validated": 2379, | |
| "all_source_parameters_accounted_for": 1199, | |
| "strict_complete_parameter_tree": "PASS using unevaluated header fixtures; no fabricated weights saved", | |
| "actual_checkpoint_samples": [ | |
| { | |
| "sample": 0, | |
| "category": "text", | |
| "checkpoint_weight": "language_model.model.layers.0.linear_attn.in_proj_a.weight", | |
| "original_shape": [ | |
| 48, | |
| 5120 | |
| ], | |
| "packed_shape": [ | |
| 48, | |
| 640 | |
| ], | |
| "native_metal_bf16": "PASS", | |
| "metal_fp32": "PASS", | |
| "bf16_max_absolute_error": 0.0004401206970214844, | |
| "fp32_max_absolute_error": 0.0 | |
| }, | |
| { | |
| "sample": 1, | |
| "category": "vision", | |
| "checkpoint_weight": "visual.blocks.0.attn.proj.weight", | |
| "original_shape": [ | |
| 1152, | |
| 1152 | |
| ], | |
| "packed_shape": [ | |
| 1152, | |
| 144 | |
| ], | |
| "native_metal_bf16": "PASS", | |
| "metal_fp32": "PASS", | |
| "bf16_max_absolute_error": 0.0002315044403076172, | |
| "fp32_max_absolute_error": 0.0 | |
| }, | |
| { | |
| "sample": 2, | |
| "category": "MTP", | |
| "checkpoint_weight": "mtp.layers.0.self_attn.k_proj.weight", | |
| "original_shape": [ | |
| 1024, | |
| 5120 | |
| ], | |
| "packed_shape": [ | |
| 1024, | |
| 640 | |
| ], | |
| "native_metal_bf16": "PASS", | |
| "metal_fp32": "PASS", | |
| "bf16_max_absolute_error": 0.0009589195251464844, | |
| "fp32_max_absolute_error": 0.0 | |
| } | |
| ], | |
| "full_model_mac_generation": "NOT_RUN: targeted native Metal format and real-weight component checks only", | |
| "peak_mlx_bytes": 4597960, | |
| "peak_process_rss_bytes": 360693760 | |
| } | |
| } | |