--- license: other license_name: swift-open-license-1.0 license_link: https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/blob/main/LICENSE base_model: ukisai/Swift-1.5-Qwen3.8-27b base_model_relation: quantized library_name: mlx pipeline_tag: text-generation tags: - mlx - quantized - 4-bit - affine - qwen3_8 ---
UkisAI
Website  •  Learn more  •  BF16 model  •  GGUF  •  GSQ-RCO GGUF  •  Evaluation  •  Enterprise licensing
# Swift 1.5 Qwen3.8-27B — 4-bit MLX **Apple MLX 4-bit affine quantization.** This is Swift 1.5 in native MLX format, converted with the official Apple MLX-LM converter using 4 bits and group size 64. Swift 1.5 is UkisAI's reasoning-efficient Qwen3.8-27B derivative, focused on stronger long-horizon, agentic and coding performance while using fewer thinking tokens. **Starting from Homebrew or seeing `Received 501 parameters not in model`?** Use [QUICKSTART.md](QUICKSTART.md) to install the required runtime and create an explicit `serve` launcher. It reuses your existing model directory, including an HF cache snapshot. A bare `mlx_lm.server` command may select a separate Homebrew installation that lacks the Swift architecture and cache patches. **Runtime compatibility:** this complete checkpoint requires the supplied architecture patch and the pinned Python installation in [USAGE.md](USAGE.md). The tested unpatched MLX-LM 0.32.0 loader rejects 501 saved vision entries. Downloading the model does not install these patches into an app's inference engine. GUI compatibility remains unverified. Full-model follow-up validation is now recorded in [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md): both complete checkpoints passed an 86k-token synthetic text conversation and two cached follow-ups on a 48 GiB M4 Pro using the patched server. GUI integration and other memory/context sizes remain outside that test. The complete weights total **15.83 GB (14.74 GiB) across 3 required shards**, plus config, index and tokenizer files. A single 5–6 GB file is not the complete model. For download checks or `404 generation thread died`, use [TROUBLESHOOTING.md](TROUBLESHOOTING.md). Swift 1.5 uses **58.5% fewer thinking tokens** than base Qwen3.8-27B while scoring **0.35% higher**, for a **9.18× speed-up** on several tasks. ## Download the complete model This repository is public. Download the **whole repository from its root**: ```bash hf download ukisai/Swift-1.5-4bit-MLX --local-dir Swift-1.5-4bit-MLX ``` Install the CLI and patched runtime using the steps in [USAGE.md](USAGE.md). The full checkpoint is **15.83 GB (14.74 GiB)** and requires all three files: | Weight shard | Size | |---|---:| | [model-00001-of-00003.safetensors](https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/resolve/main/model-00001-of-00003.safetensors?download=true) | 5,328,325,554 bytes | | [model-00002-of-00003.safetensors](https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/resolve/main/model-00002-of-00003.safetensors?download=true) | 5,354,185,158 bytes | | [model-00003-of-00003.safetensors](https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/resolve/main/model-00003-of-00003.safetensors?download=true) | 5,144,253,923 bytes | You also need the root config, index and tokenizer assets; the command above downloads them together. The previously exposed **3.87 MB `real-checkpoint-samples.safetensors` was a diagnostic sample**, not a complete model. Those test files are now archived under `compatibility/mac-check/fixtures.zip` to keep them out of model-file discovery. Model weights are unchanged. **Runtime requirement:** use the included MLX-LM patch and [loading instructions](USAGE.md). The tested stock MLX-LM 0.32.0 loader cannot load this complete checkpoint; GUI integrations must supply a compatible loader and have not been validated. Full-model generation on a 24 GB Mac has not been validated; the weights require additional memory for the runtime, cache and macOS. ## Demo We gave base Qwen3.8-27B and Swift 1.5 27B the same prompt: > create a 3d little planet globe where I (player can walk around) and it has all these biomes to explore, the globe doesn't have to be too big, but still fun to go around. It's about a boy scout who is camping and goes around exploring. Try the game yourself here: [https://ukisai.com/swift-games/27b](https://ukisai.com/swift-games/27b) Base Qwen3.8-27B took 104.6 minutes to build its game. Swift 1.5 took 11.39 minutes. ## Source and quantization The conversion used the merged Swift 1.5 BF16 export associated with [`ukisai/Swift-1.5-Qwen3.8-27b` revision `00ccd14e006897d28cb0ed5bf26390e60d274251`](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/tree/00ccd14e006897d28cb0ed5bf26390e60d274251). The source repository was subsequently completed with all 18 BF16 shards and runtime assets at [revision `5ad04445d2686f525e9fbe5c077e6fa0c7df4200`](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/tree/5ad04445d2686f525e9fbe5c077e6fa0c7df4200). The later complete-revision link does not change the actual conversion provenance. All 1,199 source tensors are accounted for, including 333 vision and 15 MTP tensors. Eligible linear and embedding weights use 4-bit affine storage; 609 remaining tensors retain their original BF16 values after the documented layout mapping. All 18 source shards and runtime assets were verified by SHA-256. No base Qwen or alternate derived checkpoint weights were substituted. The saved checkpoint contains three weight shards. The original tokenizer, chat template, processor/config assets, license and notices are included. `QUANTIZATION_MANIFEST.json` records the fixed settings and validation results, while `UPLOAD_MANIFEST.json` records release-file checksums. ## Evaluation See the [Swift 1.5 source model card](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b#evaluation) for the source model's evaluations and methodology. Those results were not independently re-run on this MLX quantization. No broad accuracy or long-context benchmark was run for this release. ## Validation and use The validation results below are preserved historical build/component evidence, not a new full-model Apple run. The approximately 15.83 GB tensor payload requires additional runtime/cache and OS memory; do not force it onto a 16 GiB Mac or raise system limits. **Install the included MLX-LM architecture patch before loading this model.** [`USAGE.md`](USAGE.md) provides the pinned official revision, patch commands and a text generation example. The source configuration declares `Qwen3_5ForConditionalGeneration` / `qwen3_5`; the patch preserves that configuration and the inherited Swift text behavior. Validation passed on Linux CPU with MLX 0.32.2 and patched MLX-LM 0.32.0: complete source hashing, strict mapping and reload, finite floating tensors, exact BF16 remainder preservation, tokenizer/chat-template/processor loading, and a short text-generation smoke test that returned `Hello from Swift.`. CPU smoke tests use FP32 floating-point arithmetic with the original packed 4-bit tensors. This avoids a reproduced accumulation issue in MLX 0.32.2's Linux BF16 quantized-matmul path; checkpoint files and stored BF16 values are unchanged. Follow the Linux branch in `USAGE.md`. Apple Silicon checks passed for all 2,379 native tensor headers. Real packed text, vision and MTP weight samples passed native BF16 Metal execution and matched the FP32 reference within BF16 tolerance. These checks verify native Mac MLX compatibility, while full-model text generation is covered separately in [FULL_MAC_VALIDATION.md](FULL_MAC_VALIDATION.md). The vision encoder and an explicit MTP step passed real-weight component checks. Integrated image/video chat and speculative generation are not implemented in this patch. Component validation does not establish those end-to-end runtime features. The `compatibility/` directory retains the MLX-LM patch, reproducible instructions, source/tensor validation evidence, and the documented runtime limitations. Internal project names, machine paths and cloud-instance details have been redacted from the current published copies; older commits remain unchanged. ## License and access Swift 1.5 is a derivative of [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) (Copyright 2026 Alibaba Cloud, [Apache License 2.0](https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/blob/main/LICENSE-APACHE-2.0)). UkisAI's contribution, including the adapted weights, is licensed under the **[Swift Open License v1.0](https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/blob/main/LICENSE)**. See [NOTICE](https://huggingface.co/ukisai/Swift-1.5-4bit-MLX/blob/main/NOTICE) for the change notice and attribution details. Personal, research, educational, evaluation and commercial use are free for individuals and organizations with gross annual revenue, including affiliates, of up to US$1,000,000. Above that threshold, commercial use requires a separate Swift Enterprise License. Contact [UkisAI](https://ukisai.com/contact) for terms. Nothing in the Swift Open License limits rights in Qwen3.8-27B itself under Apache 2.0. The accompanying Apple MLX-LM code has a separate [MIT notice](compatibility/MLX-LM-LICENSE). ## Citation ```bibtex @misc{swift-1.5-qwen3.8-27b, title = {Swift 1.5 Qwen3.8-27B}, author = {UkisAI}, year = {2026}, url = {https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b} } ``` ## Acknowledgements We acknowledge the [NVIDIA Innovation Lab](https://www.nvidia.com/en-us/data-center/innovation-lab/), [Amazon Web Services](https://aws.amazon.com/), and [Google Cloud](https://cloud.google.com/) for providing compute credits and infrastructure support for Swift's development, training, and evaluation.