Instructions to use ukisai/Swift-1.5-4bit-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ukisai/Swift-1.5-4bit-MLX with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ukisai/Swift-1.5-4bit-MLX") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ukisai/Swift-1.5-4bit-MLX with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-1.5-4bit-MLX" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ukisai/Swift-1.5-4bit-MLX with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ukisai/Swift-1.5-4bit-MLX"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-1.5-4bit-MLX", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ukisai/Swift-1.5-4bit-MLX with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-1.5-4bit-MLX
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-1.5-4bit-MLX with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-4bit-MLX"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-1.5-4bit-MLX" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 2,465 Bytes
9fd3d5f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 | # Swift 1.5 MLX compatibility: validated complete checkpoint
Source: `ukisai/Swift-1.5-Qwen3.8-27b` at `00ccd14e006897d28cb0ed5bf26390e60d274251`.
All 18 BF16 shards and original runtime assets passed full SHA256 verification.
Source: 55563006776 weight bytes, 1199 BF16 tensors.
The official Qwen loader dropped 333 vision and 15 MTP tensors. The isolated patch
adds `mlx_lm/models/qwen3_5_full.py` with a real vision encoder and explicit MTP
module, plus strict dispatch/index checks and config/asset preservation in
`mlx_lm/utils.py`. `tests/test_qwen3_5_full.py` exercises mapping failures,
Transformers numerical agreement, cache behavior, native nonquantized conversion,
and exact-target quantized saving/reloading.
All 16 nonquantized architecture tests and the additional fixed affine test passed
on Linux CPU. The real source loaded strictly with 851 text, 333 vision and 15 MTP
tensors. There are zero ignored or unexplained tensors. The native quant has
2379 saved tensors because quantized weights have scales and biases.
All 609 BF16 remainder tensors equal the source values.
Text generation passed. Vision encoder execution and an MTP step using real text
hidden states passed. Image/video text integration and speculative generation are
not implemented. The original tokenizer, chat template, context, processor,
untied output head/shared MTP embeddings, norms and gating configuration are retained.
CPU runtime checks promote only in-memory floating values to FP32. The original
Linux BF16 QMM kernel produced 256 when summing 8192 exact ones; FP32 returned
8192. This reproducible backend issue and the runtime workaround are recorded in
cpu-quantized-matmul-diagnostic.json and ../USAGE.md. Stored weights were not changed.
Apple Silicon Metal checks passed for all 2379 native parameter headers and
real packed Q4 samples from text, vision and MTP. Both native BF16 Metal and
FP32 Metal executions matched their references. Only small actual samples were
evaluated on the 16 GiB Mac; no full 27B Mac generation is claimed.
The original source was read without modification. The source tensor layout
transposes are recorded individually in `quant-tensor-mapping-manifest.json`.
No custom quantization algorithm or alternate quantization configuration was used.
Patch SHA256: `f6f1d0bdafa45863bfbf93dac0398c481c993ea04fdf38b9bae98c643f89eaec`. Apply to official MLX-LM commit
`c69d1288440a0dc4e6401fc417098b07598dccd5`; see `../USAGE.md`.
|