Instructions to use jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS # Run inference directly in the terminal: llama cli -hf jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS # Run inference directly in the terminal: llama cli -hf jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS # Run inference directly in the terminal: ./llama-cli -hf jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
Use Docker
docker model run hf.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
- LM Studio
- Jan
- vLLM
How to use jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
- Ollama
How to use jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF with Ollama:
ollama run hf.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
- Unsloth Desktop
- Pi
How to use jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF with Docker Model Runner:
docker model run hf.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
- Lemonade
How to use jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
Run and chat with the model
lemonade run user.ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF-IQ4_XS
List all available models
lemonade list
- Hermes Agent
How to use jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF:IQ4_XS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Swift-Qwen3.8-27B-Uncensored ATX-IQ4_XS-M (GGUF)
The ATX-IQ4_XS-M recipe applied to d0xin/Swift-Qwen3.8-27B-Uncensored-BF16. That model is an uncensored (abliterated) derivative of UkisAI's Swift-Qwen3.8-27B, made by rank-1 directional residual-stream ablation. The architecture, the 27.32B parameters and the 866 text and MTP tensors all match the base Qwen3.8-27B, so the recipe carries over unchanged. This build adds an importance matrix computed on this model.
The recipe was designed for a single RTX 3090 / 3090 Ti (24 GB). The goal is to fit a populated 200K-token context with an 8-bit key cache and the model's own MTP speculative head, and to decode faster per speculative round than the stock Q4 mixes on that card. The original build and its measurements are at jakeatx/Qwen3.8-27B-ATX-IQ4_XS-M-GGUF. The same recipe applied to Jackrong's Qwopus fine-tune is at jakeatx/Qwopus3.8-27B-Flash-ATX-IQ4_XS-M-GGUF.
This is an uncensored model. Its refusal behavior was deliberately removed upstream, so it may produce content that the original Swift model would decline to produce. Review its output before you rely on it, and make sure your use complies with the license and applicable law. The upstream card describes the ablation and its refusal and intelligence-preservation evaluations.
| file | size | notes |
|---|---|---|
ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M.gguf |
14.5 GiB (15,588,550,816 bytes) | 4.56 bits per weight, MTP layer included, text only (no vision tower) |
imatrix_swift_unc.gguf |
importance matrix computed from this model, see below | |
tensor_types_ATX-4-XS.txt |
the per-tensor type map used by llama-quantize (identical to the base ATX-IQ4_XS-M build) |
On the name. ATX-IQ4_XS-M means IQ4_XS as the base format for the bulk tensors, with upgrade pattern M. In llama.cpp's naming, the XS in IQ4_XS belongs to the format's name, not to a mix size: it is the 4.25 bits-per-weight super-block layout of the IQ4 codebook. On the K-quants, by contrast, the S/M/L suffix says how many tensors are lifted above the base format. This file lifts the same tensors Q4_K_M does, so it is an M-pattern mix on an IQ4_XS base.
Recipe
The Swift uncensored safetensors were converted to a BF16 GGUF with llama.cpp's converter, keeping the MTP layer. That file was then quantized with the per-tensor type map from the base ATX-IQ4_XS-M build:
| tensors | format | share of weight bytes |
|---|---|---|
| everything not listed below | IQ4_XS | ~75% |
| attn_output, ssm_out, ffn_down in the layers Unsloth's tier ladder upgrades first | Q5_0 | ~17% |
| attention K/V projections, output head | Q6_K | |
| token embedding (host side) | Q4_K | |
| the eight attention K/V tensors Q4_K_M keeps at Q8_0 (V in layers 11, 27, 31, 51, 55, 59, 63; K in 31) | Q8_0 | |
| MTP draft layer (blk.64) | Q5_0 | |
| GDN alpha / beta vectors (96 tiny) | Q8_0 |
llama-quantize --imatrix imatrix_swift_unc.gguf --tensor-type-file tensor_types_ATX-4-XS.txt \
--token-embedding-type q4_K Swift-Qwen3.8-27B-Uncensored-BF16.gguf ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M.gguf iq4_xs
After quantization, every tensor's name and type was checked against the base ATX-IQ4_XS-M file, and all of them were identical. The Qwopus build carries eleven extra early-layer Q5_0 tensors to match the protections in Jackrong's own Q4_K_S. No such reference mix exists for Swift, so this build uses the base map unchanged.
Importance matrix. This build does not reuse the base-model imatrix. It computes its own, from this uncensored Swift model at Q8_0 precision, over about 226K tokens of our calibration text:
- Bartowski's
calibration_datav3(mixed prose, code, multilingual text and chat). - 64K tokens of agentic and coding prompts, plus a 32K-token retrieval-analysis prompt, from the production prompt set the recipe was tuned on.
The calibration text is the same one used for the Qwopus build. Ablation and fine-tuning shift activation statistics, so measuring them on the model being quantized translates the recipe more faithfully.
Why this mix:
- On SM86, the fastest weight format per tensor at speculative verification widths 1-5 is IQ4_XS. The 2-3 bit codebook types (IQ3_S, IQ3_XXS, IQ2_S) are instruction-bound, so they run slower despite using fewer bytes.
- Q5_0 is about 16% cheaper than Q5_K.
- Q8_0 is the only format near the memory roof.
Extra bits go where Unsloth's tier ladder puts them: attention V/K, attention output, GDN output and FFN down.
sha256 of the GGUF: 51880ce0f15aebcad7abc8e4d6273b7bf270e15353a5ddcf63e8139c571b4425.
Run it
Runtime: llamAmpere, the Qwen3.8 / RTX 3090 fork with the SM86 kernel and memory work. It is the successor of https://github.com/JakeATX/llama-cpp-qwen-ampere. Build with -DGGML_CUDA=ON -DGGML_CUDA_FA=ON -DCMAKE_CUDA_ARCHITECTURES=86. The file also loads on stock TurboQuant+ and on mainline llama.cpp, with less context headroom; mainline has no turbo3 cache. The fastest llamAmpere v0.3 configuration:
GGML_Q8_TURBO3_MMA_FUSED=1 llama-server -m ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M.gguf \
-c 245760 -b 4096 -ub 1024 -t 8 -tb 8 -ngl 99 -fa on -ctk q8_0 -ctv turbo3 \
--parallel 1 --jinja --fit off \
--cache-prompt --cache-ram 8192 --ctx-checkpoints 24 --checkpoint-min-step 10240 \
--spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0 \
--spec-draft-type-k q8_0 --spec-draft-type-v q8_0 \
--spec-draft-vocab-map docs/mtp-vocab/atx_65536.txt
This is a single-user configuration that serves one request at a time. The base model's card explains the cache flags: they keep long conversations from re-prefilling, and they let you edit or regenerate a turn on this partly recurrent architecture. Lower -c if you need VRAM for something else. The upstream recommended sampling settings for thinking mode are temperature 1.0, top-p 0.95, top-k 20 and min-p 0.
This GGUF holds only the text model and its MTP head; the vision tower is not included. Speed and quality were not re-measured for this build. The recipe's measurements are on the base model's card, and the upstream card documents the Swift model's own evaluations.
LiveCodeBench v6 evaluation
An independent four-seed direct evaluation on a pinned 100-task LiveCodeBench v6 subset scored 357/400 = 89.25% pass@1 (seed scores: 90%, 87%, 90%, 90%). This is 1.05 percentage points below Qwen's published 90.3% BF16 reference figure.
The run used the exact GGUF on this card (SHA-256 51880ce0f15aebcad7abc8e4d6273b7bf270e15353a5ddcf63e8139c571b4425) with llamAmpere on one RTX 3090 Ti. Requests omitted max_tokens; generation was limited only by the server's approximately 150K context. Because this is a quantized uncensored Swift derivative tested with a custom runtime, a local pinned subset, four stochastic seeds, and an xhigh reasoning setting, the BF16 comparison is a reference rather than a controlled estimate of quantization loss.
Full sanitized per-sample metrics, protocol, token/timing data, hashes, difficulty breakdowns, and the nine formerly 32K-capped retry outcomes are published at jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-LiveCodeBench-v6.
License
These weights are distributed under the Swift Open License v1.0, inherited from the upstream model. Personal, research, educational, evaluation and commercial use is free for individuals and organizations with annual recurring revenue of up to US$1,000,000, including affiliates. Above that threshold, commercial use requires a separate Swift Enterprise License. Contact UkisAI for terms.
Credits
- UkisAI, for the original Swift model,
ukisai/Swift-Qwen3.8-27b. - d0xin, for the uncensored BF16 derivative,
d0xin/Swift-Qwen3.8-27B-Uncensored-BF16. - The Qwen team, for Qwen3.8.
- Unsloth, for the dynamic-quant tier ladder this recipe follows.
- Bartowski, for the calibration set.
- TheTom, for TurboQuant+.
- Downloads last month
- 2,271
4-bit
Model tree for jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF
Base model
Qwen/Qwen3.8-27BSpace using jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF 1
Collection including jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF
Evaluation results
- pass@1 (4-seed mean) on LiveCodeBench v6 (pinned 100-task subset, 4 seeds)self-reported89.250