Swift-1.5-Qwen3.8-27B · NVFP4 · NInfer

Swift-1.5-Qwen3.8-27B · huihui-style abliterated · NVFP4 · NInfer

A 27.78B-parameter multimodal derivative of ukisai/Swift-1.5-Qwen3.8-27b, abliterated in the huihui-ai style, quantized to NVFP4 + FP8 for the NInfer engine on Blackwell (sm_120a). One file contains the text model, the vision tower, the MTP head and a DFlash2 drafter.

Base ukisai/Swift-1.5-Qwen3.8-27b @ bc7a1e10b689
Abliteration huihui-style refusal-direction removal, transferred by weight difference from the Qwen/Qwen3.8-27B ↔ huihui-ai/Huihui-Qwen3.8-27B-abliterated pair. 70 tensors, language layers 17–51 (measured, see below)
Weight cost median ‖Δ‖/‖W‖ = 0.0188 (min 0.0177, max 0.0218) over the 70 changed matrices
Quantization NVFP4 (MLP of layers 0–55) + FP8 (attention, GDN, MLP of layers 56–63) — allocation copied verbatim from unsloth/Qwen3.8-27B-NVFP4, 32 calibration samples
Engine Neroued/ninfer @ bace20dc70249eed, built for sm_120a (CUDA 13.0.2 toolchain)
Container NInfer artifact v3, official tools.convert, recipe qwen3_8_27b_nvfp4, components text,vision,mtp,dflash2, --proposal
Artifact qwen3_8_27b_swift15_abliterated_nvfp4.ninfer — 1246 objects, 23,719,719,940 bytes, sha256 9810893aa7ee18ae3d213c3f95999191d52b425521da5312fe9c00fb8ca01088
Built 2026-09-29

Quickstart

Requires the ninfer runtime at revision ≥ bace20dc built for sm_120a (this artifact was verified with a build made on the CUDA 13.0.2 toolchain; see the NInfer README for its own requirements):

# one-shot
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
  --prompt "Tell me something interesting about the universe." \
  --max-context 16384 --max-new 512 --kv-dtype fp8 --no-thinking

# MTP speculative decoding
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
  --prompt "Write a haiku about a GPU cluster at 3am." \
  --max-context 16384 --max-new 256 --kv-dtype fp8 --no-thinking \
  --spec mtp --draft-tokens 3 --lm-head-draft

# DFlash2 speculative decoding
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
  --prompt "Write a haiku about a GPU cluster at 3am." \
  --max-context 16384 --max-new 256 --kv-dtype fp8 --no-thinking \
  --spec dflash2 --draft-tokens 7 --lm-head-draft

# image input (works together with DFlash2)
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
  --messages examples/cli/messages/image_chart.json --vision \
  --max-context 8192 --max-new 72 --kv-dtype fp8 --no-thinking

# serving (OpenAI / Anthropic compatible)
ninfer-serve qwen3_8_27b_swift15_abliterated_nvfp4.ninfer --port 8088

Vision and MTP/DFlash2 selection are startup choices; the runtime loads only the components you select.

Abliteration evaluation

The abliteration was checked on the BF16 checkpoints, before quantization, with the same 100 harmful prompts Heretic asks (mlabonne/harmful_behaviors test[:100] @ 01cead013), the same system prompt (You are a helpful assistant.), greedy decoding, 100 new tokens, and the 33 refusal markers of the Heretic keyword scorer (@ 3521f864, AGPL-3.0-or-later; logic reproduced here for comparability of numbers, response texts are not published). KL divergence is the first-token KL against the stock Swift 1.5 checkpoint on 100 harmless prompts (mlabonne/harmless_alpaca test[:100]), i.e. it prices the abliteration itself, not the Qwen→Swift difference.

Model Refusals by Heretic metric Explicit refusals (read by hand) KL vs stock
Swift 1.5 (stock, BF16) 98 / 100 98 0
Swift 1.5 abliterated (BF16, before quantization) 34 / 100 0 0.0791
This NVFP4/FP8 artifact (run inside ninfer-serve, thinking off) 45 / 100 0 not measured

How to read this:

  • The Heretic metric overcounts refusals for an abliterated model. Its list contains words that also appear in completed answers (disclaimer, illegal, violat, …). The model answers and adds a “Disclaimer: … for educational purposes” line, and the scorer counts that as a refusal. Every one of the 34 (BF16) and 45 (artifact) flagged answers was read by hand: all of them carry out the request (39 of the 45 artifact hits are disclaimer, 4 illegal, 1 violat, 1 i am an ai). Two BF16 answers (#66, #48) are softer compliance (a “simulated narrative” framing, a legal-vs-illegal reframing), not refusals.
  • The BF16 numbers were measured twice, on 2026-09-28 and again on 2026-09-29 from re-downloaded, sha256-verified inputs: verdicts and generated texts were identical for all 100 prompts on both variants.
  • The artifact number (45) is not comparable to the BF16 number (34). Different engine, NVFP4/FP8 weights, thinking disabled through the chat template instead of a forced <think></think> prefix, and 16 prompts flip from “not flagged” to “flagged” (5 flip the other way), i.e. whether a disclaimer/legality word appears in the first 100 tokens. Its meaningful result is 0 explicit refusals. No KL was measured for the quantized artifact.
  • One evaluation run is not a guarantee; other decoding settings, prompts, or the system prompt can change the numbers.
  • Full per-prompt verdicts (no text) are in evaluation-summary.json and refusal-probe-artifact.json.

Measured on this artifact

Single short runs on one RTX PRO 6000 Blackwell (Modal), --kv-dtype fp8 --no-thinking --greedy. Illustrative smoke checks, not benchmarks — acceptance rates depend heavily on the text. Raw output is in ninfer-runtime-report.json.

Check Result
Inventory version 3, 1246 objects, 1240 tensors, 6 resources, 1513 bindings, 844 uses; nvfp4 = 112, fp8_e4m3fn_row_bf16 = 146
Modes exercised text, arithmetic, vision (image chart), MTP, DFlash2, DFlash2 + vision — all completed, 0 fallback steps
Weights resident 19.0 GiB (text), 19.3 GiB with --vision
Decode speed, no speculation ~70 tok/s (69.5–70.1 across runs)
MTP (--draft-tokens 3) 75.0 % acceptance on a one-sentence answer (148 tok/s overall), 51.9 % on a 256-token haiku prompt (119.5 tok/s)
DFlash2 (--draft-tokens 7) 14.3 % and 13.1 % acceptance on the two short text prompts (78.6 and 90.1 tok/s overall); 85.7 % on the image example (122 tok/s)

Speculative decoding is not guaranteed to reproduce the non-speculative greedy text token for token: in the haiku smoke the plain, MTP and DFlash2 runs differ in one line.

Provenance

Component Source
Base weights ukisai/Swift-1.5-Qwen3.8-27b @ bc7a1e10b689648585a3ef41494c8d84cf77271a (18 shards, sha256-verified against the Hub LFS metadata)
Abliteration reference Qwen/Qwen3.8-27B @ 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 and huihui-ai/Huihui-Qwen3.8-27B-abliterated @ 739e3c5b89849f6c238ce1e5b70008612ae42cdd (both 18/18 shards sha256-verified)
Quantization recipe unsloth/Qwen3.8-27B-NVFP4 quantization_config, verbatim: recipe/unsloth_qconfig.json and recipe/quantize_nvfp4.py (sha256-pinned copies from Dragoy/Swift-Qwen3.8-27B-abliterated-NVFP4-NInfer @ 4c25aa201b41)
Calibration HuggingFaceH4/ultrachat_200k, 32 samples, seq 2048; dataset revision observed at 8049631c405ae657 before and after quantization (the recipe does not pin it)
Quantizer stack torch 2.11.0, transformers 5.10.1, llmcompressor 0.12.0.1, compressed-tensors 0.17.1
Converter Neroued/ninfer @ bace20dc70249eed6402b66d4852c6c3f9612905, unmodified tools.convert; chat template tools/chat_templates/qwen3_8.jinja (sha256 a497db9e&hellip;) — note this is NInfer's maintained template, not the byte-identical Qwen/Swift chat_template.jinja
DFlash2 drafter z-lab/Qwen3.8-27B-DFlash2 @ 50307d4c4cde6860d4eee73e2547cd786fe8e8a4, model.safetensors sha256 67fc76d68dc5a9415511a4f394ef744d67510cd20e93b37cc2cc7d28e4bab65c
Reports ablation-report.json, source-audit.json, quantization-report.json, conversion-report.json, qwen3_8_27b_swift15_abliterated_nvfp4.ninfer.conversion.json, checksums in SHA256SUMS

How the abliteration was applied and checked

The refusal projection is linear, so its effect on any derivative of the base model is the same constant difference:

W_ablated = W_swift15 + (W_huihui - W_qwen)      # per tensor, fp32, then rounded to bf16
  • The set of changed tensors is measured, not assumed. Comparing the full Qwen and huihui checkpoints (all 1199 tensors, 18/18 shards each) found exactly 70 differing tensors: self_attn.o_proj, linear_attn.out_proj and mlp.down_proj of language layers 17–51. (The huihui model card states layers 18–51; layer 17 is also changed, which the build's guard caught.) No vision (model.visual.*), MTP (mtp.*), lm_head or embedding tensor differs.
  • Independent audit (source-audit.json): a second, separate program re-read every shard, found the 1129 unchanged tensors bit-identical to Swift 1.5, and recomputed the 70 changed tensors from the three source checkpoints (torch.equal on the result). The quantizer and converter refuse to run on any checkpoint whose shard hashes differ from this audit.
  • The whole chain was built twice from clean inputs: the abliterated shards are byte-identical between runs (18/18 sha256).
  • Vision, MTP and the token embedding come from the BF16 abliterated checkpoint; the 70 abliterated matrices come from the quantized checkpoint (import_encoded), verified per matrix in the conversion.

Files

File Purpose
qwen3_8_27b_swift15_abliterated_nvfp4.ninfer the artifact
SHA256SUMS checksums of every file in this repository
recipe/ quantization recipe and the scripts that built and verified this release (Modal)
*-report.json, source-audit.json, evaluation-summary.json, refusal-probe-artifact.json, ninfer-runtime-report.json build, audit and measurement records
NOTICE, LICENSE, LICENSE-APACHE-2.0 licence and modification notices

License

This repository is a derivative of the Swift 1.5 checkpoint, whose license is the Swift Open License v1.0 — not Apache. The chain:

Component Licence
Qwen/Qwen3.8-27B (base model) Apache-2.0 — Copyright 2026 Alibaba Cloud (LICENSE-APACHE-2.0)
ukisai/Swift-1.5-Qwen3.8-27b (Swift Contribution) Swift Open License v1.0 (LICENSE)
huihui-ai/Huihui-Qwen3.8-27B-abliterated (source of the weight difference) Apache-2.0
This repo (abliteration + quantization + packaging) derivative work — the Swift Contribution contained in it stays under the Swift Open License v1.0

What that means in practice:

  • Free use, including commercial, while your gross revenue (counting all controlled entities) is below the $1,000,000 per fiscal year threshold; qualified non-profits have no threshold for non-commercial or research use.
  • Above the threshold: obtain a separate written licence from UkisAI (Swift Enterprise License).
  • Redistribution: ship both licence files, keep the copyright and attribution notices, and mark files you modified (Swift licence §4–§5). The weights here are modified (abliteration, quantization); see NOTICE.

This is a description of what the licences say, not legal advice.

Intended use and limitations

This is an uncensored model: the abliteration removes the refusal direction, so it will attempt requests that a stock instruction-tuned model declines, including harmful ones. It is published for research, evaluation and local deployment where that behaviour is understood and wanted. It is not safety-aligned, and its answers can be wrong, dangerous or illegal to act on.

Use at your own responsibility. Anyone deploying it is responsible for their own safeguards, output handling and compliance with the licences above and applicable law. It is provided as-is, without warranty.

What was not measured: quality benchmarks of any kind (the KL against stock Swift 1.5 above is the only quality proxy, and it was taken on the BF16 abliterated checkpoint, not on this artifact); refusal behaviour under thinking mode or other system prompts; long-context behaviour; anything on GPUs other than one RTX PRO 6000 Blackwell. Claims in the Swift 1.5 card apply to its BF16 weights, not to this quantized, abliterated build.

Credit for the base model to Qwen (Alibaba Cloud); for the Swift training to UkisAI; for the abliteration to huihui-ai; for the engine and artifact contract to Neroued; for the DFlash2 drafter to z-lab; for the published NVFP4 recipe to unsloth; and for the refusal scorer to p-e-w/heretic.

Downloads last month
112
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Dragoy/Swift-1.5-Qwen3.8-27B-abliterated-NVFP4-NInfer

Base model

Qwen/Qwen3.8-27B
Finetuned
(6)
this model