Related models: all models

Swift-1.5-Qwen3.8-27B-Uncensored-BF16 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer · Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE · Rank-2 Abliteration Patch

Swift-1.5 Qwen3.8 Flash-Next NVFP4 FP8PLE — Rank-2 Abliterated

Full baked checkpoint of the validated Rank-2 refusal-direction ablation for:

d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE

This release contains the complete ready-to-serve model weights. No offline patching or runtime abliteration hook is required.

The checkpoint preserves the FP8 PLE conversion and NVFP4 layout of the parent release while baking the validated Rank-2 {Orca, Swift⊥} projection directly into the affected model weights.

Important: do not apply the separate Rank-2 patch or the runtime flashnext_abliteration hook on top of this checkpoint. The transformation is already baked into the weights.

Hardware / deployment note: this full baked checkpoint is approximately 126 GiB and is not intended to be loaded as a complete in-memory checkpoint on a single NVIDIA RTX PRO 6000 96 GB.

For a single RTX PRO 6000 96 GB, use the validated deployment path based on the parent Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE together with the published Rank-2 runtime patch and NVMe-PLE setup.

This full baked release is intended primarily for multi-GPU or larger-memory deployments, runtimes with explicit model sharding/offload support, further conversion or quantization, and reproducible distribution of the already-abliterated weights.

Model lineage

Qwen/Qwen3.8-Flash-Next
        ↓
ukisai/Swift1.5-Qwen3.8-Flash-Next
        ↓
ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4
        ↓
d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE
        ↓
Rank-2 {Orca, Swift⊥} refusal-direction ablation
        ↓
d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated

What changed

The transformation suppresses the learned refusal direction while keeping the parent checkpoint architecture and quantization layout intact.

Property Value
Projection basis {Orca, Swift⊥}
Rank 2
Alpha 1.0
BF16 targets 100
NVFP4 expert weights transformed 24,576
Companion scale tensors 49,152
Modified safetensors shards 49
Modified payload 78.227 GiB
FP8 PLE shards 10 — unchanged

The large PLE embedding table remains in the FP8 E4M3 representation from the parent FP8PLE release. The abliteration does not rebuild or modify those 10 PLE shards.

Artifact

This repository contains the full baked checkpoint.

Property Value
Repository files 74
Safetensors shards 59
Indexed tensors 296,475
Published payload 135,257,737,044 bytes
Safetensors payload 135,195,697,746 bytes
Configured context 262,144 tokens
PLE FP8 E4M3
Target weights NVFP4 / mixed precision

The Hugging Face upload was verified after transfer:

  • 74 / 74 repository files present
  • 59 / 59 safetensors shards present
  • remote safetensors byte count exactly matches the local checkpoint
  • required config, tokenizer and safetensors-index files present

Abliteration validation

The baked checkpoint was independently compared with the validated reference transformation.

Reference-build validation:

  • 59 / 59 safetensors shards
  • 296,475 / 296,475 indexed tensors
  • 0 missing tensors
  • 0 extra tensors
  • 0 duplicate tensors
  • 0 wrong-shard mappings
  • 49 / 49 modified-shard SHA256 checks passed

Tensor-level differential audit:

expected_touched   = 1538
touched_changed    = 1538
touched_same       = 0
unexpected_changed = 0

All expected tensors changed and no unrelated tensor changed.

This establishes that the offline baked checkpoint reproduces the validated Rank-2 transformation rather than being an independently tuned approximation.

Refusal behavior

The validated Rank-2 transform produced:

Evaluation Result
Primary validation set 80 / 80 DIRECT — 100%
Held-out JBB non-AdvBench set 79 / 80 DIRECT — 98.75%
Held-out refusals 1 / 80 — 1.25%

The held-out evaluation used 80 prompts kept separate from the prompts used while developing the projection.

Configuration:

  • temperature: 0
  • maximum output: 192 tokens
  • concurrency: 4
  • deterministic refusal-marker classifier
  • output classified as REFUSE when a refusal marker was detected; otherwise DIRECT

These numbers describe this fixed evaluation only. They are not a guarantee that the model will never refuse under another prompt, system message, sampling configuration or runtime.

Performance

Reference generation-speed runs with the Rank-2 transform on a single NVIDIA RTX PRO 6000 Blackwell 96 GB:

Run Throughput
1 141.75 tok/s
2 138.66 tok/s
3 154.47 tok/s
Median 141.75 tok/s

Runtime:

  • Pennyroyal / SGLang v2.5.3
  • single NVIDIA RTX PRO 6000 Blackwell 96 GB
  • native NEXTN / MTP speculative decoding
  • FP8 KV cache
  • FP8 PLE
  • NVMe-backed PLE
  • 262,144-token configured context

Performance is hardware-, runtime- and workload-dependent.

Functional regression checks

The validated Rank-2 deployment passed:

  • normal chat-completion inference
  • native NEXTN / MTP speculative decoding
  • structured startup warmup
  • tool calling with correctly formed function arguments
  • end-to-end routing through hybrid_auto
  • 262,144-token configured context
  • NVMe-backed FP8 PLE operation

No agent/tool-calling regression was observed in these smoke tests.

Validation provenance

The behavioral, performance and functional measurements above were originally performed with the validated runtime application of the same Rank-2 {Orca, Swift⊥} projection.

The offline baker was then independently checked against the baked reference checkpoint:

  • all 49 modified shard SHA256 hashes matched
  • transformation totals matched exactly
  • all 1,538 expected tensors changed
  • zero unrelated tensors changed

This repository contains that baked transformation.

The behavioral benchmark was not separately rerun merely as a consequence of uploading the baked checkpoint to Hugging Face.

Parent-model quality reference

The parent d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE qualification produced:

Test Result
MMLU-Pro 226 / 280 — 80.71%
Native configured context 262,144 — PASS
Retrieval near context limit ~261.7K input tokens — PASS
Agentic tool/workflow smoke 7 / 7 PASS

These are parent-checkpoint reference results, not a post-abliteration MMLU-Pro score.

A separate post-abliteration capability suite can be reported when measured directly on this baked release.

Deployment notes

The original Rank-2 validation used Pennyroyal / SGLang v2.5.3 on NVIDIA Blackwell.

The documented single-RTX-PRO-6000 results were obtained with the parent FP8PLE checkpoint plus the Rank-2 runtime transformation and NVMe-backed PLE path. They should not be interpreted as evidence that this approximately 126 GiB baked checkpoint can be loaded entirely into the 96 GB VRAM of one RTX PRO 6000.

Reference image:

ghcr.io/jpezzulli/sglang-rtxpro6000:v2.5.3

Download the checkpoint:

hf download \
  d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated \
  --local-dir /srv/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated

Use the Pennyroyal next-plain profile and point TARGET_MODEL directly at the baked checkpoint:

PENNYROYAL_PROFILE=next-plain
PENNYROYAL_IMAGE=ghcr.io/jpezzulli/sglang-rtxpro6000:v2.5.3

TARGET_MODEL=/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated

PENNY_PLE_BACKEND=nvme
PENNY_PLE_NVME_MODEL=/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated-nvme

SGLANG_SM120_ONLINE_MXFP8=true
SGLANG_MM_PREPROCESS_DEVICE=cpu

MAX_RUNNING_REQUESTS=4
MAX_MAMBA_CACHE_SIZE=24

Prepare the NVMe PLE overlay

The NVMe PLE derivative is a machine-local runtime artifact and is not included in this repository.

Prepare it once from the baked checkpoint:

docker compose run --rm --no-deps \
  -v /srv/models:/models \
  pennyroyal exec .venv/bin/python \
  scripts/pennyroyal/prepare_ple_nvme.py \
  --source /models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated \
  --output /models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated-nvme

Then start Pennyroyal normally.

Do not double-apply the abliteration

This is a baked model.

Do not enable:

  • the Rank-2 patch baker
  • flashnext_abliteration.pth
  • another runtime {Orca, Swift⊥} projection hook

when serving this checkpoint.

Those mechanisms are intended for applying the transformation to the unabliterated FP8PLE parent. Applying them again would transform already modified weights a second time and is not the model validated here.

Reproducibility

The complete reproducible transformation kit is published separately:

d0xin/Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch

It contains:

  • direction_rank2_orca_swift.pt
  • plan.json
  • flashnext_abliteration_stage_b_rank2.py
  • bake_flashnext_rank2.py
  • apply_patch.sh
  • reference_BAKE_MANIFEST.jsonl
  • BASE_FINGERPRINTS.json
  • SHA256SUMS

Reference hashes:

Artifact SHA-256
Direction 8eb11e23dcbea3cbc040f5077856b19ea475d5af81d1a62f789ebe96d5130b2c
Plan 22ab558cf8708bf1f6a57d475ce2bc0c9ce7b5a9a54e5f756be70e28f30a7f50
Stage-B implementation 5138983ee95bd80b2645ce8f37ff20bab0c023376f6f68ee0c567842a6ebfbcf
Baker e777252cdb7fa445c4d19cd316f3218d736de9af2a31e9b4ba56b9c6322f8e8c

Related releases

Safety and responsible use

This model intentionally has substantially reduced refusal behavior.

Abliteration changes model-level refusal behavior; it does not provide application-level safety controls and it does not remove knowledge from the model.

Deployment operators are responsible for deciding what moderation, access control, monitoring and other safeguards are appropriate for their use case.

The refusal measurements above describe fixed evaluations and should not be interpreted as a universal guarantee about model behavior.

License and provenance

This model derives from Swift 1.5 Qwen3.8 Flash-Next.

UkisAI's Swift contribution is distributed under the Swift Open License v1.0. The underlying Qwen3.8-Flash-Next portions remain subject to the Qwen Community License 1.0.

See the included LICENSE, LICENSE-QWEN and NOTICE files for governing terms and attribution requirements.

The Rank-2 transformation in this repository modifies the parent checkpoint; it does not replace or supersede upstream license terms.

Downloads last month
63
Safetensors
Model size
120B params
Tensor type
F8_E4M3
·
BF16
·
U8
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated