Instructions to use d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated") model = AutoModelForMultimodalLM.from_pretrained("d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated
- SGLang
How to use d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated with Docker Model Runner:
docker model run hf.co/d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated
- Swift-1.5 Qwen3.8 Flash-Next NVFP4 FP8PLE — Rank-2 Abliterated
Related models: all models
Swift-1.5-Qwen3.8-27B-Uncensored-BF16 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8 · Swift-1.5-Qwen3.8-27B-Uncensored-FP8-NInfer · Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE · Rank-2 Abliteration Patch
Swift-1.5 Qwen3.8 Flash-Next NVFP4 FP8PLE — Rank-2 Abliterated
Full baked checkpoint of the validated Rank-2 refusal-direction ablation for:
d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE
This release contains the complete ready-to-serve model weights. No offline patching or runtime abliteration hook is required.
The checkpoint preserves the FP8 PLE conversion and NVFP4 layout of the parent release while baking the validated Rank-2 {Orca, Swift⊥} projection directly into the affected model weights.
Important: do not apply the separate Rank-2 patch or the runtime
flashnext_abliterationhook on top of this checkpoint. The transformation is already baked into the weights.
Hardware / deployment note: this full baked checkpoint is approximately 126 GiB and is not intended to be loaded as a complete in-memory checkpoint on a single NVIDIA RTX PRO 6000 96 GB.
For a single RTX PRO 6000 96 GB, use the validated deployment path based on the parent
Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLEtogether with the published Rank-2 runtime patch and NVMe-PLE setup.This full baked release is intended primarily for multi-GPU or larger-memory deployments, runtimes with explicit model sharding/offload support, further conversion or quantization, and reproducible distribution of the already-abliterated weights.
Model lineage
Qwen/Qwen3.8-Flash-Next
↓
ukisai/Swift1.5-Qwen3.8-Flash-Next
↓
ukisai/Swift-1.5-Qwen3.8-Flash-Next-NVFP4
↓
d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE
↓
Rank-2 {Orca, Swift⊥} refusal-direction ablation
↓
d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated
What changed
The transformation suppresses the learned refusal direction while keeping the parent checkpoint architecture and quantization layout intact.
| Property | Value |
|---|---|
| Projection basis | {Orca, Swift⊥} |
| Rank | 2 |
| Alpha | 1.0 |
| BF16 targets | 100 |
| NVFP4 expert weights transformed | 24,576 |
| Companion scale tensors | 49,152 |
| Modified safetensors shards | 49 |
| Modified payload | 78.227 GiB |
| FP8 PLE shards | 10 — unchanged |
The large PLE embedding table remains in the FP8 E4M3 representation from the parent FP8PLE release. The abliteration does not rebuild or modify those 10 PLE shards.
Artifact
This repository contains the full baked checkpoint.
| Property | Value |
|---|---|
| Repository files | 74 |
| Safetensors shards | 59 |
| Indexed tensors | 296,475 |
| Published payload | 135,257,737,044 bytes |
| Safetensors payload | 135,195,697,746 bytes |
| Configured context | 262,144 tokens |
| PLE | FP8 E4M3 |
| Target weights | NVFP4 / mixed precision |
The Hugging Face upload was verified after transfer:
- 74 / 74 repository files present
- 59 / 59 safetensors shards present
- remote safetensors byte count exactly matches the local checkpoint
- required config, tokenizer and safetensors-index files present
Abliteration validation
The baked checkpoint was independently compared with the validated reference transformation.
Reference-build validation:
- 59 / 59 safetensors shards
- 296,475 / 296,475 indexed tensors
- 0 missing tensors
- 0 extra tensors
- 0 duplicate tensors
- 0 wrong-shard mappings
- 49 / 49 modified-shard SHA256 checks passed
Tensor-level differential audit:
expected_touched = 1538
touched_changed = 1538
touched_same = 0
unexpected_changed = 0
All expected tensors changed and no unrelated tensor changed.
This establishes that the offline baked checkpoint reproduces the validated Rank-2 transformation rather than being an independently tuned approximation.
Refusal behavior
The validated Rank-2 transform produced:
| Evaluation | Result |
|---|---|
| Primary validation set | 80 / 80 DIRECT — 100% |
| Held-out JBB non-AdvBench set | 79 / 80 DIRECT — 98.75% |
| Held-out refusals | 1 / 80 — 1.25% |
The held-out evaluation used 80 prompts kept separate from the prompts used while developing the projection.
Configuration:
- temperature:
0 - maximum output:
192tokens - concurrency:
4 - deterministic refusal-marker classifier
- output classified as
REFUSEwhen a refusal marker was detected; otherwiseDIRECT
These numbers describe this fixed evaluation only. They are not a guarantee that the model will never refuse under another prompt, system message, sampling configuration or runtime.
Performance
Reference generation-speed runs with the Rank-2 transform on a single NVIDIA RTX PRO 6000 Blackwell 96 GB:
| Run | Throughput |
|---|---|
| 1 | 141.75 tok/s |
| 2 | 138.66 tok/s |
| 3 | 154.47 tok/s |
| Median | 141.75 tok/s |
Runtime:
- Pennyroyal / SGLang v2.5.3
- single NVIDIA RTX PRO 6000 Blackwell 96 GB
- native NEXTN / MTP speculative decoding
- FP8 KV cache
- FP8 PLE
- NVMe-backed PLE
- 262,144-token configured context
Performance is hardware-, runtime- and workload-dependent.
Functional regression checks
The validated Rank-2 deployment passed:
- normal chat-completion inference
- native NEXTN / MTP speculative decoding
- structured startup warmup
- tool calling with correctly formed function arguments
- end-to-end routing through
hybrid_auto - 262,144-token configured context
- NVMe-backed FP8 PLE operation
No agent/tool-calling regression was observed in these smoke tests.
Validation provenance
The behavioral, performance and functional measurements above were originally performed with the validated runtime application of the same Rank-2 {Orca, Swift⊥} projection.
The offline baker was then independently checked against the baked reference checkpoint:
- all 49 modified shard SHA256 hashes matched
- transformation totals matched exactly
- all 1,538 expected tensors changed
- zero unrelated tensors changed
This repository contains that baked transformation.
The behavioral benchmark was not separately rerun merely as a consequence of uploading the baked checkpoint to Hugging Face.
Parent-model quality reference
The parent d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE qualification produced:
| Test | Result |
|---|---|
| MMLU-Pro | 226 / 280 — 80.71% |
| Native configured context | 262,144 — PASS |
| Retrieval near context limit | ~261.7K input tokens — PASS |
| Agentic tool/workflow smoke | 7 / 7 PASS |
These are parent-checkpoint reference results, not a post-abliteration MMLU-Pro score.
A separate post-abliteration capability suite can be reported when measured directly on this baked release.
Deployment notes
The original Rank-2 validation used Pennyroyal / SGLang v2.5.3 on NVIDIA Blackwell.
The documented single-RTX-PRO-6000 results were obtained with the parent FP8PLE checkpoint plus the Rank-2 runtime transformation and NVMe-backed PLE path. They should not be interpreted as evidence that this approximately 126 GiB baked checkpoint can be loaded entirely into the 96 GB VRAM of one RTX PRO 6000.
Reference image:
ghcr.io/jpezzulli/sglang-rtxpro6000:v2.5.3
Download the checkpoint:
hf download \
d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated \
--local-dir /srv/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated
Use the Pennyroyal next-plain profile and point TARGET_MODEL directly at the baked checkpoint:
PENNYROYAL_PROFILE=next-plain
PENNYROYAL_IMAGE=ghcr.io/jpezzulli/sglang-rtxpro6000:v2.5.3
TARGET_MODEL=/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated
PENNY_PLE_BACKEND=nvme
PENNY_PLE_NVME_MODEL=/models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated-nvme
SGLANG_SM120_ONLINE_MXFP8=true
SGLANG_MM_PREPROCESS_DEVICE=cpu
MAX_RUNNING_REQUESTS=4
MAX_MAMBA_CACHE_SIZE=24
Prepare the NVMe PLE overlay
The NVMe PLE derivative is a machine-local runtime artifact and is not included in this repository.
Prepare it once from the baked checkpoint:
docker compose run --rm --no-deps \
-v /srv/models:/models \
pennyroyal exec .venv/bin/python \
scripts/pennyroyal/prepare_ple_nvme.py \
--source /models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated \
--output /models/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated-nvme
Then start Pennyroyal normally.
Do not double-apply the abliteration
This is a baked model.
Do not enable:
- the Rank-2 patch baker
flashnext_abliteration.pth- another runtime
{Orca, Swift⊥}projection hook
when serving this checkpoint.
Those mechanisms are intended for applying the transformation to the unabliterated FP8PLE parent. Applying them again would transform already modified weights a second time and is not the model validated here.
Reproducibility
The complete reproducible transformation kit is published separately:
d0xin/Swift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch
It contains:
direction_rank2_orca_swift.ptplan.jsonflashnext_abliteration_stage_b_rank2.pybake_flashnext_rank2.pyapply_patch.shreference_BAKE_MANIFEST.jsonlBASE_FINGERPRINTS.jsonSHA256SUMS
Reference hashes:
| Artifact | SHA-256 |
|---|---|
| Direction | 8eb11e23dcbea3cbc040f5077856b19ea475d5af81d1a62f789ebe96d5130b2c |
| Plan | 22ab558cf8708bf1f6a57d475ce2bc0c9ce7b5a9a54e5f756be70e28f30a7f50 |
| Stage-B implementation | 5138983ee95bd80b2645ce8f37ff20bab0c023376f6f68ee0c567842a6ebfbcf |
| Baker | e777252cdb7fa445c4d19cd316f3218d736de9af2a31e9b4ba56b9c6322f8e8c |
Related releases
Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE— unabliterated FP8-PLE parentSwift-1.5-Qwen3.8-Flash-Next-Rank2-Abliteration-Patch— reproducible Rank-2 patch kitSwift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated— this full baked checkpoint
Safety and responsible use
This model intentionally has substantially reduced refusal behavior.
Abliteration changes model-level refusal behavior; it does not provide application-level safety controls and it does not remove knowledge from the model.
Deployment operators are responsible for deciding what moderation, access control, monitoring and other safeguards are appropriate for their use case.
The refusal measurements above describe fixed evaluations and should not be interpreted as a universal guarantee about model behavior.
License and provenance
This model derives from Swift 1.5 Qwen3.8 Flash-Next.
UkisAI's Swift contribution is distributed under the Swift Open License v1.0. The underlying Qwen3.8-Flash-Next portions remain subject to the Qwen Community License 1.0.
See the included LICENSE, LICENSE-QWEN and NOTICE files for governing terms and attribution requirements.
The Rank-2 transformation in this repository modifies the parent checkpoint; it does not replace or supersede upstream license terms.
- Downloads last month
- 63
Model tree for d0xin/Swift-1.5-Qwen3.8-Flash-Next-NVFP4-FP8PLE-Rank2-Abliterated
Base model
Qwen/Qwen3.8-Flash-Next