File size: 7,682 Bytes
f5b159e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 | ---
license: other
license_name: qwen-research
license_link: LICENSE
base_model: abenzerps/Qwen-Image-2.1-Uncensored-GGUF
base_model_relation: quantized
pipeline_tag: text-to-image
tags:
- mxfp4
- amd
- rocm
- image-editing
- rgba
---
# Qwen_Image-2.1-Uncensored-MXFP4-Paiton
**Built with Qwen.** This is EliovpAI's balanced MXFP4 conversion of the BF16
checkpoint published by [abenzerps](https://huggingface.co/abenzerps/Qwen-Image-2.1-Uncensored-GGUF),
derived from [Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1).
Fine-tune credit belongs to abenzerps; Paiton supplies the conversion and runtime
integration. This was converted from BF16, not from a lower-bit GGUF quant.
The author describes a merged LoRA in
[the source discussion](https://huggingface.co/abenzerps/Qwen-Image-2.1-Uncensored-GGUF/discussions/20).
The training recipe is not published. “Uncensored” retains the author's model
label; our ordinary-image correctness screen does not verify every claimed
fine-tuning behavior or guarantee unrestricted outputs.
## Weights and runtimes
Tensor payload: **8.674 GiB**. Storage uses E2M1 MXFP4 with one E8M0 scale per 32
columns. The balanced recipe packs 224 transformer and 372 text/vision targets.
Transformer input/output/conditioning, normalizations, biases and the VAE retain
BF16 inference precision. Activations remain BF16 in this package.
The supplied loader is required; an unmodified
`DiffusionPipeline.from_pretrained` cannot load this packed format.
| Runtime | Behavior |
| --- | --- |
| This HF package, compatible `gfx1201` GPU | Automatically uses the included, SHA-256-verified native HIP weight decoder |
| This HF package, other supported ROCm GPU | Uses framework weight reconstruction; tested on MI355X |
| Paiton plugin/container | Uses the existing optimized RDNA4 pipeline, with checkpoint selection at launch |
The included decoder is the existing released binary, unchanged. It has no Torch
or Triton dependency; the surrounding image pipeline uses the pinned framework.
The additional attention, normalization and GEMM optimizations are supplied by
the plugin/container. The compiler and generated implementation source are not
included. The portable decoder does not provide those serving optimizations.
## One container, either model
The plugin's v1.0.3 launcher selects the model before loading it:
```bash
./models/Qwen-Image-2.1/serve-docker.sh --model uncensored
./models/Qwen-Image-2.1/serve-docker.sh --model original
```
Run one worker at a time. `original` remains the default. Both selections use the
same container and native libraries, with separate pinned HF revisions and file
hashes. The uncensored selection defaults to the `exact` BF16 arithmetic profile;
the original retains its existing `schedule-int8` default. Low-precision profiles
remain explicit options and have not been graded on this fine-tune.
Without a checkout:
```bash
docker run --rm --device /dev/kfd --device /dev/dri --ipc=host -p 127.0.0.1:8191:8191 -v paiton-qwen-image21-cache:/cache ghcr.io/eliovp/paiton-vllm-plugin:qwen-image21-mxfp4-rdna4-v1.0.3 --model uncensored
```
The API identifies this checkpoint as `paiton-image-2.1-uncensored`. A generation
request may omit `model` to use the loaded checkpoint. Switching checkpoints
requires restarting the worker; two models are not resident simultaneously.
## Direct package use
Download this repository with the Hugging Face CLI. Use Python 3.12 and a ROCm
PyTorch stack compatible with the GPU. Install the pinned image dependencies
from `requirements.txt`; `requirements-mi355.lock.txt` records the tested MI355
environment. Native decoder use requires a runtime able to load the included
ROCm 10 artifact; an unavailable native library falls back to framework decoding
in `auto` mode and records the reason.
```bash
hf download EliovpAI/Qwen_Image-2.1-Uncensored-MXFP4-Paiton --local-dir model
cd model
python -m pip install -r requirements.txt
./launch.sh --model-dir . generate --prompt 'A green ceramic teapot on a wooden table' --size 2048 --seed 44 --output ../outputs/teapot.png
```
Use `--decoder framework` or `--decoder native` before `generate` to choose
explicitly. `--mode rgba` requests transparency. Editing uses `--mode edit`, one
`--image` and `--size 1024`. Supported generation sizes are 1024 and 2048 square,
batch one, 40 steps and guidance 1.0. No CPU model offload or VAE tiling is used.
Whole-pipeline graph capture is unsupported.
## Correctness and measured limits
Conversion and correctness were tested on **MI355X (`gfx950`)**, Torch
`2.10.0+rocm7.1`, HIP `7.1.25424`, Quark `0.12.post1+rocm71.torch2.10`, and
Diffusers commit `7263f3317f6b392d62f41e9d75ed9d7e21fc5a5c`.
- All 297 BF16 source tensors were extracted losslessly. All 224 converted
transformer layers matched Quark through an independent MXFP4 decoder, and
all 73 intentionally BF16 transformer tensors remained exact.
- Eight paired 2048-square images were compared with the fine-tune's BF16
reference. Mean CLIPScore was 0.9024 → 0.8688; mean LPIPS was 0.1927. Quantization
can change composition and text rendering. This is a descriptive screen, not
an image-equivalence or human-preference guarantee. See the
[paired images](evaluation/quality-pairs-1.png),
[remaining pairs](evaluation/quality-pairs-2.png) and
[complete measurements](evaluation/mi355-correctness.json).
- Non-default streams, RGBA with alpha spanning 0–255, editing and A–B–A passed.
Repeated prompt embeddings and final denoiser latents were bit-identical.
PNG pixels showed the same small VAE variability seen in BF16 controls:
repeated-image RGB PSNR was approximately 61.2 dB.
- Instrumented MI355 MXFP4 generation took 23.60 seconds for the first tested
2048 request and 21.13–21.52 seconds across seven subsequent, different prompts.
Sampled whole-device peak was 32.31 GiB for that screen and 34.59 GiB across the
mode suite. A 30 GiB framework allocator cap failed during VAE decoding;
successful correctness runs used a 64 GiB cap. These are instrumented
framework results, not R9700 serving measurements.
**This fine-tune has not been run on R9700 in this qualification.** Its tensor
layout matches the existing optimized pipeline, but the original model's R9700
latency, VRAM and low-precision quality measurements do not qualify this new
checkpoint. Its native decoder is reused from the previously qualified release.
## Provenance and license
Source revision: `40319fb15542f0ad22921e0124a191a8a935a60a`.
Source file: `qwen-image-2.1-UC-BF16.gguf`.
SHA-256: `f151c683a8aed4b310777017ebbbe3f2180f1180f7867115171adb7d50b0762a`.
The unchanged original text encoder and VAE come from Qwen revision
`790c92633540aa0cb11d9abf19eb46d861714758`; the encoder reuses our verified balanced
conversion. `result.json` and `checkpoint.lock.json` record the complete layout
and hashes. Conversion tools are included under `tools/` and require the pinned
MI355 Quark environment; inference does not require Quark.
The original [RDNA4 package](https://huggingface.co/EliovpAI/Qwen_Image-2.1-MXFP4-Paiton-RDNA4)
and [portable package](https://huggingface.co/EliovpAI/Qwen_Image-2.1-MXFP4) contain
the same original-model tensor data, with different loaders and metadata. This
repository contains the distinct fine-tune and keeps the Paiton storage format
for plugin compatibility.
The weights remain subject to the [Qwen Research License](LICENSE), which allows
non-commercial research and evaluation; commercial use requires a separate
upstream license. See [NOTICE](NOTICE) and `artifacts/NOTICE` for attribution and
native-runtime terms.
|