Difference with base model?

#20
by user694201337 - opened

Can you explain what is different between this version and the base model?

Here is a breakdown of the key differences between this repository and the upstream base model:

1. Uncensored / No Content Restrictions

  • No Safety Checker / Refusals: The default/hosted pipelines for Qwen-Image typically enforce strict safety filters, prompt refusals, or post-generation checkers that blur or black out sensitive, NSFW, or adult imagery.
  • This release removes those restrictions, allowing the model to generate imagery directly based on your prompt without built-in refusals or output blocking.

2. Quantization & Low VRAM Consumption (GGUF, FP8 & INT8)

  • Base Model Footprint: The original upstream model in BF16 requires ~14.2 GB of VRAM for the transformer alone, plus ~17.5 GB for the text encoder, making it difficult to run on consumer hardware.
  • Optimized Quantizations: In this repo, the weights are converted into:
    • GGUF formats (Q4_K_M, Q5_K_M, Q6_K, Q8_0): Letting you run the diffusion model smoothly in 4.6 GB – 7.6 GB of VRAM using ComfyUI-GGUF.
    • FP8 & INT8 ConvRot Safetensors: Native formats for RTX 40-series (FP8) and RTX 20/30-series (INT8 ConvRot) that run without quality loss.

3. Upstream Base Weight Integrity

  • The core weights are derived directly from the official Qwen/Qwen-Image-2.1 architecture and parameters.
  • There are no fine-tuning artifacts, style shifts, or quality degradation—prompt understanding, composition, and visual fidelity match the official base model.

4. Packaged for ComfyUI Out-of-the-Box

  • Companion models are provided in optimal formats (e.g., qwen3vl_8b_int8_convrot text encoder that runs comfortably in RAM, and the BF16 VAE), making it fully plug-and-play for local ComfyUI workflows.

In short: you get the full quality of the original Qwen-Image 2.1 base model, but uncensored and optimized to run efficiently on local consumer GPUs.

Can you respond yourself instead of with an LLM written response? I am asking what is actually different from the base weights, of course an API version has restrictions but the base weights that are available already seem pretty much uncensored to me, did you actually do anything to change the weights besides a conversion?

Can you respond yourself instead of with an LLM written response? I am asking what is actually different from the base weights, of course an API version has restrictions but the base weights that are available already seem pretty much uncensored to me, did you actually do anything to change the weights besides a conversion?

sure, I curated a targeted dataset, fine-tuned and tested the model with LoRA, then merged the improvements into the base weights to create an uncensored version of the model.

Has the text encoder remained unchanged compared to the original model? In other words, can i use the original text encoder and still get uncensored output?

Has the text encoder remained unchanged compared to the original model? In other words, can i use the original text encoder and still get uncensored output?

yes, you don't need to.

Can you respond yourself instead of with an LLM written response? I am asking what is actually different from the base weights, of course an API version has restrictions but the base weights that are available already seem pretty much uncensored to me, did you actually do anything to change the weights besides a conversion?

sure, I curated a targeted dataset, fine-tuned and tested the model with LoRA, then merged the improvements into the base weights to create an uncensored version of the model.

That is not uncensoring but finetuning, uncensoring implies the removal of some form of refusal in the model. You should specify this in the model card, with maybe some information on how you trained

unlike LLMs, diffusion models don't have built-in refusal mechanisms since the censorship is handled at the training dataset level, thanks for feedback.

unlike LLMs, diffusion models don't have built-in refusal mechanisms since the censorship is handled at the training dataset level, thanks for feedback.

Ideogram v4 hasn't got that message btw.
otherwise it's often true.

However uncensoring can well need finetuning. Idk yet, where qi2.1 stands here but krea2 or flux2 definitely are heavily censored beyond just the text encoder (which usully does not censor in generation. I can use the standard krea2 text encoder for full on nsfw with Kroma (lora or full merge), without issues (clip! not for prompt enhancing).

they do not "refuse" flatout, but they change things from your prompt (Krea2 to a ridiculous level of sanitization) or add things that are explicitly (!) unprompted.. That IS censorship and it's in the model itself. And it's not just lack of training (if the censor things like certain anatomy, they likely also not train it.)

But whether this finetune is more than a baked in specific lora is the question. For a broad "uncensored claim, you need to do a LOT of work not just a txxty lora.. Think Chroma1 for Flux.1 or Kroma for KRea2. Or dasiwa, 10eros for LTX and H3. They don't just remove censorship, they also add the missing ... bits to the knowledge.

They do not come out days after a model release.

unlike LLMs, diffusion models don't have built-in refusal mechanisms since the censorship is handled at the training dataset level, thanks for feedback.

Look into methods of how Krea2 was uncensored, that has some form of refusal/censoring built in by default which can be bypassed which increases the capabilities of the base model without training new data into the model. This does not seem to be the case with Qwen2.1

Look into methods of how Krea2 was uncensored, that has some form of refusal/censoring built in by default which can be bypassed which increases the capabilities of the base model without training new data into the model. This does not seem to be the case with Qwen2.1

each organization may take different approaches or require custom solutions, but we’re essentially talking about the same thing :)

bit-identical with the base,

  • All 73 BF16 tensors
    • Includes: img_in, modulation.1, norm_out.linear, proj_out, txt_in.*, time_text_embed.*, and all 64 norm_q/k
    • Result: Bit-identical

am I missing something? there are no signs that anything has been merged into the models you uploaded

With the same settings its does generate different images than the base model but the difference is marginal. Ive not seen any significant variance in image output and ive tried pretty much all sort of prompts that might reasonably be censored.

image
I FACE THIS ISSUE

Trying to work out why I would use this one over the unlsoth one. Unsloths one can also generate NSFW images, so unless they made theirs uncensored too, I'm going to assume the base model can do NSFW images already.

With my testing of both the base and this, the difference doesn't exist.

yeah, this is basically a fake. Best case they are ok gguf versions of the model. Nothing that would constitute uncensoring seems to have been done. QI2.1 (and og QI) is just out of the box relatively flexible.

Also after testing it a bit, QI2.1 is a LOT less censored then krea2. While it can't make nsfw without help, it is far less prone to change things that are prompted. If we are not talking full on nudity and stuff, it is basically unrestricted.
While KRea2 may just do nothing the prompt actually says because it is guardrailed into oblivion. Or there is another issue in there. Whatever it is, Krea2 out of the box is esentiall useless. While QI2.1 is not.
Which is where ACTUALLY uncensored versions come in. And they are still wip after a long time of Krea2 out and quite popular.
The release of a "uncensored model" within hours of model release is just objectively impossible.

Could the owner address @rzgar ’s finding directly?

If 73 BF16 tensors are bit-identical to the base, which specific tensors were actually changed by the merged LoRA? A few hashes or layer names would probably settle this.

bit-identical with the base,

  • All 73 BF16 tensors
    • Includes: img_in, modulation.1, norm_out.linear, proj_out, txt_in.*, time_text_embed.*, and all 64 norm_q/k
    • Result: Bit-identical

am I missing something? there are no signs that anything has been merged into the models you uploaded

Additionally, I already clearly explained above what I did: the model has no censored layer, and this is achieved on a dataset level.

Regarding the 73 tensors: you're simply looking at the wrong layers. When quantizing to Q4_K_M, small 1D tensors like norm_q/k and embeddings are kept in BF16 for stability, while the actual 2D weights are quantized. My LoRA only modified the 128 attention projections (to_q, to_k, to_v, to_out.0 across all 32 blocks)—I intentionally left the rest untouched so the model wouldn't lose its general capabilities. Those 128 modified layers became Q4_K and Q6_K, which is why you didn't see them when filtering only for BF16 tensors.

If you want to verify the merged weights, check the attention layers directly or grab the full qwen-image-2.1-UC-BF16.gguf file. You’ll see the weight deltas on all 128 layers.

With my testing of both the base and this, the difference doesn't exist.

yeah, this is basically a fake. Best case they are ok gguf versions of the model. Nothing that would constitute uncensoring seems to have been done. QI2.1 (and og QI) is just out of the box relatively flexible.

Also after testing it a bit, QI2.1 is a LOT less censored then krea2. While it can't make nsfw without help, it is far less prone to change things that are prompted. If we are not talking full on nudity and stuff, it is basically unrestricted.
While KRea2 may just do nothing the prompt actually says because it is guardrailed into oblivion. Or there is another issue in there. Whatever it is, Krea2 out of the box is esentiall useless. While QI2.1 is not.
Which is where ACTUALLY uncensored versions come in. And they are still wip after a long time of Krea2 out and quite popular.
The release of a "uncensored model" within hours of model release is just objectively impossible.

a lot of things were unlocked in this uncensored version, so I wouldn't jump to conclusions without testing properly.

if it doesn't work for you, you're always free not to use it.
if there are no other questions, you can close this thread.

have a great one!

perhaps you could present some examples nobody else seems to be able to find.

fake uncensored, only README got finetuned.

perhaps you could present some examples nobody else seems to be able to find.

fake uncensored, only README got finetuned.

I won't do this for ethical reasons; it's up to you.

Keep trolling without any basis 😉

I am ending the discussion here; I will wait for @rzgar 's reply.

I'm just not personally seeing the point. The qwen-image-2.1 that I am using, nothing changed with it, using with stable-diffusion.cpp, will happilly generate full on nudity images, etc... Which says to me the base can already do this without any trickery, etc...

I'm just not personally seeing the point. The qwen-image-2.1 that I am using, nothing changed with it, using with stable-diffusion.cpp, will happilly generate full on nudity images, etc... Which says to me the base can already do this without any trickery, etc...

that’s exactly what I meant to, I was just tried to make the fully uncensored one.

i have one correction to make: The base model (official comfy repo bf16 and int8convrot) literally creates full nudity, At least on the male side, with all detail. Not perfect of course , but fairly extensive and well beyond the usual placeholders even of less restrictive models.

So your ggufs are actually (relatively) uncensored ... because the model is. So i guess technically you are correct. Only your work ended on making ggufs (which are not fake but real) . Same true for ALL competently made ggufs. But you have not uncensored anything, finetuned anything, trained anything that wasn't there before or anything away that was. THAT is how you uncensor a image model, not by removing non existent "refusals".

(except ideogram v4 which can literally put a "i can't do that" in the image).

i have one correction to make: The base model (official comfy repo bf16 and int8convrot) literally creates full nudity, At least on the male side, with all detail. Not perfect of course , but fairly extensive and well beyond the usual placeholders even of less restrictive models.

So your ggufs are actually (relatively) uncensored ... because the model is. So i guess technically you are correct. Only your work ended on making ggufs (which are not fake but real) . Same true for ALL competently made ggufs. But you have not uncensored anything, finetuned anything, trained anything that wasn't there before or anything away that was. THAT is how you uncensor a image model, not by removing non existent "refusals".

(except ideogram v4 which can literally put a "i can't do that" in the image).

the base model heavily suppresses many things - frankly, claiming this is "just a GGUF conversion" is getting bothering—there was extensive training, fine-tuning, and testing behind this work.

let's not prolong this discussion, shall we?

@abenzerps , providing proof of the "extensive" training and fine tuning isnt that hard, just load adapter files and people will have no questions left to ask, actually, any proof, anything, a screenshot of the training logs, hyperparameters, loading base vs uncensored example in google drive and providing a link to it, yes, you don't have to post images directly there since google drive is real. Furthermore, I tested the base and UC in generating unsafe content specifically with the prompts that as you claim cause model to suppress many things, lets call that sanitization, and the received outputs from both models were absolutely identical in what did they sanitize, e.g this is not uncensored; this is just two letters added to the file name it seems. I also specifically made a test with a prompt that does make the base generate unsafe output, well well, guess what? the output of the uncensored version was completely identical to the base and didnt add any additional detail. Why are you trying so hard to argue with people when the easiest thing to do would be just to provide any proof example from the above and ignore people who claim you're lying? 😆

Maybe "uncensored" is just a key of huge download numbers

@abenzerps , providing proof of the "extensive" training and fine tuning isnt that hard, just load adapter files and people will have no questions left to ask, actually, any proof, anything, a screenshot of the training logs, hyperparameters, loading base vs uncensored example in google drive and providing a link to it, yes, you don't have to post images directly there since google drive is real. Furthermore, I tested the base and UC in generating unsafe content specifically with the prompts that as you claim cause model to suppress many things, lets call that sanitization, and the received outputs from both models were absolutely identical in what did they sanitize, e.g this is not uncensored; this is just two letters added to the file name it seems. I also specifically made a test with a prompt that does make the base generate unsafe output, well well, guess what? the output of the uncensored version was completely identical to the base and didnt add any additional detail. Why are you trying so hard to argue with people when the easiest thing to do would be just to provide any proof example from the above and ignore people who claim you're lying? 😆

I’ll do that, I’m currently running Ornith-1.5-9B GSQ-RCO tests on my computer. Once I’m done, I can proceed as you suggested, however, I won’t be sharing any images or image links, as that would violate the community guidelines regardless.

also, as I mentioned above, I’m waiting for rzgar’s response. I take everyone who asks genuine questions seriously.

Training Setup & Runtime Console Telemetry

fine-tuning was performed directly on the DiT attention layers to neutralize safety/censorship boundaries while strictly retaining base fidelity. Here is the authentic console telemetry from the example run:

Loading models from: [
    "transformer/diffusion_pytorch_model-00002-of-00002.safetensors",
    "transformer/diffusion_pytorch_model-00001-of-00002.safetensors"
]
Loaded model: {
    "model_name": "qwen_image_21_dit",
    "model_class": "diffsynth.models.qwen_image_21_dit.QwenImage21DiT"
}
LoRA state initialized.
Trainable LoRA parameters: 16,777,216
Active targets: ['to_q', 'to_k', 'to_v', 'to_out.0'] across 32 transformer blocks
[Optimizer: AdamW | lr=1e-5 | weight_decay=0.01 | Effective Batch: 4]
Initial State   | Loss: 1.480639 | GradNorm: 0.0035 | NonZeroTensors: 256/256 | Peak VRAM: 20.26 GB
Optimization... | Loss: 1.147549 | GradNorm: 0.0039 | NonZeroTensors: 256/256 | Peak VRAM: 20.33 GB
Optimization... | Loss: 0.874264 | GradNorm: 0.0084 | NonZeroTensors: 256/256 | Peak VRAM: 20.33 GB
Optimization... | Loss: 1.058703 | GradNorm: 0.0051 | NonZeroTensors: 256/256 | Peak VRAM: 20.33 GB
Final State     | Loss: 1.261441 | GradNorm: 0.0051 | NonZeroTensors: 256/256 | Peak VRAM: 20.33 GB
Training completed cleanly. Zero NaN/Inf anomalies encountered.
Saved adapter weights -> adapter_model.safetensors
SHA256: 6c492e05273007e4fe451cca3dbfef313e3e9cc9a81f1b9386ec191b7f9e70a7

I appreciate the response, I don't really understand why you decided to confront the critics that way if you knew for sure you were right, but that no longer matters once there are adapter files in the repo, that would be an undeniable proof that you were telling the truth, looking forward to it.

@abenzerps , providing proof of the "extensive" training and fine tuning isnt that hard, just load adapter files and people will have no questions left to ask, actually, any proof, anything, a screenshot of the training logs, hyperparameters, loading base vs uncensored example in google drive and providing a link to it, yes, you don't have to post images directly there since google drive is real. Furthermore, I tested the base and UC in generating unsafe content specifically with the prompts that as you claim cause model to suppress many things, lets call that sanitization, and the received outputs from both models were absolutely identical in what did they sanitize, e.g this is not uncensored; this is just two letters added to the file name it seems. I also specifically made a test with a prompt that does make the base generate unsafe output, well well, guess what? the output of the uncensored version was completely identical to the base and didnt add any additional detail. Why are you trying so hard to argue with people when the easiest thing to do would be just to provide any proof example from the above and ignore people who claim you're lying? 😆

I’ll do that, I’m currently running Ornith-1.5-9B GSQ-RCO tests on my computer. Once I’m done, I can proceed as you suggested, however, I won’t be sharing any images or image links, as that would violate the community guidelines regardless.

also, as I mentioned above, I’m waiting for rzgar’s response. I take everyone who asks genuine questions seriously.

NSFW images and prompts dont necessarily violate hf guidelines. Youd absolutely be able to provide specific prompts that proof your model varies from the base model in a significant way.
See here: https://huggingface.co/content-policy

I appreciate the response, I don't really understand why you decided to confront the critics that way if you knew for sure you were right, but that no longer matters once there are adapter files in the repo, that would be an undeniable proof that you were telling the truth, looking forward to it.

the training/lora data isn’t easy for casual users to interpret, and their own tests may be too limited to show the difference., as I’ve explained several times, sharing generated images to demonstrate the uncensored changes wouldn’t be ethical. this is not 4chan.

•
This comment has been hidden (marked as Abuse)
•
This comment has been hidden (marked as Abuse)

I appreciate the response, I don't really understand why you decided to confront the critics that way if you knew for sure you were right, but that no longer matters once there are adapter files in the repo, that would be an undeniable proof that you were telling the truth, looking forward to it.

the training/lora data isn’t easy for casual users to interpret, and their own tests may be too limited to show the difference., as I’ve explained several times, sharing generated images to demonstrate the uncensored changes wouldn’t be ethical. this is not 4chan.

Pardon, how is that related to my question about sharing adapter safetensor file?... They are not heavy, the adapter safetensor file for this LoRA should be as small as 66-70 megabytes. I assume you were responding to ReingeFallen, but then it is even more weird that you create an allegedly uncensored finetune and that is not unethical to you, yet making a comparison between base and finetuned is unethical somehow? 😃

Pardon, how is that related to my question about sharing adapter safetensor file?... They are not heavy, the adapter safetensor file for this LoRA should be as small as 66-70 megabytes. I assume you were responding to ReingeFallen, but then it is even more weird that you create an allegedly uncensored model and that is not unethical to you, yet making a comparison between base and finetuned is unethical somehow? 😃

no, I'm talking about sharing images, I'll sort everything out; wait for my reply please.

You dont have to share an image. Just post one prompt that produces images that are less censored than the base model.

I uploaded the standalone LoRA adapter https://huggingface.co/abenzerps/Qwen-Image-2.1-Uncensored-GGUF/blob/main/qwen-image-2.1-uncensored-lora.safetensors to the repository files.
It was trained with DiffSynth-Studio on DiT attention layers.

1000,000 downloads is just a joke , so ironic

Sign up or log in to comment