How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training", dtype=torch.bfloat16, device_map="cuda")

prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]

Qwen-Image 2.1 NF4 for LoRA Training

Pre-quantized Qwen-Image 2.1 packaged for AcademiaSD Qwen-Image 2.1 LoRAlab, a LoRA trainer for consumer GPUs (8 GB and up). It is downloaded automatically by the LoRAlab; you do not need to fetch it by hand.

This is not an inference checkpoint. The transformer is stored as a per-layer NF4 cache that the LoRAlab rebuilds directly; it does not load with from_pretrained. For generation, use the original model or its ComfyUI release.

What is inside

Folder Size Content
transformer/ 3.87 GB 7B DiT as an NF4 cache: 224 NF4 layers + 8 precision-critical layers in BF16 + norms (others.safetensors)
text_encoder_BF16/ 16.29 GB Qwen3-VL-8B in BF16, exact. lm_head is not stored (only hidden states are used)
text_encoder_INT8/ 8.78 GB Qwen3-VL-8B in LLM.int8
text_encoder_NF4/ 5.45 GB Qwen3-VL-8B in NF4, with lm_head: also used as the dataset auto-captioner
vae/ 1.35 GB Original 64-channel RGBA VAE, unchanged
scheduler/, processor/, model_index.json — Original files, unchanged
LoRAs/ 0.34 GB Viggle Turbo LoRA (4 steps), used for fast training previews

The LoRAlab downloads only the text encoder variant you select.

Transformer (NF4)

  • Every attention and MLP linear of the 32 blocks is NF4 (bitsandbytes, double quantization).
  • img_in, txt_in, the shared modulation, the timestep embedder, norm_out and proj_out stay in BF16: they are small and every block depends on them.
  • Measured against the BF16 transformer on the same inputs: cosine similarity 0.998–0.9996 on the predicted velocity, with practically identical loss.
  • Loads in 2 s and takes **3.9 GB of VRAM**, with no BF16 copy of the model needed.

Text encoder variants

The text encoder of Qwen-Image 2.1 is exactly Qwen3-VL-8B-Instruct (verified tensor by tensor against the official repository). It is only used during pre-caching, never while training.

Variant Peak VRAM Error per token vs BF16 Notes
BF16 CPU offload adapts to free VRAM 0% Default. Layers that do not fit run from RAM
BF16 ~15.5 GB 0% Needs 16+ GB
INT8 ~8.5 GB ~9%
NF4 ~5.5 GB ~24% Also the auto-captioner (~6 GB peak)

Error measured on real captions against the BF16 text encoder, which is what ComfyUI uses at inference. Pre-caching with an exact encoder means the LoRA trains on the same conditioning it will see when you generate.

What the LoRAlab trains with it

  • Characters, objects and styles: one image + caption per sample.
  • Edit LoRAs: name_before / name_after pairs with an imperative instruction (e.g. "make it TOSTIOK style"). The "before" is encoded by Qwen3-VL together with the instruction and prepended as clean latents, exactly as the ComfyUI TextEncodeQwenImage21 node does.

Exported LoRAs use the diffusers/PEFT key format with per-layer alpha and load directly in ComfyUI (verified with ComfyUI's own LoRA loader).

License

Qwen-Image 2.1 and its derivatives are released under the Qwen Research License (see LICENSE): non-commercial use only (research or evaluation). The same license applies to this conversion and to the Viggle Turbo LoRA included in LoRAs/. By downloading these files you accept the terms of that license.

Credits

  • Qwen for Qwen-Image 2.1 and Qwen3-VL.
  • Viggle for the Qwen-Image 2.1 Turbo LoRA.
  • NF4 conversion and LoRAlab by AcademiaSD: YouTube · X · Discord · Ko-Fi

Resumen en español

Qwen-Image 2.1 cuantizado a NF4 para el LoRAlab de AcademiaSD, que lo descarga solo. Transformer de 7B en NF4 (~3,9 GB de VRAM), text encoder Qwen3-VL-8B en tres variantes (BF16 exacto, INT8, NF4; solo se usa en el pre-caché), VAE original y el Turbo LoRA de Viggle para previews rápidas. Sirve para entrenar LoRAs de personajes, objetos, estilos y de edición (pares antes/después). No es un checkpoint de inferencia. Licencia Qwen Research: solo uso no comercial.

Downloads last month
20
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training

Finetuned
(49)
this model