Instructions to use AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Qwen-Image 2.1 NF4 for LoRA Training
Pre-quantized Qwen-Image 2.1 packaged for AcademiaSD Qwen-Image 2.1 LoRAlab, a LoRA trainer for consumer GPUs (8 GB and up). It is downloaded automatically by the LoRAlab; you do not need to fetch it by hand.
This is not an inference checkpoint. The transformer is stored as a per-layer NF4 cache that the LoRAlab rebuilds directly; it does not load with
from_pretrained. For generation, use the original model or its ComfyUI release.
What is inside
| Folder | Size | Content |
|---|---|---|
transformer/ |
3.87 GB | 7B DiT as an NF4 cache: 224 NF4 layers + 8 precision-critical layers in BF16 + norms (others.safetensors) |
text_encoder_BF16/ |
16.29 GB | Qwen3-VL-8B in BF16, exact. lm_head is not stored (only hidden states are used) |
text_encoder_INT8/ |
8.78 GB | Qwen3-VL-8B in LLM.int8 |
text_encoder_NF4/ |
5.45 GB | Qwen3-VL-8B in NF4, with lm_head: also used as the dataset auto-captioner |
vae/ |
1.35 GB | Original 64-channel RGBA VAE, unchanged |
scheduler/, processor/, model_index.json |
— | Original files, unchanged |
LoRAs/ |
0.34 GB | Viggle Turbo LoRA (4 steps), used for fast training previews |
The LoRAlab downloads only the text encoder variant you select.
Transformer (NF4)
- Every attention and MLP linear of the 32 blocks is NF4 (bitsandbytes, double quantization).
img_in,txt_in, the sharedmodulation, the timestep embedder,norm_outandproj_outstay in BF16: they are small and every block depends on them.- Measured against the BF16 transformer on the same inputs: cosine similarity 0.998–0.9996 on the predicted velocity, with practically identical loss.
- Loads in
2 s and takes **3.9 GB of VRAM**, with no BF16 copy of the model needed.
Text encoder variants
The text encoder of Qwen-Image 2.1 is exactly Qwen3-VL-8B-Instruct (verified tensor by tensor against the official repository). It is only used during pre-caching, never while training.
| Variant | Peak VRAM | Error per token vs BF16 | Notes |
|---|---|---|---|
| BF16 CPU offload | adapts to free VRAM | 0% | Default. Layers that do not fit run from RAM |
| BF16 | ~15.5 GB | 0% | Needs 16+ GB |
| INT8 | ~8.5 GB | ~9% | |
| NF4 | ~5.5 GB | ~24% | Also the auto-captioner (~6 GB peak) |
Error measured on real captions against the BF16 text encoder, which is what ComfyUI uses at inference. Pre-caching with an exact encoder means the LoRA trains on the same conditioning it will see when you generate.
What the LoRAlab trains with it
- Characters, objects and styles: one image + caption per sample.
- Edit LoRAs:
name_before/name_afterpairs with an imperative instruction (e.g. "make it TOSTIOK style"). The "before" is encoded by Qwen3-VL together with the instruction and prepended as clean latents, exactly as the ComfyUITextEncodeQwenImage21node does.
Exported LoRAs use the diffusers/PEFT key format with per-layer alpha and load directly in
ComfyUI (verified with ComfyUI's own LoRA loader).
License
Qwen-Image 2.1 and its derivatives are released under the Qwen Research License
(see LICENSE): non-commercial use only (research or evaluation). The same license applies
to this conversion and to the Viggle Turbo LoRA included in LoRAs/. By downloading these files
you accept the terms of that license.
Credits
- Qwen for Qwen-Image 2.1 and Qwen3-VL.
- Viggle for the Qwen-Image 2.1 Turbo LoRA.
- NF4 conversion and LoRAlab by AcademiaSD: YouTube · X · Discord · Ko-Fi
Resumen en español
Qwen-Image 2.1 cuantizado a NF4 para el LoRAlab de AcademiaSD, que lo descarga solo. Transformer de 7B en NF4 (~3,9 GB de VRAM), text encoder Qwen3-VL-8B en tres variantes (BF16 exacto, INT8, NF4; solo se usa en el pre-caché), VAE original y el Turbo LoRA de Viggle para previews rápidas. Sirve para entrenar LoRAs de personajes, objetos, estilos y de edición (pares antes/después). No es un checkpoint de inferencia. Licencia Qwen Research: solo uso no comercial.
- Downloads last month
- 19
Model tree for AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training
Base model
Qwen/Qwen-Image-2.1