Text-to-Image
Diffusers
Safetensors
QwenImage21Pipeline
qwen-image
qwen-image-2.1
nf4
bitsandbytes
lora-training
image-editing
Instructions to use AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
|
Download README.md from AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training: direct link, hf CLI and curl.
- Browser
- Download file 4.66 kB
-
https://huggingface.co/AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training/resolve/main/README.md
- Command line
-
hf download hf://AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training/README.md
-
curl -L -o README.md https://huggingface.co/AcademiaSD/Qwen-Image-2.1-NF4-for-LoRA-Training/resolve/main/README.md
4.66 kB
| license: other | |
| license_name: qwen-research | |
| license_link: LICENSE | |
| base_model: Qwen/Qwen-Image-2.1 | |
| tags: | |
| - qwen-image | |
| - qwen-image-2.1 | |
| - nf4 | |
| - bitsandbytes | |
| - lora-training | |
| - text-to-image | |
| - image-editing | |
| # Qwen-Image 2.1 NF4 for LoRA Training | |
| Pre-quantized [Qwen-Image 2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) packaged for | |
| **AcademiaSD Qwen-Image 2.1 LoRAlab**, a LoRA trainer for consumer GPUs (8 GB and up). | |
| It is downloaded automatically by the LoRAlab; you do not need to fetch it by hand. | |
| > **This is not an inference checkpoint.** The transformer is stored as a per-layer NF4 cache | |
| > that the LoRAlab rebuilds directly; it does not load with `from_pretrained`. For generation, | |
| > use the original model or its ComfyUI release. | |
| ## What is inside | |
| | Folder | Size | Content | | |
| | :--- | ---: | :--- | | |
| | `transformer/` | 3.87 GB | 7B DiT as an NF4 cache: 224 NF4 layers + 8 precision-critical layers in BF16 + norms (`others.safetensors`) | | |
| | `text_encoder_BF16/` | 16.29 GB | Qwen3-VL-8B in BF16, exact. `lm_head` is not stored (only hidden states are used) | | |
| | `text_encoder_INT8/` | 8.78 GB | Qwen3-VL-8B in LLM.int8 | | |
| | `text_encoder_NF4/` | 5.45 GB | Qwen3-VL-8B in NF4, with `lm_head`: also used as the dataset **auto-captioner** | | |
| | `vae/` | 1.35 GB | Original 64-channel RGBA VAE, unchanged | | |
| | `scheduler/`, `processor/`, `model_index.json` | — | Original files, unchanged | | |
| | `LoRAs/` | 0.34 GB | [Viggle Turbo LoRA](https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo) (4 steps), used for fast training previews | | |
| The LoRAlab downloads only the text encoder variant you select. | |
| ### Transformer (NF4) | |
| - Every attention and MLP linear of the 32 blocks is NF4 (bitsandbytes, double quantization). | |
| - `img_in`, `txt_in`, the shared `modulation`, the timestep embedder, `norm_out` and `proj_out` | |
| stay in BF16: they are small and every block depends on them. | |
| - Measured against the BF16 transformer on the same inputs: cosine similarity **0.998–0.9996** | |
| on the predicted velocity, with practically identical loss. | |
| - Loads in ~2 s and takes **~3.9 GB of VRAM**, with no BF16 copy of the model needed. | |
| ### Text encoder variants | |
| The text encoder of Qwen-Image 2.1 is exactly **Qwen3-VL-8B-Instruct** (verified tensor by tensor | |
| against the official repository). It is only used during pre-caching, never while training. | |
| | Variant | Peak VRAM | Error per token vs BF16 | Notes | | |
| | :--- | :--- | :--- | :--- | | |
| | BF16 CPU offload | adapts to free VRAM | 0% | **Default.** Layers that do not fit run from RAM | | |
| | BF16 | ~15.5 GB | 0% | Needs 16+ GB | | |
| | INT8 | ~8.5 GB | ~9% | | | |
| | NF4 | ~5.5 GB | ~24% | Also the auto-captioner (~6 GB peak) | | |
| Error measured on real captions against the BF16 text encoder, which is what ComfyUI uses at | |
| inference. Pre-caching with an exact encoder means the LoRA trains on the same conditioning it | |
| will see when you generate. | |
| ## What the LoRAlab trains with it | |
| - **Characters, objects and styles**: one image + caption per sample. | |
| - **Edit LoRAs**: `name_before` / `name_after` pairs with an imperative instruction | |
| (e.g. *"make it TOSTIOK style"*). The "before" is encoded by Qwen3-VL together with the | |
| instruction and prepended as clean latents, exactly as the ComfyUI `TextEncodeQwenImage21` node does. | |
| Exported LoRAs use the diffusers/PEFT key format with per-layer `alpha` and load directly in | |
| ComfyUI (verified with ComfyUI's own LoRA loader). | |
| ## License | |
| Qwen-Image 2.1 and its derivatives are released under the **Qwen Research License** | |
| (see `LICENSE`): **non-commercial use only (research or evaluation)**. The same license applies | |
| to this conversion and to the Viggle Turbo LoRA included in `LoRAs/`. By downloading these files | |
| you accept the terms of that license. | |
| ## Credits | |
| - [Qwen](https://huggingface.co/Qwen) for Qwen-Image 2.1 and Qwen3-VL. | |
| - [Viggle](https://huggingface.co/Viggle) for the Qwen-Image 2.1 Turbo LoRA. | |
| - NF4 conversion and LoRAlab by **AcademiaSD**: | |
| [YouTube](https://www.youtube.com/@Academia_SD) · | |
| [X](https://twitter.com/Academia_S_D) · | |
| [Discord](https://discord.gg/Syuaduy678) · | |
| [Ko-Fi](https://ko-fi.com/academiasd) | |
| --- | |
| ### Resumen en español | |
| Qwen-Image 2.1 cuantizado a NF4 para el **LoRAlab de AcademiaSD**, que lo descarga solo. | |
| Transformer de 7B en NF4 (~3,9 GB de VRAM), text encoder Qwen3-VL-8B en tres variantes | |
| (BF16 exacto, INT8, NF4; solo se usa en el pre-caché), VAE original y el Turbo LoRA de Viggle | |
| para previews rápidas. Sirve para entrenar LoRAs de personajes, objetos, estilos y de edición | |
| (pares antes/después). No es un checkpoint de inferencia. Licencia Qwen Research: | |
| **solo uso no comercial**. | |