--- license: other license_name: qwen-research license_link: LICENSE base_model: Qwen/Qwen-Image-2.1 tags: - qwen-image - qwen-image-2.1 - nf4 - bitsandbytes - lora-training - text-to-image - image-editing --- # Qwen-Image 2.1 NF4 for LoRA Training Pre-quantized [Qwen-Image 2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) packaged for **AcademiaSD Qwen-Image 2.1 LoRAlab**, a LoRA trainer for consumer GPUs (8 GB and up). It is downloaded automatically by the LoRAlab; you do not need to fetch it by hand. > **This is not an inference checkpoint.** The transformer is stored as a per-layer NF4 cache > that the LoRAlab rebuilds directly; it does not load with `from_pretrained`. For generation, > use the original model or its ComfyUI release. ## What is inside | Folder | Size | Content | | :--- | ---: | :--- | | `transformer/` | 3.87 GB | 7B DiT as an NF4 cache: 224 NF4 layers + 8 precision-critical layers in BF16 + norms (`others.safetensors`) | | `text_encoder_BF16/` | 16.29 GB | Qwen3-VL-8B in BF16, exact. `lm_head` is not stored (only hidden states are used) | | `text_encoder_INT8/` | 8.78 GB | Qwen3-VL-8B in LLM.int8 | | `text_encoder_NF4/` | 5.45 GB | Qwen3-VL-8B in NF4, with `lm_head`: also used as the dataset **auto-captioner** | | `vae/` | 1.35 GB | Original 64-channel RGBA VAE, unchanged | | `scheduler/`, `processor/`, `model_index.json` | — | Original files, unchanged | | `LoRAs/` | 0.34 GB | [Viggle Turbo LoRA](https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo) (4 steps), used for fast training previews | The LoRAlab downloads only the text encoder variant you select. ### Transformer (NF4) - Every attention and MLP linear of the 32 blocks is NF4 (bitsandbytes, double quantization). - `img_in`, `txt_in`, the shared `modulation`, the timestep embedder, `norm_out` and `proj_out` stay in BF16: they are small and every block depends on them. - Measured against the BF16 transformer on the same inputs: cosine similarity **0.998–0.9996** on the predicted velocity, with practically identical loss. - Loads in ~2 s and takes **~3.9 GB of VRAM**, with no BF16 copy of the model needed. ### Text encoder variants The text encoder of Qwen-Image 2.1 is exactly **Qwen3-VL-8B-Instruct** (verified tensor by tensor against the official repository). It is only used during pre-caching, never while training. | Variant | Peak VRAM | Error per token vs BF16 | Notes | | :--- | :--- | :--- | :--- | | BF16 CPU offload | adapts to free VRAM | 0% | **Default.** Layers that do not fit run from RAM | | BF16 | ~15.5 GB | 0% | Needs 16+ GB | | INT8 | ~8.5 GB | ~9% | | | NF4 | ~5.5 GB | ~24% | Also the auto-captioner (~6 GB peak) | Error measured on real captions against the BF16 text encoder, which is what ComfyUI uses at inference. Pre-caching with an exact encoder means the LoRA trains on the same conditioning it will see when you generate. ## What the LoRAlab trains with it - **Characters, objects and styles**: one image + caption per sample. - **Edit LoRAs**: `name_before` / `name_after` pairs with an imperative instruction (e.g. *"make it TOSTIOK style"*). The "before" is encoded by Qwen3-VL together with the instruction and prepended as clean latents, exactly as the ComfyUI `TextEncodeQwenImage21` node does. Exported LoRAs use the diffusers/PEFT key format with per-layer `alpha` and load directly in ComfyUI (verified with ComfyUI's own LoRA loader). ## License Qwen-Image 2.1 and its derivatives are released under the **Qwen Research License** (see `LICENSE`): **non-commercial use only (research or evaluation)**. The same license applies to this conversion and to the Viggle Turbo LoRA included in `LoRAs/`. By downloading these files you accept the terms of that license. ## Credits - [Qwen](https://huggingface.co/Qwen) for Qwen-Image 2.1 and Qwen3-VL. - [Viggle](https://huggingface.co/Viggle) for the Qwen-Image 2.1 Turbo LoRA. - NF4 conversion and LoRAlab by **AcademiaSD**: [YouTube](https://www.youtube.com/@Academia_SD) · [X](https://twitter.com/Academia_S_D) · [Discord](https://discord.gg/Syuaduy678) · [Ko-Fi](https://ko-fi.com/academiasd) --- ### Resumen en español Qwen-Image 2.1 cuantizado a NF4 para el **LoRAlab de AcademiaSD**, que lo descarga solo. Transformer de 7B en NF4 (~3,9 GB de VRAM), text encoder Qwen3-VL-8B en tres variantes (BF16 exacto, INT8, NF4; solo se usa en el pre-caché), VAE original y el Turbo LoRA de Viggle para previews rápidas. Sirve para entrenar LoRAs de personajes, objetos, estilos y de edición (pares antes/después). No es un checkpoint de inferencia. Licencia Qwen Research: **solo uso no comercial**.