AcademiaSD's picture
Upload README.md
e94eadc verified
|
Raw History Blame Contribute Delete
4.66 kB
---
license: other
license_name: qwen-research
license_link: LICENSE
base_model: Qwen/Qwen-Image-2.1
tags:
- qwen-image
- qwen-image-2.1
- nf4
- bitsandbytes
- lora-training
- text-to-image
- image-editing
---
# Qwen-Image 2.1 NF4 for LoRA Training
Pre-quantized [Qwen-Image 2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) packaged for
**AcademiaSD Qwen-Image 2.1 LoRAlab**, a LoRA trainer for consumer GPUs (8 GB and up).
It is downloaded automatically by the LoRAlab; you do not need to fetch it by hand.
> **This is not an inference checkpoint.** The transformer is stored as a per-layer NF4 cache
> that the LoRAlab rebuilds directly; it does not load with `from_pretrained`. For generation,
> use the original model or its ComfyUI release.
## What is inside
| Folder | Size | Content |
| :--- | ---: | :--- |
| `transformer/` | 3.87 GB | 7B DiT as an NF4 cache: 224 NF4 layers + 8 precision-critical layers in BF16 + norms (`others.safetensors`) |
| `text_encoder_BF16/` | 16.29 GB | Qwen3-VL-8B in BF16, exact. `lm_head` is not stored (only hidden states are used) |
| `text_encoder_INT8/` | 8.78 GB | Qwen3-VL-8B in LLM.int8 |
| `text_encoder_NF4/` | 5.45 GB | Qwen3-VL-8B in NF4, with `lm_head`: also used as the dataset **auto-captioner** |
| `vae/` | 1.35 GB | Original 64-channel RGBA VAE, unchanged |
| `scheduler/`, `processor/`, `model_index.json` | — | Original files, unchanged |
| `LoRAs/` | 0.34 GB | [Viggle Turbo LoRA](https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo) (4 steps), used for fast training previews |
The LoRAlab downloads only the text encoder variant you select.
### Transformer (NF4)
- Every attention and MLP linear of the 32 blocks is NF4 (bitsandbytes, double quantization).
- `img_in`, `txt_in`, the shared `modulation`, the timestep embedder, `norm_out` and `proj_out`
stay in BF16: they are small and every block depends on them.
- Measured against the BF16 transformer on the same inputs: cosine similarity **0.998–0.9996**
on the predicted velocity, with practically identical loss.
- Loads in ~2 s and takes **~3.9 GB of VRAM**, with no BF16 copy of the model needed.
### Text encoder variants
The text encoder of Qwen-Image 2.1 is exactly **Qwen3-VL-8B-Instruct** (verified tensor by tensor
against the official repository). It is only used during pre-caching, never while training.
| Variant | Peak VRAM | Error per token vs BF16 | Notes |
| :--- | :--- | :--- | :--- |
| BF16 CPU offload | adapts to free VRAM | 0% | **Default.** Layers that do not fit run from RAM |
| BF16 | ~15.5 GB | 0% | Needs 16+ GB |
| INT8 | ~8.5 GB | ~9% | |
| NF4 | ~5.5 GB | ~24% | Also the auto-captioner (~6 GB peak) |
Error measured on real captions against the BF16 text encoder, which is what ComfyUI uses at
inference. Pre-caching with an exact encoder means the LoRA trains on the same conditioning it
will see when you generate.
## What the LoRAlab trains with it
- **Characters, objects and styles**: one image + caption per sample.
- **Edit LoRAs**: `name_before` / `name_after` pairs with an imperative instruction
(e.g. *"make it TOSTIOK style"*). The "before" is encoded by Qwen3-VL together with the
instruction and prepended as clean latents, exactly as the ComfyUI `TextEncodeQwenImage21` node does.
Exported LoRAs use the diffusers/PEFT key format with per-layer `alpha` and load directly in
ComfyUI (verified with ComfyUI's own LoRA loader).
## License
Qwen-Image 2.1 and its derivatives are released under the **Qwen Research License**
(see `LICENSE`): **non-commercial use only (research or evaluation)**. The same license applies
to this conversion and to the Viggle Turbo LoRA included in `LoRAs/`. By downloading these files
you accept the terms of that license.
## Credits
- [Qwen](https://huggingface.co/Qwen) for Qwen-Image 2.1 and Qwen3-VL.
- [Viggle](https://huggingface.co/Viggle) for the Qwen-Image 2.1 Turbo LoRA.
- NF4 conversion and LoRAlab by **AcademiaSD**:
[YouTube](https://www.youtube.com/@Academia_SD) ·
[X](https://twitter.com/Academia_S_D) ·
[Discord](https://discord.gg/Syuaduy678) ·
[Ko-Fi](https://ko-fi.com/academiasd)
---
### Resumen en español
Qwen-Image 2.1 cuantizado a NF4 para el **LoRAlab de AcademiaSD**, que lo descarga solo.
Transformer de 7B en NF4 (~3,9 GB de VRAM), text encoder Qwen3-VL-8B en tres variantes
(BF16 exacto, INT8, NF4; solo se usa en el pre-caché), VAE original y el Turbo LoRA de Viggle
para previews rápidas. Sirve para entrenar LoRAs de personajes, objetos, estilos y de edición
(pares antes/después). No es un checkpoint de inferencia. Licencia Qwen Research:
**solo uso no comercial**.