Upload outpaint LoRA
Browse files
README.md
ADDED
|
@@ -0,0 +1,114 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: qwen-research-license
|
| 4 |
+
base_model: Qwen/Qwen-Image-2.1
|
| 5 |
+
pipeline_tag: image-to-image
|
| 6 |
+
tags:
|
| 7 |
+
- lora
|
| 8 |
+
- outpainting
|
| 9 |
+
- uncrop
|
| 10 |
+
- qwen-image
|
| 11 |
+
- qwen-image-2.1
|
| 12 |
+
- comfyui
|
| 13 |
+
- ai-toolkit
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# Qwen Image 2.1 Outpaint LoRA
|
| 17 |
+
|
| 18 |
+
Extends a picture in **any direction**: one side, two sides, a corner, three
|
| 19 |
+
sides, all four, or a small picture on a big empty canvas. Pad the picture with
|
| 20 |
+
flat gray `#808080`, give the padded canvas to Qwen Image 2.1 as the reference,
|
| 21 |
+
and the LoRA fills the gray with a continuation of the scene while keeping
|
| 22 |
+
the original picture where it is.
|
| 23 |
+
|
| 24 |
+

|
| 25 |
+
|
| 26 |
+
*Held-out pictures (never trained on), rendered in ComfyUI with the INT8 Qwen
|
| 27 |
+
Image 2.1 model, 25 steps, CFG 1: source | base model | LoRA step 500 | 1250 |
|
| 28 |
+
1500 | original picture.*
|
| 29 |
+
|
| 30 |
+
## Files
|
| 31 |
+
|
| 32 |
+
| file | notes |
|
| 33 |
+
|---|---|
|
| 34 |
+
| `qwen-image-2.1-outpaint.safetensors` | **use this**: step 1500 |
|
| 35 |
+
| `checkpoints/qwen21_outpaint_v1_000001250.safetensors` | step 1250, practically identical |
|
| 36 |
+
| `checkpoints/qwen21_outpaint_v1_000000500.safetensors` | step 500, already fixes framing; plainer fills |
|
| 37 |
+
|
| 38 |
+
Rank 32, ComfyUI key format (`diffusion_model.transformer_blocks.*`), all 384
|
| 39 |
+
tensors load onto the Comfy-Org Qwen Image 2.1 weights.
|
| 40 |
+
|
| 41 |
+
## Prompt
|
| 42 |
+
|
| 43 |
+
The instruction is the trigger. Put it first:
|
| 44 |
+
|
| 45 |
+
```
|
| 46 |
+
Outpaint the image: replace the solid gray areas with a seamless continuation of the scene, keeping the existing picture unchanged.
|
| 47 |
+
```
|
| 48 |
+
|
| 49 |
+
Optionally follow it with `Scene: <description>`. The training captions
|
| 50 |
+
described the whole picture, so a description of the source (what is visible)
|
| 51 |
+
works well; in ComfyUI, **Text Generate** with the Qwen3-VL 8B encoder you
|
| 52 |
+
already load for Qwen 2.1 can write it automatically from the picture.
|
| 53 |
+
|
| 54 |
+
## ComfyUI
|
| 55 |
+
|
| 56 |
+
1. Pad the picture with flat gray `#808080` (AusBoss **Load Image + Pad**:
|
| 57 |
+
fill `color`, `#808080`, canvas multiple 32, target 1.0 MP).
|
| 58 |
+
2. **Text Encode Qwen Image 2.1**: the padded canvas as `image_1`,
|
| 59 |
+
`resolution` 0 (reference and output share the canvas size), plus the prompt.
|
| 60 |
+
3. **KSampler** on the encoder's `latent` output: 25 steps, CFG 1,
|
| 61 |
+
`euler` / `simple`, denoise 1.
|
| 62 |
+
4. **VAE Decode** -> **Split Image with Alpha** (the Qwen 2.1 VAE decodes RGBA)
|
| 63 |
+
-> AusBoss **Stitch Inpaint** to paste the exact original pixels back over a
|
| 64 |
+
32 px feather.
|
| 65 |
+
|
| 66 |
+
Do **not** pin the known area with *Set Latent Noise Mask*: on Qwen 2.1 that
|
| 67 |
+
draws a visible rectangle at the seam. The LoRA keeps the picture in place on
|
| 68 |
+
its own.
|
| 69 |
+
|
| 70 |
+
## What it fixes (12 held-out pictures, ComfyUI renders)
|
| 71 |
+
|
| 72 |
+
| | base Qwen 2.1 | + LoRA 500 | + LoRA 1250 | + LoRA 1500 |
|
| 73 |
+
|---|---|---|---|---|
|
| 74 |
+
| gray left unfilled | 9.1 % | 1.1 % | 1.0 % | 0.9 % |
|
| 75 |
+
| kept-area PSNR (picture stays in place) | 23.6 dB | 33.6 dB | 33.6 dB | 33.5 dB |
|
| 76 |
+
| mean colour step at the seam | +2.2 | -0.8 | -0.7 | -0.6 |
|
| 77 |
+
|
| 78 |
+
Without the LoRA, Qwen 2.1 often reframes or rescales the picture (a fisheye
|
| 79 |
+
kitchen shrank inside its frame, a street scene was recomposed) or returns a
|
| 80 |
+
gray frame untouched. With it, the known picture stays pixel-registered, which
|
| 81 |
+
is what makes a clean stitch possible. (The "gray" metric counts any flat
|
| 82 |
+
mid-gray, so gray pavement or sky can read as a few percent.)
|
| 83 |
+
|
| 84 |
+
## Training
|
| 85 |
+
|
| 86 |
+
- [ostris/ai-toolkit](https://github.com/ostris/ai-toolkit), `arch: qwen_image_2`,
|
| 87 |
+
Comfy-Org INT8 convrot base (the same weights ComfyUI runs), text encoder
|
| 88 |
+
embeddings cached with the reference image.
|
| 89 |
+
- 924 pairs from 231 pictures, 4 layouts each: all-sides frame, one side,
|
| 90 |
+
opposite sides, corner, three sides, small window (12-30 % kept). Target =
|
| 91 |
+
the picture at <= 1 MP on the /32 grid (never upscaled), source = same canvas
|
| 92 |
+
with everything outside the kept rectangle painted `#808080`.
|
| 93 |
+
- Captions: the instruction (70 %) or one of five paraphrases (30 %), and for
|
| 94 |
+
75 % of pairs `Scene:` + a Qwen3-VL-8B description of the whole picture.
|
| 95 |
+
Caption dropout 0 (in this trainer a dropped caption also drops the reference).
|
| 96 |
+
- Rank 32 / alpha 32, AdamW8bit, LR 1e-4 constant, batch 1, `shift` timesteps,
|
| 97 |
+
1500 steps. Converged by about step 1250; step 500 already works.
|
| 98 |
+
- Pictures: a hand-picked set plus openly licensed images: Flickr photos from
|
| 99 |
+
[CommonCatalog CC-BY](https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by),
|
| 100 |
+
museum art from [PD12M](https://huggingface.co/datasets/Spawning/PD12M)
|
| 101 |
+
(CC0 / public domain), anime from
|
| 102 |
+
[anime-with-caption-cc0](https://huggingface.co/datasets/alfredplpl/anime-with-caption-cc0).
|
| 103 |
+
|
| 104 |
+
## Limits
|
| 105 |
+
|
| 106 |
+
- Very large extensions (the kept picture under ~15 % of the canvas) invent a
|
| 107 |
+
lot; results vary more by seed.
|
| 108 |
+
- Trained at about 1 MP. Larger canvases work but are untested.
|
| 109 |
+
- Text in the new area is plausible, not legible.
|
| 110 |
+
|
| 111 |
+
## License
|
| 112 |
+
|
| 113 |
+
A LoRA for Qwen Image 2.1, which is released under the Qwen Research License;
|
| 114 |
+
use of the base model, and of this LoRA with it, follows that license.
|