ausboss commited on
Commit
b96e7bb
·
verified ·
1 Parent(s): af57533

Upload outpaint LoRA

Browse files
Files changed (1) hide show
  1. README.md +114 -0
README.md ADDED
@@ -0,0 +1,114 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: qwen-research-license
4
+ base_model: Qwen/Qwen-Image-2.1
5
+ pipeline_tag: image-to-image
6
+ tags:
7
+ - lora
8
+ - outpainting
9
+ - uncrop
10
+ - qwen-image
11
+ - qwen-image-2.1
12
+ - comfyui
13
+ - ai-toolkit
14
+ ---
15
+
16
+ # Qwen Image 2.1 Outpaint LoRA
17
+
18
+ Extends a picture in **any direction**: one side, two sides, a corner, three
19
+ sides, all four, or a small picture on a big empty canvas. Pad the picture with
20
+ flat gray `#808080`, give the padded canvas to Qwen Image 2.1 as the reference,
21
+ and the LoRA fills the gray with a continuation of the scene while keeping
22
+ the original picture where it is.
23
+
24
+ ![source | base | step 500 | step 1250 | step 1500 | ground truth](comparison.jpg)
25
+
26
+ *Held-out pictures (never trained on), rendered in ComfyUI with the INT8 Qwen
27
+ Image 2.1 model, 25 steps, CFG 1: source | base model | LoRA step 500 | 1250 |
28
+ 1500 | original picture.*
29
+
30
+ ## Files
31
+
32
+ | file | notes |
33
+ |---|---|
34
+ | `qwen-image-2.1-outpaint.safetensors` | **use this**: step 1500 |
35
+ | `checkpoints/qwen21_outpaint_v1_000001250.safetensors` | step 1250, practically identical |
36
+ | `checkpoints/qwen21_outpaint_v1_000000500.safetensors` | step 500, already fixes framing; plainer fills |
37
+
38
+ Rank 32, ComfyUI key format (`diffusion_model.transformer_blocks.*`), all 384
39
+ tensors load onto the Comfy-Org Qwen Image 2.1 weights.
40
+
41
+ ## Prompt
42
+
43
+ The instruction is the trigger. Put it first:
44
+
45
+ ```
46
+ Outpaint the image: replace the solid gray areas with a seamless continuation of the scene, keeping the existing picture unchanged.
47
+ ```
48
+
49
+ Optionally follow it with `Scene: <description>`. The training captions
50
+ described the whole picture, so a description of the source (what is visible)
51
+ works well; in ComfyUI, **Text Generate** with the Qwen3-VL 8B encoder you
52
+ already load for Qwen 2.1 can write it automatically from the picture.
53
+
54
+ ## ComfyUI
55
+
56
+ 1. Pad the picture with flat gray `#808080` (AusBoss **Load Image + Pad**:
57
+ fill `color`, `#808080`, canvas multiple 32, target 1.0 MP).
58
+ 2. **Text Encode Qwen Image 2.1**: the padded canvas as `image_1`,
59
+ `resolution` 0 (reference and output share the canvas size), plus the prompt.
60
+ 3. **KSampler** on the encoder's `latent` output: 25 steps, CFG 1,
61
+ `euler` / `simple`, denoise 1.
62
+ 4. **VAE Decode** -> **Split Image with Alpha** (the Qwen 2.1 VAE decodes RGBA)
63
+ -> AusBoss **Stitch Inpaint** to paste the exact original pixels back over a
64
+ 32 px feather.
65
+
66
+ Do **not** pin the known area with *Set Latent Noise Mask*: on Qwen 2.1 that
67
+ draws a visible rectangle at the seam. The LoRA keeps the picture in place on
68
+ its own.
69
+
70
+ ## What it fixes (12 held-out pictures, ComfyUI renders)
71
+
72
+ | | base Qwen 2.1 | + LoRA 500 | + LoRA 1250 | + LoRA 1500 |
73
+ |---|---|---|---|---|
74
+ | gray left unfilled | 9.1 % | 1.1 % | 1.0 % | 0.9 % |
75
+ | kept-area PSNR (picture stays in place) | 23.6 dB | 33.6 dB | 33.6 dB | 33.5 dB |
76
+ | mean colour step at the seam | +2.2 | -0.8 | -0.7 | -0.6 |
77
+
78
+ Without the LoRA, Qwen 2.1 often reframes or rescales the picture (a fisheye
79
+ kitchen shrank inside its frame, a street scene was recomposed) or returns a
80
+ gray frame untouched. With it, the known picture stays pixel-registered, which
81
+ is what makes a clean stitch possible. (The "gray" metric counts any flat
82
+ mid-gray, so gray pavement or sky can read as a few percent.)
83
+
84
+ ## Training
85
+
86
+ - [ostris/ai-toolkit](https://github.com/ostris/ai-toolkit), `arch: qwen_image_2`,
87
+ Comfy-Org INT8 convrot base (the same weights ComfyUI runs), text encoder
88
+ embeddings cached with the reference image.
89
+ - 924 pairs from 231 pictures, 4 layouts each: all-sides frame, one side,
90
+ opposite sides, corner, three sides, small window (12-30 % kept). Target =
91
+ the picture at <= 1 MP on the /32 grid (never upscaled), source = same canvas
92
+ with everything outside the kept rectangle painted `#808080`.
93
+ - Captions: the instruction (70 %) or one of five paraphrases (30 %), and for
94
+ 75 % of pairs `Scene:` + a Qwen3-VL-8B description of the whole picture.
95
+ Caption dropout 0 (in this trainer a dropped caption also drops the reference).
96
+ - Rank 32 / alpha 32, AdamW8bit, LR 1e-4 constant, batch 1, `shift` timesteps,
97
+ 1500 steps. Converged by about step 1250; step 500 already works.
98
+ - Pictures: a hand-picked set plus openly licensed images: Flickr photos from
99
+ [CommonCatalog CC-BY](https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by),
100
+ museum art from [PD12M](https://huggingface.co/datasets/Spawning/PD12M)
101
+ (CC0 / public domain), anime from
102
+ [anime-with-caption-cc0](https://huggingface.co/datasets/alfredplpl/anime-with-caption-cc0).
103
+
104
+ ## Limits
105
+
106
+ - Very large extensions (the kept picture under ~15 % of the canvas) invent a
107
+ lot; results vary more by seed.
108
+ - Trained at about 1 MP. Larger canvases work but are untested.
109
+ - Text in the new area is plausible, not legible.
110
+
111
+ ## License
112
+
113
+ A LoRA for Qwen Image 2.1, which is released under the Qwen Research License;
114
+ use of the base model, and of this LoRA with it, follows that license.