ausboss's picture
Upload outpaint LoRA
b96e7bb verified
|
Raw History Blame
4.94 kB
metadata
license: other
license_name: qwen-research-license
base_model: Qwen/Qwen-Image-2.1
pipeline_tag: image-to-image
tags:
  - lora
  - outpainting
  - uncrop
  - qwen-image
  - qwen-image-2.1
  - comfyui
  - ai-toolkit

Qwen Image 2.1 Outpaint LoRA

Extends a picture in any direction: one side, two sides, a corner, three sides, all four, or a small picture on a big empty canvas. Pad the picture with flat gray #808080, give the padded canvas to Qwen Image 2.1 as the reference, and the LoRA fills the gray with a continuation of the scene while keeping the original picture where it is.

source | base | step 500 | step 1250 | step 1500 | ground truth

Held-out pictures (never trained on), rendered in ComfyUI with the INT8 Qwen Image 2.1 model, 25 steps, CFG 1: source | base model | LoRA step 500 | 1250 | 1500 | original picture.

Files

file notes
qwen-image-2.1-outpaint.safetensors use this: step 1500
checkpoints/qwen21_outpaint_v1_000001250.safetensors step 1250, practically identical
checkpoints/qwen21_outpaint_v1_000000500.safetensors step 500, already fixes framing; plainer fills

Rank 32, ComfyUI key format (diffusion_model.transformer_blocks.*), all 384 tensors load onto the Comfy-Org Qwen Image 2.1 weights.

Prompt

The instruction is the trigger. Put it first:

Outpaint the image: replace the solid gray areas with a seamless continuation of the scene, keeping the existing picture unchanged.

Optionally follow it with Scene: <description>. The training captions described the whole picture, so a description of the source (what is visible) works well; in ComfyUI, Text Generate with the Qwen3-VL 8B encoder you already load for Qwen 2.1 can write it automatically from the picture.

ComfyUI

  1. Pad the picture with flat gray #808080 (AusBoss Load Image + Pad: fill color, #808080, canvas multiple 32, target 1.0 MP).
  2. Text Encode Qwen Image 2.1: the padded canvas as image_1, resolution 0 (reference and output share the canvas size), plus the prompt.
  3. KSampler on the encoder's latent output: 25 steps, CFG 1, euler / simple, denoise 1.
  4. VAE Decode -> Split Image with Alpha (the Qwen 2.1 VAE decodes RGBA) -> AusBoss Stitch Inpaint to paste the exact original pixels back over a 32 px feather.

Do not pin the known area with Set Latent Noise Mask: on Qwen 2.1 that draws a visible rectangle at the seam. The LoRA keeps the picture in place on its own.

What it fixes (12 held-out pictures, ComfyUI renders)

base Qwen 2.1 + LoRA 500 + LoRA 1250 + LoRA 1500
gray left unfilled 9.1 % 1.1 % 1.0 % 0.9 %
kept-area PSNR (picture stays in place) 23.6 dB 33.6 dB 33.6 dB 33.5 dB
mean colour step at the seam +2.2 -0.8 -0.7 -0.6

Without the LoRA, Qwen 2.1 often reframes or rescales the picture (a fisheye kitchen shrank inside its frame, a street scene was recomposed) or returns a gray frame untouched. With it, the known picture stays pixel-registered, which is what makes a clean stitch possible. (The "gray" metric counts any flat mid-gray, so gray pavement or sky can read as a few percent.)

Training

  • ostris/ai-toolkit, arch: qwen_image_2, Comfy-Org INT8 convrot base (the same weights ComfyUI runs), text encoder embeddings cached with the reference image.
  • 924 pairs from 231 pictures, 4 layouts each: all-sides frame, one side, opposite sides, corner, three sides, small window (12-30 % kept). Target = the picture at <= 1 MP on the /32 grid (never upscaled), source = same canvas with everything outside the kept rectangle painted #808080.
  • Captions: the instruction (70 %) or one of five paraphrases (30 %), and for 75 % of pairs Scene: + a Qwen3-VL-8B description of the whole picture. Caption dropout 0 (in this trainer a dropped caption also drops the reference).
  • Rank 32 / alpha 32, AdamW8bit, LR 1e-4 constant, batch 1, shift timesteps, 1500 steps. Converged by about step 1250; step 500 already works.
  • Pictures: a hand-picked set plus openly licensed images: Flickr photos from CommonCatalog CC-BY, museum art from PD12M (CC0 / public domain), anime from anime-with-caption-cc0.

Limits

  • Very large extensions (the kept picture under ~15 % of the canvas) invent a lot; results vary more by seed.
  • Trained at about 1 MP. Larger canvases work but are untested.
  • Text in the new area is plausible, not legible.

License

A LoRA for Qwen Image 2.1, which is released under the Qwen Research License; use of the base model, and of this LoRA with it, follows that license.