--- license: other license_name: qwen-research-license base_model: Qwen/Qwen-Image-2.1 pipeline_tag: image-to-image tags: - lora - outpainting - uncrop - qwen-image - qwen-image-2.1 - comfyui - ai-toolkit --- # Qwen Image 2.1 Outpaint LoRA Extends a picture in **any direction**: one side, two sides, a corner, three sides, all four, or a small picture on a big empty canvas. Pad the picture with flat gray `#808080`, give the padded canvas to Qwen Image 2.1 as the reference, and the LoRA fills the gray with a continuation of the scene while keeping the original picture where it is. **Ready-made ComfyUI workflow:** [Qwen Image 2.1 Outpaint: auto prompt, any side, optional LoRA](https://civitai.com/models/2964810/qwen-image-21-outpaint-auto-prompt-any-side-optional-lora) on Civitai. ![padded input | no LoRA | v1 | v2](comparison.jpg) *Pictures neither LoRA was trained on, rendered in ComfyUI with the INT8 Qwen Image 2.1 model: 25 steps, CFG 1, seed 42, the same auto-caption prompt in every column, original pixels pasted back with Stitch Inpaint. Each row: padded input | no LoRA | v1 | v2. The amber box marks what goes wrong without the LoRA: visible seams where the original is pasted back, and a duplicated person.* ## Files | file | notes | |---|---| | `qwen-image-2.1-outpaint.safetensors` | **v1**, step 1500: any extension size, trained at ~1 MP | | `qwen-image-2.1-outpaint-v2.safetensors` | **v2**, step 2000: everyday extensions at 1-2 MP, people and outfits | | `checkpoints/qwen21_outpaint_v1_000001250.safetensors` | step 1250, practically identical | | `checkpoints/qwen21_outpaint_v1_000000500.safetensors` | step 500, already fixes framing; plainer fills | Both are rank 32, ComfyUI key format (`diffusion_model.transformer_blocks.*`), all 384 tensors load onto the Comfy-Org Qwen Image 2.1 weights. ## v1 or v2 Both fix the same thing (the picture stays in place, so the stitch is clean) and measure almost the same. v2 is a little better on small and medium extensions and was trained at 1-2 MP on a people- and fashion-heavy set; v1 saw more extreme zoom-outs. | unseen pictures | no LoRA | v1 | v2 | |---|---|---|---| | subtle crops (10 pictures, ~75 % kept): picture stays in place | 16.4 dB | 34.6 dB | 34.7 dB | | subtle crops: fill error vs the real photo (lower is better) | 16.3 | 10.5 | 10.2 | | big crops (8 pictures, 36-50 % kept): picture stays in place | 25.5 dB | 33.9 dB | 34.0 dB | | big crops: fill error | 29.8 | 27.0 | 27.2 | ## Prompt The instruction is the trigger. Put it first: ``` Outpaint the image: replace the solid gray areas with a seamless continuation of the scene, keeping the existing picture unchanged. ``` Optionally follow it with `Scene: `. The training captions described the whole picture, so a description of the source (what is visible) works well; in ComfyUI, **Text Generate** with the Qwen3-VL 8B encoder you already load for Qwen 2.1 can write it automatically from the picture. ## ComfyUI The quickest start is the [ready-made workflow on Civitai](https://civitai.com/models/2964810/qwen-image-21-outpaint-auto-prompt-any-side-optional-lora). To build it yourself: 1. Pad the picture with flat gray `#808080` (AusBoss **Load Image + Pad**: fill `color`, `#808080`, canvas multiple 32, target 1.0 MP for v1, 1-2 MP for v2). For a tilted picture, use **Image Crop + Rotate + Pad** with `feather` 0 instead (see [Tilted pictures](#tilted-pictures)). 2. **Text Encode Qwen Image 2.1**: the padded canvas as `image_1`, `resolution` 0 (reference and output share the canvas size), plus the prompt. 3. **KSampler** on the encoder's `latent` output: 25 steps, CFG 1, `euler` / `simple`, denoise 1. 4. **VAE Decode** -> **Split Image with Alpha** (the Qwen 2.1 VAE decodes RGBA) -> AusBoss **Stitch Inpaint** to paste the exact original pixels back over a 32 px feather. Do **not** pin the known area with *Set Latent Noise Mask*: on Qwen 2.1 that draws a visible rectangle at the seam. The LoRA keeps the picture in place on its own. The Civitai workflow's LoRA Loader skips a missing file instead of stopping, so without the download it quietly runs plain Qwen 2.1. If your results show a lighter box or doubled edges where the original was pasted back, check that the LoRA row is loaded. ## What v1 fixes (12 held-out pictures, ComfyUI renders) | | base Qwen 2.1 | + LoRA 500 | + LoRA 1250 | + LoRA 1500 | |---|---|---|---|---| | gray left unfilled | 9.1 % | 1.1 % | 1.0 % | 0.9 % | | kept-area PSNR (picture stays in place) | 23.6 dB | 33.6 dB | 33.6 dB | 33.5 dB | | mean colour step at the seam | +2.2 | -0.8 | -0.7 | -0.6 | Without the LoRA, Qwen 2.1 often reframes or rescales the picture (a fisheye kitchen shrank inside its frame, a street scene was recomposed) or returns a gray frame untouched. With it, the known picture stays pixel-registered, which is what makes a clean stitch possible. (The "gray" metric counts any flat mid-gray, so gray pavement or sky can read as a few percent.) ## Tilted pictures Straightening or tilting a picture before outpainting leaves gray wedges in the corners instead of straight borders. Neither LoRA was trained on diagonal borders, and both handle them. ![tilted input | no LoRA | v1 | v2](rotation.jpg) *36 ComfyUI renders: a sports photo, a flat lawn and the pier above, each rotated 17° with AusBoss Image Crop + Rotate + Pad, 2 seeds, `feather` 24 and 0, the same prompt and seed in every column, original pasted back with Stitch Inpaint.* | rotated 17° | no LoRA | v1 | v2 | |---|---|---|---| | renders where the picture did not stay in place near the tilted edges (> 2 px) | 12 / 12 | 0 / 12 | 0 / 12 | | largest shift (sports photo, pier) | 95 px | 0 px | 0 px | Without the LoRA, Qwen 2.1 moves the scene while it fills the wedges, so the pasted-back original no longer lines up: a lighter tilted box, cut or doubled edges and ghosted objects along the tilt. With either LoRA the picture stays exactly in place. Set `feather` to **0** on Image Crop + Rotate + Pad (AusBoss nodes 2.2.0). Its feather also fades the picture itself into the gray fill (Load Image + Pad only feathers the mask). The model paints that ramp as a darker band along the tilt, and Stitch Inpaint's color match reads the faded edge and pulls flat areas grayer. Stitch Inpaint still blends the paste-back over 32 px at feather 0. ## Training ### v1 - [ostris/ai-toolkit](https://github.com/ostris/ai-toolkit), `arch: qwen_image_2`, Comfy-Org INT8 convrot base (the same weights ComfyUI runs), text encoder embeddings cached with the reference image. - 924 pairs from 231 pictures, 4 layouts each: all-sides frame, one side, opposite sides, corner, three sides, small window (12-30 % kept). Target = the picture at <= 1 MP on the /32 grid (never upscaled), source = same canvas with everything outside the kept rectangle painted `#808080`. - Captions: the instruction (70 %) or one of five paraphrases (30 %), and for 75 % of pairs `Scene:` + a Qwen3-VL-8B description of the whole picture. Caption dropout 0 (in this trainer a dropped caption also drops the reference). - Rank 32 / alpha 32, AdamW8bit, LR 1e-4 constant, batch 1, `shift` timesteps, 1500 steps. Converged by about step 1250; step 500 already works. - Pictures: a hand-picked set plus openly licensed images: Flickr photos from [CommonCatalog CC-BY](https://huggingface.co/datasets/common-canvas/commoncatalog-cc-by), museum art from [PD12M](https://huggingface.co/datasets/Spawning/PD12M) (CC0 / public domain), anime from [anime-with-caption-cc0](https://huggingface.co/datasets/alfredplpl/anime-with-caption-cc0). ### v2 - 68 hand-picked pictures: 47 of people and outfits (Unsplash photos under the [Unsplash License](https://unsplash.com/license) via [unsplash-lite](https://huggingface.co/datasets/1aurent/unsplash-lite), plus the author's own), 21 of scenes, interiors, anime and paintings. One random layout per picture, 68 pairs. - Smaller extensions than v1: the kept picture covers 45-93 % of the canvas (median ~74 %), about one layout in seven bolder; no small-window zoom-outs. - Targets at 1-2 MP (never upscaled), ai-toolkit `resolution: 1408`. - Same captions recipe, rank 32 / alpha 32, AdamW8bit, LR 1e-4, 2000 steps. Works from step 500; step 2000 had the lowest fill error. ## Limits - Very large extensions (the kept picture under ~15 % of the canvas) invent a lot; results vary more by seed. - v1 was trained at about 1 MP, v2 at 1-2 MP; much larger canvases are untested. - Text in the new area is plausible, not legible. - On shallow-focus photos, tilt wedges come out a little softer than an in-focus subject: the model continues the blur it sees. Qwen 2.1 without the LoRA, Krea 2 and FLUX.2 Klein did the same in the same test. ## License A LoRA for Qwen Image 2.1, which is released under the Qwen Research License; use of the base model, and of this LoRA with it, follows that license.