recoilme commited on
Commit
88a93c1
·
1 Parent(s): d264d52

Fix the editing convention: the FIRST image is the target, not the last

Browse files

The tag/build order in the usage example was inverted, which made swaps silently fail: the model
edits <image1>, so passing the reference first asks it to edit the reference. Verified on one pair
both ways (identity transfers only with the target first); the two edit examples in media/ are
regenerated with the correct order.

Files changed (4) hide show
  1. README.md +12 -4
  2. example.py +4 -3
  3. media/edit_char.jpg +2 -2
  4. media/edit_swap.jpg +2 -2
README.md CHANGED
@@ -57,7 +57,8 @@ Every image below is generated by this pipeline with 30 steps at 1024 px.
57
 
58
  ![edit single](media/edit_single.jpg)
59
 
60
- **Edit — two condition images** (character replacement: identity from `<image1>`, pose/clothing/scene from `<image2>`)
 
61
 
62
  | | |
63
  |---|---|
@@ -86,14 +87,21 @@ image = pipe(prompt="a red fox in a snowy forest at dusk, cinematic, 85mm",
86
  output_resolution=1024, num_inference_steps=30,
87
  generator=torch.Generator("cuda").manual_seed(1234)).images[0]
88
 
89
- # editing: 1..N condition images, referenced in the prompt by TAG <image1>, <image2>, ...
90
- image = pipe(prompt="Replace the woman in <image2> with the woman from <image1>; keep <image2> pose, "
 
91
  "clothing and background unchanged.",
92
- image=[ref_image, scene_image],
93
  output_resolution=1024, num_inference_steps=30,
94
  generator=torch.Generator("cuda").manual_seed(1234)).images[0]
95
  ```
96
 
 
 
 
 
 
 
97
  `custom_pipeline="pipeline"` builds the shipped `pipeline.py` and `trust_remote_code=True` lets it run,
98
  so no clone is needed. (`_class_name` is kept a plain string in `model_index.json` because that is what
99
  Hub tooling expects; the `[file, class]` form diffusers also accepts makes the Hub print a
 
57
 
58
  ![edit single](media/edit_single.jpg)
59
 
60
+ **Edit — two condition images** (character replacement: `<image1>` is the edit target and keeps its
61
+ pose, clothing and scene; the identity is copied from `<image2>`)
62
 
63
  | | |
64
  |---|---|
 
87
  output_resolution=1024, num_inference_steps=30,
88
  generator=torch.Generator("cuda").manual_seed(1234)).images[0]
89
 
90
+ # editing: 1..N condition images. The FIRST one is the edit target, the rest are references;
91
+ # reference them in the prompt by TAG <image1>, <image2>, ...
92
+ image = pipe(prompt="Replace the woman in <image1> with the woman from <image2>; keep <image1> pose, "
93
  "clothing and background unchanged.",
94
+ image=[scene_image, ref_image],
95
  output_resolution=1024, num_inference_steps=30,
96
  generator=torch.Generator("cuda").manual_seed(1234)).images[0]
97
  ```
98
 
99
+ Editing convention: the **first** image is the one being edited (`<image1>`), everything after it is a
100
+ reference. That is the model's own convention and what the stock ComfyUI node documents; feeding the
101
+ reference first is the usual reason a swap "does not happen" (the model then edits the reference).
102
+ Note that the *canvas size* still comes from the last image's aspect ratio — pass `height`/`width`
103
+ explicitly to pin it.
104
+
105
  `custom_pipeline="pipeline"` builds the shipped `pipeline.py` and `trust_remote_code=True` lets it run,
106
  so no clone is needed. (`_class_name` is kept a plain string in `model_index.json` because that is what
107
  Hub tooling expects; the `[file, class]` form diffusers also accepts makes the Hub print a
example.py CHANGED
@@ -7,9 +7,10 @@ On the GPU: Qwen3.5-0.8B (~1.7 GB), the DiT with the adapter inside (~14.5 GB) a
7
  # text-to-image
8
  python example.py --prompt "a red fox in a snowy forest at dusk, cinematic, 85mm" --out fox.png
9
 
10
- # editing: 1..N condition images, referenced in the prompt by TAG <image1>, <image2>, ...
11
- python example.py --image ref.png scene.png \
12
- --prompt "Replace the woman in <image2> with the woman from <image1>; keep <image2> pose, \\
 
13
  clothing and background unchanged." --out swap.png
14
 
15
  # a batch from a text file: one prompt per line, '#' starts a comment, blank lines are skipped
 
7
  # text-to-image
8
  python example.py --prompt "a red fox in a snowy forest at dusk, cinematic, 85mm" --out fox.png
9
 
10
+ # editing: 1..N condition images. The FIRST one is the edit target, the rest are references;
11
+ # the prompt refers to them as <image1>, <image2>, ...
12
+ python example.py --image scene.png ref.png \
13
+ --prompt "Replace the woman in <image1> with the woman from <image2>; keep <image1> pose, \\
14
  clothing and background unchanged." --out swap.png
15
 
16
  # a batch from a text file: one prompt per line, '#' starts a comment, blank lines are skipped
media/edit_char.jpg CHANGED

Git LFS Details

  • SHA256: 3df5faa4989ab0351c5f3614b3d5c2cc9a557a3888abb2d3da62b2073b7537e5
  • Pointer size: 131 Bytes
  • Size of remote file: 136 kB

Git LFS Details

  • SHA256: 720fe2624336626620c4da29614ed0cd55a56574c3c29b708aa4934f7232ef13
  • Pointer size: 131 Bytes
  • Size of remote file: 123 kB
media/edit_swap.jpg CHANGED

Git LFS Details

  • SHA256: 69f98269f59a92042bd0d29776669bf214a4c1bc7435fa3f7e5cfc4135518d76
  • Pointer size: 131 Bytes
  • Size of remote file: 242 kB

Git LFS Details

  • SHA256: f09e6690c9ac7953b84795956aebd8b7ddad72c0b87f1f4053f0366899f34dd5
  • Pointer size: 132 Bytes
  • Size of remote file: 1.63 MB