recoilme commited on
Commit
0e81ff1
Β·
1 Parent(s): 88a93c1

Adapter v12 in the DiT: 2304-slot position table, trained at the real edit geometry

Browse files

The bundled text_fusion is now revision v12 (vision cosine against the native Qwen3-VL-8B encoder
0.930 -> 0.971, text unchanged at 0.951/24.2%). Two things moved: the attention-branch position table
covers 2304 slots, and the fine-tune ran at the real inference geometry (--geom 1024), so the ~2000
token condition of a two-image edit no longer falls into a zero-padded tail.

transformer/config.json text_fusion_config.max_len goes 512 -> 2304 to match the table; the model card,
the conditioning row and the card media are re-rendered with the shipped weights.

README.md CHANGED
@@ -28,7 +28,7 @@ Text-to-image, character and scene editing, and transparent (RGBA) generation in
28
  |---|---|
29
  | transformer | Qwen-Image-2.1 DiT β€” 32 layers, 14.5 GB fp16, plus a **158M text-fusion adapter** inside |
30
  | text encoder | **Qwen3.5-0.8B**, 1.7 GB fp16 β€” upstream checkpoint re-saved to fp16, tokenizer/processor files unchanged (native: Qwen3-VL-8B, 17.5 GB) |
31
- | conditioning | cosine **0.94** against the native Qwen3-VL-8B encoder (text positions) |
32
  | VAE | Qwen-Image-2.1, 16Γ— spatial, fp32 |
33
  | scheduler | `FlowMatchEulerDiscreteScheduler`, plain static shift 5.0 (dynamic shifting off) |
34
  | resolution | `output_resolution`, 1024 by default; follows the condition image aspect ratio |
@@ -43,6 +43,11 @@ images (**Improved using Qwen**). The adapter lives *inside* the DiT as its text
43
  whole model is one self-contained diffusers folder and no 17.5 GB encoder is needed anywhere. The
44
  sampler runs a plain static shift of 5.0 instead of the original dynamic shifting.
45
 
 
 
 
 
 
46
  ### Examples
47
 
48
  Every image below is generated by this pipeline with 30 steps at 1024 px.
@@ -64,7 +69,8 @@ pose, clothing and scene; the identity is copied from `<image2>`)
64
  |---|---|
65
  | ![edit swap](media/edit_swap.jpg) | ![edit char](media/edit_char.jpg) |
66
 
67
- **Edit β€” three condition images** (subject from `<image1>`, scene from `<image2>`, lighting from `<image3>`)
 
68
 
69
  ![edit three](media/edit_three.jpg)
70
 
@@ -170,6 +176,8 @@ would not find the DiT class.
170
  * **English only** β€” that is all the adapter was trained and tested on; other languages drift.
171
  * **Numerals on signage** come out wrong: "OPEN 24 HOURS" renders as "OPEN **26** HOURS" on every
172
  seed tried. Words are fine. ![numbers](media/limit_numbers.jpg)
 
 
173
  * Batch size >1 at 1024 px peaks near 28 GB; one prompt per call is the safe mode.
174
 
175
  ### NOTICE
 
28
  |---|---|
29
  | transformer | Qwen-Image-2.1 DiT β€” 32 layers, 14.5 GB fp16, plus a **158M text-fusion adapter** inside |
30
  | text encoder | **Qwen3.5-0.8B**, 1.7 GB fp16 β€” upstream checkpoint re-saved to fp16, tokenizer/processor files unchanged (native: Qwen3-VL-8B, 17.5 GB) |
31
+ | conditioning | cosine **0.95** on text, **0.97** on the vision positions of edit prompts, against the native Qwen3-VL-8B encoder |
32
  | VAE | Qwen-Image-2.1, 16Γ— spatial, fp32 |
33
  | scheduler | `FlowMatchEulerDiscreteScheduler`, plain static shift 5.0 (dynamic shifting off) |
34
  | resolution | `output_resolution`, 1024 by default; follows the condition image aspect ratio |
 
43
  whole model is one self-contained diffusers folder and no 17.5 GB encoder is needed anywhere. The
44
  sampler runs a plain static shift of 5.0 instead of the original dynamic shifting.
45
 
46
+ The bundled adapter is revision **v12**: its attention-branch position table covers 2304 slots and it
47
+ was fine-tuned at the real inference geometry (~2000-token conditions at 1024 px), so reference images
48
+ keep their positions instead of falling into a zero-padded tail β€” the vision cosine against the native
49
+ encoder moved 0.93 β†’ 0.97, text stayed at 0.95.
50
+
51
  ### Examples
52
 
53
  Every image below is generated by this pipeline with 30 steps at 1024 px.
 
69
  |---|---|
70
  | ![edit swap](media/edit_swap.jpg) | ![edit char](media/edit_char.jpg) |
71
 
72
+ **Edit β€” three condition images** (target and composition from `<image1>`, the person from `<image2>`,
73
+ colour and lighting from `<image3>`)
74
 
75
  ![edit three](media/edit_three.jpg)
76
 
 
176
  * **English only** β€” that is all the adapter was trained and tested on; other languages drift.
177
  * **Numerals on signage** come out wrong: "OPEN 24 HOURS" renders as "OPEN **26** HOURS" on every
178
  seed tried. Words are fine. ![numbers](media/limit_numbers.jpg)
179
+ * **Non-photo references transfer less faithfully** than photographic ones: the adapter imitates the
180
+ native encoder, so its ceiling is the native encoder's ceiling.
181
  * Batch size >1 at 1024 px peaks near 28 GB; one prompt per call is the safe mode.
182
 
183
  ### NOTICE
media/edit_char.jpg CHANGED

Git LFS Details

  • SHA256: 720fe2624336626620c4da29614ed0cd55a56574c3c29b708aa4934f7232ef13
  • Pointer size: 131 Bytes
  • Size of remote file: 123 kB

Git LFS Details

  • SHA256: 975128b005fa3777eebb49a05fab2cb13432cdb542414ce71ef4df51143cba31
  • Pointer size: 131 Bytes
  • Size of remote file: 123 kB
media/edit_single.jpg CHANGED

Git LFS Details

  • SHA256: 5b80383b19098d44902e5ee46a6a81aae290c2963eee57e516865f55abfc90bc
  • Pointer size: 131 Bytes
  • Size of remote file: 232 kB

Git LFS Details

  • SHA256: faad03123d37fe40efab9948957fb41da76629d1b58f4d48273805f7890580c2
  • Pointer size: 131 Bytes
  • Size of remote file: 215 kB
media/edit_swap.jpg CHANGED

Git LFS Details

  • SHA256: f09e6690c9ac7953b84795956aebd8b7ddad72c0b87f1f4053f0366899f34dd5
  • Pointer size: 132 Bytes
  • Size of remote file: 1.63 MB

Git LFS Details

  • SHA256: b44be5d3a1a545a44441812879e89964b45c077f96673cd45ec4f0d3e8411c6c
  • Pointer size: 131 Bytes
  • Size of remote file: 245 kB
media/edit_three.jpg CHANGED

Git LFS Details

  • SHA256: 620b943ef69f2003bbaf34d89ba14d789edd66a8ef12684461f95814eb7a9c25
  • Pointer size: 131 Bytes
  • Size of remote file: 234 kB

Git LFS Details

  • SHA256: 878fc3bef5b8563f9e67a5f35de56abaf3f2251e169768db711732939b42b12b
  • Pointer size: 131 Bytes
  • Size of remote file: 252 kB
media/limit_numbers.jpg CHANGED

Git LFS Details

  • SHA256: 130073a15d88dda8b064cbe2fcc55c851e3b9398853228a8833024a8d7068e69
  • Pointer size: 131 Bytes
  • Size of remote file: 406 kB

Git LFS Details

  • SHA256: db9fb272fbea7873292e1f4fdbdcaea958fa1b65c873987f9dcb7951099bda31
  • Pointer size: 131 Bytes
  • Size of remote file: 354 kB
media/t2i.jpg CHANGED

Git LFS Details

  • SHA256: 84bca850d238fb742290e693cf912ac62b8f4950073decd99c1624012013a64a
  • Pointer size: 131 Bytes
  • Size of remote file: 154 kB

Git LFS Details

  • SHA256: feca10a657ada0f4e4289912f487a5d953e8cc543c40bbf074983147600716f2
  • Pointer size: 131 Bytes
  • Size of remote file: 150 kB
media/transparent.png CHANGED

Git LFS Details

  • SHA256: c7affd1e6125b1b621982e9b1656cf5774399fb678cd9ad61c5c934f40bdc62c
  • Pointer size: 131 Bytes
  • Size of remote file: 855 kB

Git LFS Details

  • SHA256: 18534b03b3ca368e9c3619e7074ab947aac4d8c9c3ecb47139aae1d0d13a7ce0
  • Pointer size: 131 Bytes
  • Size of remote file: 773 kB
transformer/config.json CHANGED
@@ -26,7 +26,7 @@
26
  "attention": 2,
27
  "attn_dim": 1024,
28
  "attn_heads": 8,
29
- "max_len": 512,
30
  "mixer": 2,
31
  "mixer_heads": 8,
32
  "mixer_ffn": 2,
 
26
  "attention": 2,
27
  "attn_dim": 1024,
28
  "attn_heads": 8,
29
+ "max_len": 2304,
30
  "mixer": 2,
31
  "mixer_heads": 8,
32
  "mixer_ffn": 2,
transformer/diffusion_pytorch_model-00001-of-00002.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a71f3c4eca68c67df8c8fd288cc1daf27ac0a89d6eee9f0ab7aa2ec1031adbe4
3
- size 10284159452
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3e706b011865bbcf5e36c9a62e189549179de9fcbf8c886afc110ce2b30992bc
3
+ size 10287829484