Image-to-Image
Diffusers
Safetensors
ZenImageEditPipeline
text-to-image
image-editing
qwen-image
text-encoder
adapter
Instructions to use AiArtLab/zen-image-edit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use AiArtLab/zen-image-edit with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("AiArtLab/zen-image-edit", dtype=torch.bfloat16, device_map="cuda") prompt = "Turn this cat into a dog" input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png") image = pipe(image=input_image, prompt=prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Default scheduler: plain static shift 5.0 instead of dynamic shifting (sdxs-micro schedule)
Browse filesVerified bit-identical to the approved A/B frame; --scheduler-test now compares the shipped static
shift against Qwen-Image-2.1's original dynamic schedule. Also fixed a silent config trap: a loaded
scheduler config carries _use_default_values, which makes from_config drop every key set afterwards.
- NOTICE +2 -1
- README.md +5 -4
- example.py +35 -14
- gens/0001.png +3 -0
- scheduler/scheduler_config.json +2 -13
NOTICE
CHANGED
|
@@ -14,7 +14,8 @@ Modified files, as required by section 2.b of the agreement:
|
|
| 14 |
transformer/*.safetensors re-saved in fp16, plus the 66 `text_fusion.*` tensors of the
|
| 15 |
adapter from the image21-08b-text-encoder-adapter project (adapter_v11)
|
| 16 |
vae/ copied unchanged (fp32)
|
| 17 |
-
scheduler/
|
|
|
|
| 18 |
text_encoder/ different model: Qwen3.5-0.8B instead of Qwen3-VL-8B. Upstream files are
|
| 19 |
re-saved in fp16 (bf16 -> fp16, single `model.safetensors` instead of
|
| 20 |
the sharded bf16 file); `config.json` is unchanged. Modified under the
|
|
|
|
| 14 |
transformer/*.safetensors re-saved in fp16, plus the 66 `text_fusion.*` tensors of the
|
| 15 |
adapter from the image21-08b-text-encoder-adapter project (adapter_v11)
|
| 16 |
vae/ copied unchanged (fp32)
|
| 17 |
+
scheduler/ modified: dynamic shifting off, plain static `shift` 5.0, `shift_terminal`
|
| 18 |
+
removed (was 0.02) — the sdxs-micro schedule
|
| 19 |
text_encoder/ different model: Qwen3.5-0.8B instead of Qwen3-VL-8B. Upstream files are
|
| 20 |
re-saved in fp16 (bf16 -> fp16, single `model.safetensors` instead of
|
| 21 |
the sharded bf16 file); `config.json` is unchanged. Modified under the
|
README.md
CHANGED
|
@@ -30,7 +30,7 @@ Text-to-image, character and scene editing, and transparent (RGBA) generation in
|
|
| 30 |
| text encoder | **Qwen3.5-0.8B**, 1.7 GB fp16 — upstream checkpoint re-saved to fp16, tokenizer/processor files unchanged (native: Qwen3-VL-8B, 17.5 GB) |
|
| 31 |
| conditioning | cosine **0.94** against the native Qwen3-VL-8B encoder (text positions) |
|
| 32 |
| VAE | Qwen-Image-2.1, 16× spatial, fp32 |
|
| 33 |
-
| scheduler | `FlowMatchEulerDiscreteScheduler` |
|
| 34 |
| resolution | `output_resolution`, 1024 by default; follows the condition image aspect ratio |
|
| 35 |
| precision | fp16 everywhere except the VAE |
|
| 36 |
| peak VRAM | ~17.5 GB resident, less with `enable_model_cpu_offload()` |
|
|
@@ -40,7 +40,8 @@ Text-to-image, character and scene editing, and transparent (RGBA) generation in
|
|
| 40 |
The text encoder is replaced by **Qwen3.5-0.8B** plus a 158M adapter, fine-tuned to reproduce what
|
| 41 |
the native encoder produced — both from plain text and from text read together with the reference
|
| 42 |
images (**Improved using Qwen**). The adapter lives *inside* the DiT as its text-fusion block, so the
|
| 43 |
-
whole model is one self-contained diffusers folder and no 17.5 GB encoder is needed anywhere.
|
|
|
|
| 44 |
|
| 45 |
### Examples
|
| 46 |
|
|
@@ -112,8 +113,8 @@ python example.py --prompt "..." --negative "low quality, blurry, watermark" --c
|
|
| 112 |
python example.py --prompt "..." --scheduler-test --shift 5 --out ab.png
|
| 113 |
```
|
| 114 |
|
| 115 |
-
`--scheduler-test` renders every prompt twice with the same seed — the
|
| 116 |
-
and
|
| 117 |
change can be judged without rerunning anything by hand.
|
| 118 |
|
| 119 |
`--size` sets a square frame (or the frame *area* when `--image` supplies the aspect ratio);
|
|
|
|
| 30 |
| text encoder | **Qwen3.5-0.8B**, 1.7 GB fp16 — upstream checkpoint re-saved to fp16, tokenizer/processor files unchanged (native: Qwen3-VL-8B, 17.5 GB) |
|
| 31 |
| conditioning | cosine **0.94** against the native Qwen3-VL-8B encoder (text positions) |
|
| 32 |
| VAE | Qwen-Image-2.1, 16× spatial, fp32 |
|
| 33 |
+
| scheduler | `FlowMatchEulerDiscreteScheduler`, plain static shift 5.0 (dynamic shifting off) |
|
| 34 |
| resolution | `output_resolution`, 1024 by default; follows the condition image aspect ratio |
|
| 35 |
| precision | fp16 everywhere except the VAE |
|
| 36 |
| peak VRAM | ~17.5 GB resident, less with `enable_model_cpu_offload()` |
|
|
|
|
| 40 |
The text encoder is replaced by **Qwen3.5-0.8B** plus a 158M adapter, fine-tuned to reproduce what
|
| 41 |
the native encoder produced — both from plain text and from text read together with the reference
|
| 42 |
images (**Improved using Qwen**). The adapter lives *inside* the DiT as its text-fusion block, so the
|
| 43 |
+
whole model is one self-contained diffusers folder and no 17.5 GB encoder is needed anywhere. The
|
| 44 |
+
sampler runs a plain static shift of 5.0 instead of the original dynamic shifting.
|
| 45 |
|
| 46 |
### Examples
|
| 47 |
|
|
|
|
| 113 |
python example.py --prompt "..." --scheduler-test --shift 5 --out ab.png
|
| 114 |
```
|
| 115 |
|
| 116 |
+
`--scheduler-test` renders every prompt twice with the same seed — the shipped static `--shift` (5.0)
|
| 117 |
+
and Qwen-Image-2.1's original dynamic-shift schedule — and glues the pair with labels, so a schedule
|
| 118 |
change can be judged without rerunning anything by hand.
|
| 119 |
|
| 120 |
`--size` sets a square frame (or the frame *area* when `--image` supplies the aspect ratio);
|
example.py
CHANGED
|
@@ -19,8 +19,8 @@ clothing and background unchanged." --out swap.png
|
|
| 19 |
python example.py --prompt "..." --width 1280 --height 768 --out wide.png
|
| 20 |
python example.py --prompt "..." --negative "low quality, blurry, watermark" --cfg 3 --out cfg.png
|
| 21 |
|
| 22 |
-
# scheduler A/B: the same seed and prompt rendered twice
|
| 23 |
-
# glued side by side with labels
|
| 24 |
python example.py --prompt "..." --scheduler-test --shift 5 --out ab.png
|
| 25 |
|
| 26 |
The pipeline is loaded once, so a batch pays the ~17 GB load a single time; every prompt uses the
|
|
@@ -47,19 +47,40 @@ def read_prompts(path):
|
|
| 47 |
return [line for line in lines if line and not line.startswith("#")]
|
| 48 |
|
| 49 |
|
| 50 |
-
def
|
| 51 |
-
"""
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 52 |
|
| 53 |
sdxs-micro's config is exactly `{shift: 5.0, use_dynamic_shifting: false}`, so `shift_terminal`
|
| 54 |
-
(which stretches the schedule to end at a fixed sigma) is switched off
|
| 55 |
"""
|
| 56 |
from diffusers import FlowMatchEulerDiscreteScheduler
|
| 57 |
|
| 58 |
-
config =
|
| 59 |
config.update(use_dynamic_shifting=False, shift=shift, shift_terminal=None)
|
| 60 |
return FlowMatchEulerDiscreteScheduler.from_config(config)
|
| 61 |
|
| 62 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 63 |
def run(pipe, scheduler, args, prompt, call):
|
| 64 |
"""One generation on a fresh generator with the same seed; the scheduler is swapped for the call."""
|
| 65 |
previous = pipe.scheduler
|
|
@@ -108,9 +129,9 @@ def main():
|
|
| 108 |
ap.add_argument("--cfg", type=float, default=1.0,
|
| 109 |
help="true_cfg_scale: 1.0 = no guidance, which is how this model is meant to run")
|
| 110 |
ap.add_argument("--scheduler-test", action="store_true",
|
| 111 |
-
help="render
|
| 112 |
ap.add_argument("--shift", type=float, default=5.0,
|
| 113 |
-
help="static shift
|
| 114 |
ap.add_argument("--steps", type=int, default=30)
|
| 115 |
ap.add_argument("--seed", type=int, default=1234)
|
| 116 |
ap.add_argument("--device", default="cuda")
|
|
@@ -144,17 +165,17 @@ def main():
|
|
| 144 |
else:
|
| 145 |
pipe.to(args.device)
|
| 146 |
|
| 147 |
-
|
|
|
|
| 148 |
call = dict(image=condition, negative_prompt=args.negative, output_resolution=args.size,
|
| 149 |
height=args.height, width=args.width, num_inference_steps=args.steps,
|
| 150 |
true_cfg_scale=args.cfg, output_type="pil")
|
| 151 |
|
| 152 |
for index, prompt in enumerate(prompts, start=1):
|
| 153 |
-
image = run(pipe,
|
| 154 |
-
if
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
f"static shift {args.shift:g} (sdxs-micro style)")
|
| 158 |
path = os.path.join(out, f"{index:04d}.png") if batch else out
|
| 159 |
image.save(path)
|
| 160 |
print(f"[{index}/{len(prompts)}] {path} -> {image.size} {prompt[:70]}", flush=True)
|
|
|
|
| 19 |
python example.py --prompt "..." --width 1280 --height 768 --out wide.png
|
| 20 |
python example.py --prompt "..." --negative "low quality, blurry, watermark" --cfg 3 --out cfg.png
|
| 21 |
|
| 22 |
+
# scheduler A/B: the same seed and prompt rendered twice — the shipped static shift versus
|
| 23 |
+
# Qwen-Image-2.1's original dynamic-shift schedule — glued side by side with labels
|
| 24 |
python example.py --prompt "..." --scheduler-test --shift 5 --out ab.png
|
| 25 |
|
| 26 |
The pipeline is loaded once, so a batch pays the ~17 GB load a single time; every prompt uses the
|
|
|
|
| 47 |
return [line for line in lines if line and not line.startswith("#")]
|
| 48 |
|
| 49 |
|
| 50 |
+
def _scheduler_config(pipe):
|
| 51 |
+
"""Scheduler config as plain values, with the service key dropped.
|
| 52 |
+
|
| 53 |
+
A loaded config carries `_use_default_values`, and `ConfigMixin.extract_init_dict` *removes* those
|
| 54 |
+
keys from a dict passed to `from_config`. Left in, every field we set afterwards (base_shift,
|
| 55 |
+
max_shift, shift_terminal, ...) would be silently dropped and replaced by library defaults.
|
| 56 |
+
"""
|
| 57 |
+
return {k: v for k, v in dict(pipe.scheduler.config).items() if k != "_use_default_values"}
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
def static_scheduler(pipe, shift):
|
| 61 |
+
"""Pipeline scheduler with a plain static shift — the shipped default (sdxs-micro uses 5.0).
|
| 62 |
|
| 63 |
sdxs-micro's config is exactly `{shift: 5.0, use_dynamic_shifting: false}`, so `shift_terminal`
|
| 64 |
+
(which stretches the schedule to end at a fixed sigma) is switched off as well.
|
| 65 |
"""
|
| 66 |
from diffusers import FlowMatchEulerDiscreteScheduler
|
| 67 |
|
| 68 |
+
config = _scheduler_config(pipe)
|
| 69 |
config.update(use_dynamic_shifting=False, shift=shift, shift_terminal=None)
|
| 70 |
return FlowMatchEulerDiscreteScheduler.from_config(config)
|
| 71 |
|
| 72 |
|
| 73 |
+
def dynamic_scheduler(pipe):
|
| 74 |
+
"""Qwen-Image-2.1's original schedule (dynamic shifting), kept for the `--scheduler-test` A/B."""
|
| 75 |
+
from diffusers import FlowMatchEulerDiscreteScheduler
|
| 76 |
+
|
| 77 |
+
config = _scheduler_config(pipe)
|
| 78 |
+
config.update(use_dynamic_shifting=True, shift=1.0, shift_terminal=0.02, base_shift=0.5,
|
| 79 |
+
max_shift=0.9, base_image_seq_len=256, max_image_seq_len=8192,
|
| 80 |
+
time_shift_type="exponential")
|
| 81 |
+
return FlowMatchEulerDiscreteScheduler.from_config(config)
|
| 82 |
+
|
| 83 |
+
|
| 84 |
def run(pipe, scheduler, args, prompt, call):
|
| 85 |
"""One generation on a fresh generator with the same seed; the scheduler is swapped for the call."""
|
| 86 |
previous = pipe.scheduler
|
|
|
|
| 129 |
ap.add_argument("--cfg", type=float, default=1.0,
|
| 130 |
help="true_cfg_scale: 1.0 = no guidance, which is how this model is meant to run")
|
| 131 |
ap.add_argument("--scheduler-test", action="store_true",
|
| 132 |
+
help="also render Qwen-Image-2.1's original dynamic-shift schedule and glue the pair")
|
| 133 |
ap.add_argument("--shift", type=float, default=5.0,
|
| 134 |
+
help="static shift of the shipped scheduler; sdxs-micro uses 5.0")
|
| 135 |
ap.add_argument("--steps", type=int, default=30)
|
| 136 |
ap.add_argument("--seed", type=int, default=1234)
|
| 137 |
ap.add_argument("--device", default="cuda")
|
|
|
|
| 165 |
else:
|
| 166 |
pipe.to(args.device)
|
| 167 |
|
| 168 |
+
static = static_scheduler(pipe, args.shift)
|
| 169 |
+
dynamic = dynamic_scheduler(pipe) if args.scheduler_test else None
|
| 170 |
call = dict(image=condition, negative_prompt=args.negative, output_resolution=args.size,
|
| 171 |
height=args.height, width=args.width, num_inference_steps=args.steps,
|
| 172 |
true_cfg_scale=args.cfg, output_type="pil")
|
| 173 |
|
| 174 |
for index, prompt in enumerate(prompts, start=1):
|
| 175 |
+
image = run(pipe, static, args, prompt, call)
|
| 176 |
+
if dynamic is not None:
|
| 177 |
+
image = side_by_side(image, run(pipe, dynamic, args, prompt, call),
|
| 178 |
+
f"static shift {args.shift:g} (default)", "dynamic shift (Qwen 2.1)")
|
|
|
|
| 179 |
path = os.path.join(out, f"{index:04d}.png") if batch else out
|
| 180 |
image.save(path)
|
| 181 |
print(f"[{index}/{len(prompts)}] {path} -> {image.size} {prompt[:70]}", flush=True)
|
gens/0001.png
ADDED
|
Git LFS Details
|
scheduler/scheduler_config.json
CHANGED
|
@@ -1,18 +1,7 @@
|
|
| 1 |
{
|
| 2 |
"_class_name": "FlowMatchEulerDiscreteScheduler",
|
| 3 |
"_diffusers_version": "0.37.0.dev0",
|
| 4 |
-
"base_image_seq_len": 256,
|
| 5 |
-
"base_shift": 0.5,
|
| 6 |
-
"invert_sigmas": false,
|
| 7 |
-
"max_image_seq_len": 8192,
|
| 8 |
-
"max_shift": 0.9,
|
| 9 |
"num_train_timesteps": 1000,
|
| 10 |
-
"shift":
|
| 11 |
-
"
|
| 12 |
-
"stochastic_sampling": false,
|
| 13 |
-
"time_shift_type": "exponential",
|
| 14 |
-
"use_beta_sigmas": false,
|
| 15 |
-
"use_dynamic_shifting": true,
|
| 16 |
-
"use_exponential_sigmas": false,
|
| 17 |
-
"use_karras_sigmas": false
|
| 18 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"_class_name": "FlowMatchEulerDiscreteScheduler",
|
| 3 |
"_diffusers_version": "0.37.0.dev0",
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
"num_train_timesteps": 1000,
|
| 5 |
+
"shift": 5.0,
|
| 6 |
+
"use_dynamic_shifting": false
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 7 |
}
|