recoilme commited on
Commit
6e9de70
·
1 Parent(s): 6adfac7

Default scheduler: plain static shift 5.0 instead of dynamic shifting (sdxs-micro schedule)

Browse files

Verified bit-identical to the approved A/B frame; --scheduler-test now compares the shipped static
shift against Qwen-Image-2.1's original dynamic schedule. Also fixed a silent config trap: a loaded
scheduler config carries _use_default_values, which makes from_config drop every key set afterwards.

Files changed (5) hide show
  1. NOTICE +2 -1
  2. README.md +5 -4
  3. example.py +35 -14
  4. gens/0001.png +3 -0
  5. scheduler/scheduler_config.json +2 -13
NOTICE CHANGED
@@ -14,7 +14,8 @@ Modified files, as required by section 2.b of the agreement:
14
  transformer/*.safetensors re-saved in fp16, plus the 66 `text_fusion.*` tensors of the
15
  adapter from the image21-08b-text-encoder-adapter project (adapter_v11)
16
  vae/ copied unchanged (fp32)
17
- scheduler/ copied unchanged
 
18
  text_encoder/ different model: Qwen3.5-0.8B instead of Qwen3-VL-8B. Upstream files are
19
  re-saved in fp16 (bf16 -> fp16, single `model.safetensors` instead of
20
  the sharded bf16 file); `config.json` is unchanged. Modified under the
 
14
  transformer/*.safetensors re-saved in fp16, plus the 66 `text_fusion.*` tensors of the
15
  adapter from the image21-08b-text-encoder-adapter project (adapter_v11)
16
  vae/ copied unchanged (fp32)
17
+ scheduler/ modified: dynamic shifting off, plain static `shift` 5.0, `shift_terminal`
18
+ removed (was 0.02) — the sdxs-micro schedule
19
  text_encoder/ different model: Qwen3.5-0.8B instead of Qwen3-VL-8B. Upstream files are
20
  re-saved in fp16 (bf16 -> fp16, single `model.safetensors` instead of
21
  the sharded bf16 file); `config.json` is unchanged. Modified under the
README.md CHANGED
@@ -30,7 +30,7 @@ Text-to-image, character and scene editing, and transparent (RGBA) generation in
30
  | text encoder | **Qwen3.5-0.8B**, 1.7 GB fp16 — upstream checkpoint re-saved to fp16, tokenizer/processor files unchanged (native: Qwen3-VL-8B, 17.5 GB) |
31
  | conditioning | cosine **0.94** against the native Qwen3-VL-8B encoder (text positions) |
32
  | VAE | Qwen-Image-2.1, 16× spatial, fp32 |
33
- | scheduler | `FlowMatchEulerDiscreteScheduler` |
34
  | resolution | `output_resolution`, 1024 by default; follows the condition image aspect ratio |
35
  | precision | fp16 everywhere except the VAE |
36
  | peak VRAM | ~17.5 GB resident, less with `enable_model_cpu_offload()` |
@@ -40,7 +40,8 @@ Text-to-image, character and scene editing, and transparent (RGBA) generation in
40
  The text encoder is replaced by **Qwen3.5-0.8B** plus a 158M adapter, fine-tuned to reproduce what
41
  the native encoder produced — both from plain text and from text read together with the reference
42
  images (**Improved using Qwen**). The adapter lives *inside* the DiT as its text-fusion block, so the
43
- whole model is one self-contained diffusers folder and no 17.5 GB encoder is needed anywhere.
 
44
 
45
  ### Examples
46
 
@@ -112,8 +113,8 @@ python example.py --prompt "..." --negative "low quality, blurry, watermark" --c
112
  python example.py --prompt "..." --scheduler-test --shift 5 --out ab.png
113
  ```
114
 
115
- `--scheduler-test` renders every prompt twice with the same seed — the default dynamic-shift schedule
116
- and a plain static `--shift` (sdxs-micro uses 5.0) — and glues the pair with labels, so a schedule
117
  change can be judged without rerunning anything by hand.
118
 
119
  `--size` sets a square frame (or the frame *area* when `--image` supplies the aspect ratio);
 
30
  | text encoder | **Qwen3.5-0.8B**, 1.7 GB fp16 — upstream checkpoint re-saved to fp16, tokenizer/processor files unchanged (native: Qwen3-VL-8B, 17.5 GB) |
31
  | conditioning | cosine **0.94** against the native Qwen3-VL-8B encoder (text positions) |
32
  | VAE | Qwen-Image-2.1, 16× spatial, fp32 |
33
+ | scheduler | `FlowMatchEulerDiscreteScheduler`, plain static shift 5.0 (dynamic shifting off) |
34
  | resolution | `output_resolution`, 1024 by default; follows the condition image aspect ratio |
35
  | precision | fp16 everywhere except the VAE |
36
  | peak VRAM | ~17.5 GB resident, less with `enable_model_cpu_offload()` |
 
40
  The text encoder is replaced by **Qwen3.5-0.8B** plus a 158M adapter, fine-tuned to reproduce what
41
  the native encoder produced — both from plain text and from text read together with the reference
42
  images (**Improved using Qwen**). The adapter lives *inside* the DiT as its text-fusion block, so the
43
+ whole model is one self-contained diffusers folder and no 17.5 GB encoder is needed anywhere. The
44
+ sampler runs a plain static shift of 5.0 instead of the original dynamic shifting.
45
 
46
  ### Examples
47
 
 
113
  python example.py --prompt "..." --scheduler-test --shift 5 --out ab.png
114
  ```
115
 
116
+ `--scheduler-test` renders every prompt twice with the same seed — the shipped static `--shift` (5.0)
117
+ and Qwen-Image-2.1's original dynamic-shift schedule — and glues the pair with labels, so a schedule
118
  change can be judged without rerunning anything by hand.
119
 
120
  `--size` sets a square frame (or the frame *area* when `--image` supplies the aspect ratio);
example.py CHANGED
@@ -19,8 +19,8 @@ clothing and background unchanged." --out swap.png
19
  python example.py --prompt "..." --width 1280 --height 768 --out wide.png
20
  python example.py --prompt "..." --negative "low quality, blurry, watermark" --cfg 3 --out cfg.png
21
 
22
- # scheduler A/B: the same seed and prompt rendered twice, default dynamic shift vs a static shift,
23
- # glued side by side with labels (sdxs-micro uses a plain static shift of 5.0)
24
  python example.py --prompt "..." --scheduler-test --shift 5 --out ab.png
25
 
26
  The pipeline is loaded once, so a batch pays the ~17 GB load a single time; every prompt uses the
@@ -47,19 +47,40 @@ def read_prompts(path):
47
  return [line for line in lines if line and not line.startswith("#")]
48
 
49
 
50
- def build_shift_scheduler(pipe, shift):
51
- """Copy of the pipeline scheduler with a plain static shift instead of dynamic shifting.
 
 
 
 
 
 
 
 
 
 
52
 
53
  sdxs-micro's config is exactly `{shift: 5.0, use_dynamic_shifting: false}`, so `shift_terminal`
54
- (which stretches the schedule to end at a fixed sigma) is switched off here as well.
55
  """
56
  from diffusers import FlowMatchEulerDiscreteScheduler
57
 
58
- config = dict(pipe.scheduler.config)
59
  config.update(use_dynamic_shifting=False, shift=shift, shift_terminal=None)
60
  return FlowMatchEulerDiscreteScheduler.from_config(config)
61
 
62
 
 
 
 
 
 
 
 
 
 
 
 
63
  def run(pipe, scheduler, args, prompt, call):
64
  """One generation on a fresh generator with the same seed; the scheduler is swapped for the call."""
65
  previous = pipe.scheduler
@@ -108,9 +129,9 @@ def main():
108
  ap.add_argument("--cfg", type=float, default=1.0,
109
  help="true_cfg_scale: 1.0 = no guidance, which is how this model is meant to run")
110
  ap.add_argument("--scheduler-test", action="store_true",
111
- help="render two frames per prompt (default dynamic shift vs static shift) and glue them")
112
  ap.add_argument("--shift", type=float, default=5.0,
113
- help="static shift for the second frame of --scheduler-test; sdxs-micro uses 5.0")
114
  ap.add_argument("--steps", type=int, default=30)
115
  ap.add_argument("--seed", type=int, default=1234)
116
  ap.add_argument("--device", default="cuda")
@@ -144,17 +165,17 @@ def main():
144
  else:
145
  pipe.to(args.device)
146
 
147
- shift_scheduler = build_shift_scheduler(pipe, args.shift) if args.scheduler_test else None
 
148
  call = dict(image=condition, negative_prompt=args.negative, output_resolution=args.size,
149
  height=args.height, width=args.width, num_inference_steps=args.steps,
150
  true_cfg_scale=args.cfg, output_type="pil")
151
 
152
  for index, prompt in enumerate(prompts, start=1):
153
- image = run(pipe, pipe.scheduler, args, prompt, call)
154
- if shift_scheduler is not None:
155
- shifted = run(pipe, shift_scheduler, args, prompt, call)
156
- image = side_by_side(image, shifted, "default scheduler",
157
- f"static shift {args.shift:g} (sdxs-micro style)")
158
  path = os.path.join(out, f"{index:04d}.png") if batch else out
159
  image.save(path)
160
  print(f"[{index}/{len(prompts)}] {path} -> {image.size} {prompt[:70]}", flush=True)
 
19
  python example.py --prompt "..." --width 1280 --height 768 --out wide.png
20
  python example.py --prompt "..." --negative "low quality, blurry, watermark" --cfg 3 --out cfg.png
21
 
22
+ # scheduler A/B: the same seed and prompt rendered twice — the shipped static shift versus
23
+ # Qwen-Image-2.1's original dynamic-shift schedule — glued side by side with labels
24
  python example.py --prompt "..." --scheduler-test --shift 5 --out ab.png
25
 
26
  The pipeline is loaded once, so a batch pays the ~17 GB load a single time; every prompt uses the
 
47
  return [line for line in lines if line and not line.startswith("#")]
48
 
49
 
50
+ def _scheduler_config(pipe):
51
+ """Scheduler config as plain values, with the service key dropped.
52
+
53
+ A loaded config carries `_use_default_values`, and `ConfigMixin.extract_init_dict` *removes* those
54
+ keys from a dict passed to `from_config`. Left in, every field we set afterwards (base_shift,
55
+ max_shift, shift_terminal, ...) would be silently dropped and replaced by library defaults.
56
+ """
57
+ return {k: v for k, v in dict(pipe.scheduler.config).items() if k != "_use_default_values"}
58
+
59
+
60
+ def static_scheduler(pipe, shift):
61
+ """Pipeline scheduler with a plain static shift — the shipped default (sdxs-micro uses 5.0).
62
 
63
  sdxs-micro's config is exactly `{shift: 5.0, use_dynamic_shifting: false}`, so `shift_terminal`
64
+ (which stretches the schedule to end at a fixed sigma) is switched off as well.
65
  """
66
  from diffusers import FlowMatchEulerDiscreteScheduler
67
 
68
+ config = _scheduler_config(pipe)
69
  config.update(use_dynamic_shifting=False, shift=shift, shift_terminal=None)
70
  return FlowMatchEulerDiscreteScheduler.from_config(config)
71
 
72
 
73
+ def dynamic_scheduler(pipe):
74
+ """Qwen-Image-2.1's original schedule (dynamic shifting), kept for the `--scheduler-test` A/B."""
75
+ from diffusers import FlowMatchEulerDiscreteScheduler
76
+
77
+ config = _scheduler_config(pipe)
78
+ config.update(use_dynamic_shifting=True, shift=1.0, shift_terminal=0.02, base_shift=0.5,
79
+ max_shift=0.9, base_image_seq_len=256, max_image_seq_len=8192,
80
+ time_shift_type="exponential")
81
+ return FlowMatchEulerDiscreteScheduler.from_config(config)
82
+
83
+
84
  def run(pipe, scheduler, args, prompt, call):
85
  """One generation on a fresh generator with the same seed; the scheduler is swapped for the call."""
86
  previous = pipe.scheduler
 
129
  ap.add_argument("--cfg", type=float, default=1.0,
130
  help="true_cfg_scale: 1.0 = no guidance, which is how this model is meant to run")
131
  ap.add_argument("--scheduler-test", action="store_true",
132
+ help="also render Qwen-Image-2.1's original dynamic-shift schedule and glue the pair")
133
  ap.add_argument("--shift", type=float, default=5.0,
134
+ help="static shift of the shipped scheduler; sdxs-micro uses 5.0")
135
  ap.add_argument("--steps", type=int, default=30)
136
  ap.add_argument("--seed", type=int, default=1234)
137
  ap.add_argument("--device", default="cuda")
 
165
  else:
166
  pipe.to(args.device)
167
 
168
+ static = static_scheduler(pipe, args.shift)
169
+ dynamic = dynamic_scheduler(pipe) if args.scheduler_test else None
170
  call = dict(image=condition, negative_prompt=args.negative, output_resolution=args.size,
171
  height=args.height, width=args.width, num_inference_steps=args.steps,
172
  true_cfg_scale=args.cfg, output_type="pil")
173
 
174
  for index, prompt in enumerate(prompts, start=1):
175
+ image = run(pipe, static, args, prompt, call)
176
+ if dynamic is not None:
177
+ image = side_by_side(image, run(pipe, dynamic, args, prompt, call),
178
+ f"static shift {args.shift:g} (default)", "dynamic shift (Qwen 2.1)")
 
179
  path = os.path.join(out, f"{index:04d}.png") if batch else out
180
  image.save(path)
181
  print(f"[{index}/{len(prompts)}] {path} -> {image.size} {prompt[:70]}", flush=True)
gens/0001.png ADDED

Git LFS Details

  • SHA256: 889381e98bcea625e1d4e8a8f6a1b54b6ea3fc6c753277a5cfa41bf195a60e78
  • Pointer size: 132 Bytes
  • Size of remote file: 1.88 MB
scheduler/scheduler_config.json CHANGED
@@ -1,18 +1,7 @@
1
  {
2
  "_class_name": "FlowMatchEulerDiscreteScheduler",
3
  "_diffusers_version": "0.37.0.dev0",
4
- "base_image_seq_len": 256,
5
- "base_shift": 0.5,
6
- "invert_sigmas": false,
7
- "max_image_seq_len": 8192,
8
- "max_shift": 0.9,
9
  "num_train_timesteps": 1000,
10
- "shift": 1.0,
11
- "shift_terminal": 0.02,
12
- "stochastic_sampling": false,
13
- "time_shift_type": "exponential",
14
- "use_beta_sigmas": false,
15
- "use_dynamic_shifting": true,
16
- "use_exponential_sigmas": false,
17
- "use_karras_sigmas": false
18
  }
 
1
  {
2
  "_class_name": "FlowMatchEulerDiscreteScheduler",
3
  "_diffusers_version": "0.37.0.dev0",
 
 
 
 
 
4
  "num_train_timesteps": 1000,
5
+ "shift": 5.0,
6
+ "use_dynamic_shifting": false
 
 
 
 
 
 
7
  }