infosave commited on
Commit
30ddac3
Β·
verified Β·
1 Parent(s): 1ea56cc

Z-Image CMF demo (cortiq 0.7.5)

Browse files
Files changed (4) hide show
  1. README.md +20 -6
  2. app.py +607 -0
  3. packages.txt +5 -0
  4. requirements.txt +3 -0
README.md CHANGED
@@ -1,13 +1,27 @@
1
  ---
2
- title: Z Image Cmf
3
- emoji: πŸ†
4
- colorFrom: red
5
- colorTo: yellow
6
  sdk: gradio
7
  sdk_version: 6.28.0
8
- python_version: '3.13'
9
  app_file: app.py
10
  pinned: false
 
 
 
 
 
11
  ---
12
 
13
- Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ title: Z-Image CMF
3
+ emoji: 🎨
4
+ colorFrom: indigo
5
+ colorTo: pink
6
  sdk: gradio
7
  sdk_version: 6.28.0
8
+ python_version: "3.12"
9
  app_file: app.py
10
  pinned: false
11
+ license: apache-2.0
12
+ short_description: Z-Image and Z-Image-Turbo run by cortiq, a Rust engine
13
+ models:
14
+ - infosave/Z-Image-Turbo-cmf
15
+ - infosave/Z-Image-cmf
16
  ---
17
 
18
+ # Z-Image Β· CMF
19
+
20
+ Text-to-image with Tongyi-MAI's Z-Image and Z-Image-Turbo, packaged as single `.cmf` files
21
+ ([infosave/Z-Image-Turbo-cmf](https://huggingface.co/infosave/Z-Image-Turbo-cmf),
22
+ [infosave/Z-Image-cmf](https://huggingface.co/infosave/Z-Image-cmf)) and run by
23
+ [cortiq](https://github.com/infosave2007/cmf), a Rust engine with no Python ML stack. The Space
24
+ downloads the cortiq 0.7.5 Linux binary and the chosen model on first use, then calls
25
+ `cortiq imagine` for each request. On the free CPU hardware a 256Β² Turbo image takes minutes and
26
+ jobs run one at a time; the app shows a live ETA, a gallery of 1024Β² samples, and the commands
27
+ to run the same files on your own GPU or Mac.
app.py ADDED
@@ -0,0 +1,607 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Z-Image CMF demo: text-to-image with the cortiq engine (Rust, no Python ML stack).
2
+
3
+ The same file runs on a Hugging Face Space and locally. Paths are set by env:
4
+
5
+ CORTIQ_BIN an existing cortiq binary (default: download the Linux x86-64 release)
6
+ ZIMAGE_TURBO_CMF an existing z-image-turbo.cmf (default: download from the Hub on first use)
7
+ ZIMAGE_BASE_CMF an existing z-image.cmf (default: download from the Hub on first use)
8
+ CACHE_DIR where downloads go (default: /data if writable, else ./cache)
9
+ CORTIQ_DEVICE auto | cpu | gpu (auto: GPU only when nvidia-smi or macOS is present)
10
+ MAX_JOB_MINUTES refuse jobs whose estimate exceeds this on the CPU (default 45)
11
+ """
12
+
13
+ import glob
14
+ import json
15
+ import os
16
+ import re
17
+ import shutil
18
+ import subprocess
19
+ import sys
20
+ import tarfile
21
+ import threading
22
+ import time
23
+ import urllib.request
24
+ import uuid
25
+
26
+ import gradio as gr
27
+ from huggingface_hub import hf_hub_download
28
+ from PIL import Image
29
+
30
+ CORTIQ_VERSION = "0.7.5"
31
+ CORTIQ_URL = os.environ.get(
32
+ "CORTIQ_URL",
33
+ f"https://github.com/infosave2007/cmf/releases/download/v{CORTIQ_VERSION}/"
34
+ "cortiq-x86_64-unknown-linux-gnu.tar.gz",
35
+ )
36
+ GITHUB = "https://github.com/infosave2007/cmf"
37
+
38
+
39
+ def _pick_cache():
40
+ if os.environ.get("CACHE_DIR"):
41
+ return os.environ["CACHE_DIR"]
42
+ if os.path.isdir("/data") and os.access("/data", os.W_OK):
43
+ return "/data/zimage-cache"
44
+ return os.path.join(os.path.dirname(os.path.abspath(__file__)), "cache")
45
+
46
+
47
+ CACHE = _pick_cache()
48
+ os.makedirs(CACHE, exist_ok=True)
49
+ OUT_DIR = os.path.join(CACHE, "out")
50
+ os.makedirs(OUT_DIR, exist_ok=True)
51
+
52
+ MODELS = {
53
+ "Z-Image-Turbo (8 steps, no CFG)": {
54
+ "key": "turbo",
55
+ "repo": "infosave/Z-Image-Turbo-cmf",
56
+ "file": "z-image-turbo.cmf",
57
+ "env": "ZIMAGE_TURBO_CMF",
58
+ "steps": 8,
59
+ "cfg": 0.0,
60
+ "size_gb": 10.46,
61
+ },
62
+ "Z-Image base (28 steps, CFG 4)": {
63
+ "key": "base",
64
+ "repo": "infosave/Z-Image-cmf",
65
+ "file": "z-image.cmf",
66
+ "env": "ZIMAGE_BASE_CMF",
67
+ "steps": 28,
68
+ "cfg": 4.0,
69
+ "size_gb": 10.46,
70
+ },
71
+ }
72
+ MODEL_NAMES = list(MODELS)
73
+
74
+ SAMPLE_PROMPTS = [
75
+ "A cat on a windowsill at sunset, photorealistic",
76
+ "A fisherman's hands tying a rope",
77
+ 'A "CORTIQ" neon sign on a brick wall',
78
+ "An aerial view of a river in autumn",
79
+ ]
80
+
81
+
82
+ # ── hardware ─────────────────────────────────────────────────────────────────
83
+
84
+ def effective_cpus():
85
+ """vCPUs this container may use: the cgroup quota, not the host's core count."""
86
+ n = os.cpu_count() or 1
87
+ try:
88
+ n = len(os.sched_getaffinity(0))
89
+ except AttributeError:
90
+ pass
91
+ try: # cgroup v2
92
+ quota, period = open("/sys/fs/cgroup/cpu.max").read().split()[:2]
93
+ if quota != "max":
94
+ n = min(n, max(1, round(int(quota) / int(period))))
95
+ except (OSError, ValueError):
96
+ try: # cgroup v1
97
+ q = int(open("/sys/fs/cgroup/cpu/cpu.cfs_quota_us").read())
98
+ p = int(open("/sys/fs/cgroup/cpu/cpu.cfs_period_us").read())
99
+ if q > 0:
100
+ n = min(n, max(1, round(q / p)))
101
+ except (OSError, ValueError):
102
+ pass
103
+ return n
104
+
105
+
106
+ def ram_gb():
107
+ try:
108
+ for f in ("/sys/fs/cgroup/memory.max", "/sys/fs/cgroup/memory/memory.limit_in_bytes"):
109
+ if os.path.exists(f):
110
+ v = open(f).read().strip()
111
+ if v != "max" and int(v) < 1 << 50:
112
+ return int(v) / 1e9
113
+ return os.sysconf("SC_PAGE_SIZE") * os.sysconf("SC_PHYS_PAGES") / 1e9
114
+ except (OSError, ValueError, AttributeError):
115
+ return 0.0
116
+
117
+
118
+ NCPU = effective_cpus()
119
+ RAM_GB = ram_gb()
120
+
121
+
122
+ def detect_device():
123
+ want = os.environ.get("CORTIQ_DEVICE", "auto").lower()
124
+ if want == "cpu":
125
+ return "cpu"
126
+ if want == "gpu":
127
+ return "metal" if sys.platform == "darwin" else "vulkan"
128
+ if sys.platform == "darwin":
129
+ return "metal"
130
+ if shutil.which("nvidia-smi"):
131
+ try:
132
+ r = subprocess.run(["nvidia-smi", "-L"], capture_output=True, text=True, timeout=10)
133
+ if r.returncode == 0 and "GPU" in r.stdout:
134
+ return "vulkan"
135
+ except Exception:
136
+ pass
137
+ return "cpu"
138
+
139
+
140
+ DEVICE = detect_device()
141
+ DEVICE_LABEL = {
142
+ "cpu": f"CPU only, {NCPU} vCPU{'s' if NCPU != 1 else ''}",
143
+ "metal": "Apple silicon GPU (Metal)",
144
+ "vulkan": "NVIDIA GPU (Vulkan)",
145
+ }[DEVICE]
146
+
147
+ # ── cost model for the ETA ───────────────────────────────────────────────────
148
+ # One DiT forward is ~2 * 6.15e9 FLOP per token row (image patches + caption)
149
+ # plus attention; 256Β² β‰ˆ 3.4 TFLOP, 512Β² β‰ˆ 13.6 TFLOP. CFG doubles it.
150
+
151
+ DIT_PARAMS = 6.15e9
152
+ CAPTION_ROWS = 64
153
+
154
+
155
+ def step_tflop(w, h, cfg):
156
+ rows = (w // 16) * (h // 16) + CAPTION_ROWS
157
+ f = 2 * DIT_PARAMS * rows + 4 * rows * rows * 3840 * 34
158
+ return f * (2 if cfg > 0 else 1) / 1e12
159
+
160
+
161
+ def vae_tflop(w, h):
162
+ return 2.5 * (w * h) / (1024 * 1024)
163
+
164
+
165
+ # Default effective throughputs (TFLOP/s) and fixed overhead (s): measured on
166
+ # the M4 and the RTX 3090; the CPU figure is a guess the first run corrects.
167
+ DEFAULT_SPEED = {
168
+ "cpu": {"dit": 0.045 * NCPU, "vae": 0.03 * NCPU, "overhead": 20 + 120 / max(NCPU, 1)},
169
+ "metal": {"dit": 3.7, "vae": 2.5, "overhead": 6.0},
170
+ "vulkan": {"dit": 50.0, "vae": 20.0, "overhead": 3.0},
171
+ }
172
+ SPEED_FILE = os.path.join(CACHE, f"speed-{DEVICE}-{NCPU}.json")
173
+
174
+
175
+ def load_speed():
176
+ s = dict(DEFAULT_SPEED[DEVICE])
177
+ s["measured"] = False
178
+ try:
179
+ s.update(json.load(open(SPEED_FILE)))
180
+ except (OSError, ValueError):
181
+ pass
182
+ return s
183
+
184
+
185
+ SPEED = load_speed()
186
+
187
+
188
+ def save_speed(dit_tflops, overhead):
189
+ SPEED["dit"] = dit_tflops if not SPEED.get("measured") else 0.5 * SPEED["dit"] + 0.5 * dit_tflops
190
+ SPEED["overhead"] = overhead if not SPEED.get("measured") else 0.5 * SPEED["overhead"] + 0.5 * overhead
191
+ SPEED["measured"] = True
192
+ try:
193
+ json.dump(SPEED, open(SPEED_FILE, "w"))
194
+ except OSError:
195
+ pass
196
+
197
+
198
+ def estimate_seconds(w, h, steps, cfg):
199
+ return (
200
+ SPEED["overhead"]
201
+ + steps * step_tflop(w, h, cfg) / SPEED["dit"]
202
+ + vae_tflop(w, h) / SPEED["vae"]
203
+ )
204
+
205
+
206
+ def fmt_dur(s):
207
+ s = max(0, int(round(s)))
208
+ if s < 90:
209
+ return f"{s} s"
210
+ if s < 3600:
211
+ return f"{s // 60} min {s % 60:02d} s"
212
+ return f"{s // 3600} h {(s % 3600) // 60:02d} min"
213
+
214
+
215
+ MAX_JOB_S = float(os.environ.get("MAX_JOB_MINUTES", "45")) * 60 if DEVICE == "cpu" else float("inf")
216
+
217
+ # ── downloads ────────────────────────────────────────────────────────────────
218
+
219
+ _bin_lock = threading.Lock()
220
+ _dl_lock = threading.Lock()
221
+ _downloads = {} # model key -> {"thread", "error", "path"}
222
+
223
+
224
+ def cortiq_bin():
225
+ env = os.environ.get("CORTIQ_BIN")
226
+ if env:
227
+ return env
228
+ path = os.path.join(CACHE, "bin", "cortiq")
229
+ with _bin_lock:
230
+ if not os.path.exists(path):
231
+ os.makedirs(os.path.dirname(path), exist_ok=True)
232
+ tgz = path + ".tar.gz"
233
+ urllib.request.urlretrieve(CORTIQ_URL, tgz)
234
+ with tarfile.open(tgz) as t:
235
+ member = next(m for m in t.getmembers() if os.path.basename(m.name) == "cortiq")
236
+ member.name = "cortiq"
237
+ t.extract(member, os.path.dirname(path))
238
+ os.remove(tgz)
239
+ os.chmod(path, 0o755)
240
+ return path
241
+
242
+
243
+ def model_local(name):
244
+ m = MODELS[name]
245
+ env = os.environ.get(m["env"])
246
+ if env and os.path.exists(env):
247
+ return env
248
+ p = os.path.join(CACHE, "models", m["key"], m["file"])
249
+ return p if os.path.exists(p) else None
250
+
251
+
252
+ def _download(name):
253
+ m = MODELS[name]
254
+ d = _downloads[name]
255
+ try:
256
+ d["path"] = hf_hub_download(
257
+ m["repo"], m["file"], local_dir=os.path.join(CACHE, "models", m["key"])
258
+ )
259
+ except Exception as e: # surfaced to the user by the generator
260
+ d["error"] = f"{type(e).__name__}: {e}"
261
+
262
+
263
+ def start_download(name):
264
+ with _dl_lock:
265
+ d = _downloads.get(name)
266
+ if d and (d["thread"].is_alive() or d.get("path")):
267
+ return d
268
+ d = {"thread": None, "error": None, "path": None, "t0": time.time()}
269
+ _downloads[name] = d
270
+ d["thread"] = threading.Thread(target=_download, args=(name,), daemon=True)
271
+ d["thread"].start()
272
+ return d
273
+
274
+
275
+ def download_bytes(name):
276
+ root = os.path.join(CACHE, "models", MODELS[name]["key"])
277
+ return sum(
278
+ os.path.getsize(f)
279
+ for f in glob.glob(os.path.join(root, "**", "*.incomplete"), recursive=True)
280
+ + glob.glob(os.path.join(root, ".cache", "**", "*.incomplete"), recursive=True)
281
+ if os.path.exists(f)
282
+ )
283
+
284
+
285
+ # ── samples for the gallery ──────────────────────────────────────────────────
286
+
287
+ def load_samples():
288
+ items = []
289
+ for label, repo in (("Z-Image-Turbo", "infosave/Z-Image-Turbo-cmf"), ("Z-Image", "infosave/Z-Image-cmf")):
290
+ for i, prompt in enumerate(SAMPLE_PROMPTS, 1):
291
+ cap = f"{label}: {prompt} (1024Β², seed 7)"
292
+ try:
293
+ p = hf_hub_download(repo, f"samples/{i}.png", local_dir=os.path.join(CACHE, "samples", label))
294
+ items.append((p, cap))
295
+ except Exception:
296
+ items.append((f"https://huggingface.co/{repo}/resolve/main/samples/{i}.png", cap))
297
+ return items
298
+
299
+
300
+ # ── generation ─────────────────────────────────────────��─────────────────────
301
+
302
+ ANSI = re.compile(r"\x1b\[[0-9;]*[A-Za-z]")
303
+ RE_STEP = re.compile(r"^zimage: image (\d+) step (\d+)/(\d+) ([\d.]+)s")
304
+ RE_TE = re.compile(r"^zimage: text-encode ([\d.]+)s")
305
+ RE_STAGES = re.compile(r"^zimage stages: (.*)")
306
+
307
+
308
+ def _reader(stream, sink):
309
+ buf = b""
310
+ while True:
311
+ ch = stream.read1(4096) if hasattr(stream, "read1") else stream.read(4096)
312
+ if not ch:
313
+ break
314
+ buf += ch
315
+ parts = re.split(rb"[\r\n]", buf)
316
+ buf = parts.pop()
317
+ for p in parts:
318
+ line = ANSI.sub("", p.decode("utf-8", "replace")).strip()
319
+ if line:
320
+ sink.append(line)
321
+ if buf.strip():
322
+ sink.append(ANSI.sub("", buf.decode("utf-8", "replace")).strip())
323
+
324
+
325
+ def _status(title, lines):
326
+ return f"**{title}**\n\n" + "\n".join(f"- {l}" for l in lines)
327
+
328
+
329
+ def on_model_change(name):
330
+ m = MODELS[name]
331
+ base = m["key"] == "base"
332
+ return (
333
+ gr.update(value=m["steps"]),
334
+ gr.update(value=m["cfg"], visible=base),
335
+ gr.update(visible=base),
336
+ )
337
+
338
+
339
+ def estimate_text(name, size, steps, cfg):
340
+ m = MODELS[name]
341
+ w = h = int(size)
342
+ cfg = float(cfg) if m["key"] == "base" else 0.0
343
+ est = estimate_seconds(w, h, int(steps), cfg)
344
+ src = "measured on this machine" if SPEED.get("measured") else "first-run guess, refined after one image"
345
+ dl = ""
346
+ if not model_local(name):
347
+ dl = f" + a one-time {m['size_gb']:.1f} GB model download"
348
+ warn = ""
349
+ if est > MAX_JOB_S:
350
+ warn = f" β€” over this Space's {fmt_dur(MAX_JOB_S)} limit: use Turbo, 256Β², or fewer steps"
351
+ return f"Estimated time on {DEVICE_LABEL}: **~{fmt_dur(est)}**{dl} ({src}){warn}"
352
+
353
+
354
+ def generate(prompt, negative, name, size, steps, seed, cfg, progress=gr.Progress()):
355
+ prompt = (prompt or "").strip()
356
+ if not prompt:
357
+ raise gr.Error("Enter a prompt.")
358
+ m = MODELS[name]
359
+ w = h = int(size)
360
+ steps = int(steps)
361
+ seed = int(seed)
362
+ base = m["key"] == "base"
363
+ cfg = float(cfg) if base else 0.0
364
+ est = estimate_seconds(w, h, steps, cfg)
365
+ if est > MAX_JOB_S:
366
+ raise gr.Error(
367
+ f"This job is estimated at {fmt_dur(est)} on {DEVICE_LABEL}, over the "
368
+ f"{fmt_dur(MAX_JOB_S)} limit. Use Z-Image-Turbo, 256Β², or fewer steps."
369
+ )
370
+
371
+ t_start = time.time()
372
+ yield None, _status("Preparing", ["fetching the cortiq engine"])
373
+ try:
374
+ exe = cortiq_bin()
375
+ except Exception as e:
376
+ raise gr.Error(f"Could not fetch the cortiq binary: {e}")
377
+
378
+ path = model_local(name)
379
+ if not path:
380
+ d = start_download(name)
381
+ total = m["size_gb"] * 1e9
382
+ while d["thread"].is_alive():
383
+ got = download_bytes(name)
384
+ el = time.time() - d["t0"]
385
+ rate = got / el if el > 5 and got else 0
386
+ eta = f", ~{fmt_dur((total - got) / rate)} left" if rate > 0 else ""
387
+ progress(min(got / total, 0.99), desc="downloading model")
388
+ yield None, _status(
389
+ f"Downloading {m['file']} (one time, {m['size_gb']:.1f} GB)",
390
+ [f"{got / 1e9:.2f} GB so far{eta}", "the image starts right after"],
391
+ )
392
+ time.sleep(2)
393
+ if d.get("error"):
394
+ raise gr.Error(f"Model download failed: {d['error']}")
395
+ path = model_local(name) or d["path"]
396
+
397
+ out = os.path.join(OUT_DIR, f"{uuid.uuid4().hex}.png")
398
+ cmd = [exe, "imagine", path, "--prompt", prompt, "--width", str(w), "--height", str(h),
399
+ "--steps", str(steps), "--seed", str(seed), "--out", out]
400
+ if base:
401
+ cmd += ["--cfg", f"{cfg:g}"]
402
+ if cfg > 0:
403
+ cmd += ["--negative-prompt", (negative or "").strip()]
404
+ env = dict(os.environ, CMF_ZIMAGE_PROF="1", XDG_RUNTIME_DIR=os.environ.get("XDG_RUNTIME_DIR", "/tmp"))
405
+ if DEVICE == "cpu":
406
+ env["CMF_GPU"] = "0"
407
+
408
+ lines = []
409
+ proc = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT, env=env)
410
+ reader = threading.Thread(target=_reader, args=(proc.stdout, lines), daemon=True)
411
+ reader.start()
412
+ t0 = time.time()
413
+ hard_limit = max(MAX_JOB_S * 1.5, 600) if MAX_JOB_S != float("inf") else None
414
+ try:
415
+ while proc.poll() is None:
416
+ el = time.time() - t0
417
+ if hard_limit and el > hard_limit:
418
+ proc.kill()
419
+ raise gr.Error(f"Stopped after {fmt_dur(el)}: over this Space's time limit.")
420
+ steps_done, step_times, te = 0, [], None
421
+ for l in lines:
422
+ s = RE_STEP.match(l)
423
+ if s:
424
+ steps_done = int(s.group(2))
425
+ step_times.append(float(s.group(4)))
426
+ t = RE_TE.match(l)
427
+ if t:
428
+ te = float(t.group(1))
429
+ per_step = step_tflop(w, h, cfg) / SPEED["dit"]
430
+ if step_times:
431
+ per_step = sorted(step_times)[len(step_times) // 2]
432
+ vae_s = vae_tflop(w, h) / SPEED["vae"]
433
+ if steps_done == 0:
434
+ phase = "loading the model and encoding the prompt" if te is None else "preparing the transformer"
435
+ remaining = max(SPEED["overhead"] - el, 5) + steps * per_step + vae_s
436
+ elif steps_done < steps:
437
+ phase = f"denoising: step {steps_done}/{steps} ({per_step:.1f} s/step)"
438
+ remaining = (steps - steps_done) * per_step + vae_s
439
+ else:
440
+ phase = "decoding the image (VAE)"
441
+ remaining = vae_s
442
+ progress((steps_done + 0.5 * (te is not None)) / (steps + 1), desc=phase)
443
+ info = [phase, f"elapsed {fmt_dur(el)}, about {fmt_dur(remaining)} left",
444
+ f"{w}Γ—{h}, {steps} steps" + (f", CFG {cfg:g}" if cfg > 0 else "") + f", seed {seed}, {DEVICE_LABEL}"]
445
+ yield None, _status("Generating", info)
446
+ time.sleep(1.0)
447
+ finally:
448
+ if proc.poll() is None:
449
+ proc.kill()
450
+ proc.wait()
451
+ reader.join(timeout=5)
452
+ total = time.time() - t_start
453
+
454
+ if proc.returncode != 0 or not os.path.exists(out):
455
+ tail = "\n".join(lines[-12:])
456
+ raise gr.Error(f"cortiq exited with code {proc.returncode}:\n{tail}")
457
+
458
+ # refine the ETA model from what this machine actually did
459
+ step_times = [float(s.group(4)) for s in (RE_STEP.match(l) for l in lines) if s]
460
+ stages = next((s.group(1) for s in (RE_STAGES.match(l) for l in lines) if s), "")
461
+ if step_times:
462
+ med = sorted(step_times)[len(step_times) // 2]
463
+ run_s = time.time() - t0
464
+ overhead = max(run_s - sum(step_times) - vae_tflop(w, h) / SPEED["vae"], 1.0)
465
+ save_speed(step_tflop(w, h, cfg) / med, overhead)
466
+
467
+ img = Image.open(out)
468
+ img.load()
469
+ os.remove(out)
470
+ summary = [f"{w}Γ—{h}, {steps} steps" + (f", CFG {cfg:g}" if cfg > 0 else "") + f", seed {seed}",
471
+ f"total {fmt_dur(total)} on {DEVICE_LABEL}"]
472
+ if stages:
473
+ summary.append("engine stages: " + stages)
474
+ yield img, _status("Done", summary)
475
+
476
+
477
+ # ── UI ───────────────────────────────────────────────────────────────────────
478
+
479
+ RUN_LOCALLY = f"""
480
+ The same files run on your own machine with `cortiq` {CORTIQ_VERSION}, a single binary with no Python:
481
+ NVIDIA GPUs through Vulkan (tensor cores), Apple silicon through Metal, and a CPU fallback everywhere.
482
+
483
+ **1. Get cortiq** β€” prebuilt binaries on the [releases page]({GITHUB}/releases), or build it:
484
+
485
+ ```bash
486
+ # Linux x86-64
487
+ curl -L {CORTIQ_URL} | tar xz
488
+ # or, with Rust installed (macOS, Linux, Windows)
489
+ cargo install cortiq-cli
490
+ ```
491
+
492
+ **2. Download a model** (10.5 GB each):
493
+
494
+ ```bash
495
+ hf download infosave/Z-Image-Turbo-cmf z-image-turbo.cmf --local-dir .
496
+ hf download infosave/Z-Image-cmf z-image.cmf --local-dir .
497
+ ```
498
+
499
+ **3. Generate:**
500
+
501
+ ```bash
502
+ ./cortiq imagine z-image-turbo.cmf --prompt "A cat sitting on a windowsill at sunset, photorealistic"
503
+ ./cortiq imagine z-image.cmf --prompt "..." --negative-prompt "blurry, low quality" --steps 28 --cfg 4
504
+ ```
505
+
506
+ No flags are needed: each file stores its recipe (Turbo: 1024Β², 8 steps, no CFG; base: 1024Β², 28 steps, CFG 4).
507
+ Useful options: `--width`/`--height` (multiples of 16), `--steps`, `--seed`, `--num-images N`, `--out file.png`.
508
+ `CMF_ZIMAGE_PROF=1` prints stage times; `CMF_GPU=0` forces the CPU. On a headless Linux box set `XDG_RUNTIME_DIR=/tmp`.
509
+
510
+ **Measured speed** (cortiq 0.7.5, one image per process, no flags):
511
+
512
+ | hardware | model | 512Γ—512 | 1024Γ—1024 |
513
+ |---|---|---:|---:|
514
+ | RTX 3090 (Vulkan) | Turbo, 8 steps | 4.6 s | 11.4 s |
515
+ | RTX 3090 (Vulkan) | base, 28 steps + CFG | 16.2 s | 60 s |
516
+ | Mac mini M4 24 GB (Metal) | Turbo, 8 steps | 30 s | 160 s |
517
+ | Mac mini M4 24 GB (Metal) | base, 28 steps + CFG | 3.9 min | 21 min |
518
+
519
+ About 8 GB of memory on the Mac; 14–15 GB of VRAM peak on the 3090 (16 GB cards fit).
520
+ On a CPU the transformer costs about 3.4 TFLOP per step at 256Β² and 13.6 at 512Β², so expect minutes per image.
521
+
522
+ Models: [Z-Image-Turbo-cmf](https://huggingface.co/infosave/Z-Image-Turbo-cmf) Β·
523
+ [Z-Image-cmf](https://huggingface.co/infosave/Z-Image-cmf) Β· engine: [{GITHUB}]({GITHUB})
524
+ """
525
+
526
+ CSS = """
527
+ #title h1 {margin-bottom: 0}
528
+ .gradio-container {max-width: 1180px !important; margin: 0 auto}
529
+ """
530
+
531
+
532
+ def build():
533
+ samples = load_samples()
534
+ default_model = MODEL_NAMES[0]
535
+ default_size = "256" if DEVICE == "cpu" else "512"
536
+ header = (
537
+ "# Z-Image Β· CMF\n"
538
+ "Text-to-image with Tongyi-MAI's Z-Image and Z-Image-Turbo (6B DiT + Qwen3-4B text encoder), "
539
+ f"run by [cortiq]({GITHUB}), a Rust engine with no Python ML stack. "
540
+ "Models: [Z-Image-Turbo-cmf](https://huggingface.co/infosave/Z-Image-Turbo-cmf) Β· "
541
+ "[Z-Image-cmf](https://huggingface.co/infosave/Z-Image-cmf).\n\n"
542
+ f"This Space runs on **{DEVICE_LABEL}**"
543
+ + (f", {RAM_GB:.0f} GB RAM" if RAM_GB else "")
544
+ + ". "
545
+ + (
546
+ "The CPU is slow for a 6B image model: a 256Β² Turbo image takes minutes, and jobs run one at a time "
547
+ "in a queue. The Gallery tab shows what the models do at 1024Β² on a GPU."
548
+ if DEVICE == "cpu"
549
+ else "Jobs run one at a time in a queue."
550
+ )
551
+ )
552
+ with gr.Blocks(title="Z-Image CMF") as demo:
553
+ gr.Markdown(header, elem_id="title")
554
+ with gr.Tabs():
555
+ with gr.Tab("Generate"):
556
+ with gr.Row():
557
+ with gr.Column(scale=5):
558
+ prompt = gr.Textbox(label="Prompt", lines=3, value=SAMPLE_PROMPTS[0])
559
+ negative = gr.Textbox(label="Negative prompt (base model only)",
560
+ value="blurry, low quality", visible=False)
561
+ model = gr.Dropdown(MODEL_NAMES, value=default_model, label="Model")
562
+ size = gr.Radio(["256", "384", "512"], value=default_size,
563
+ label="Size (square, pixels)")
564
+ with gr.Row():
565
+ steps = gr.Slider(1, 50, value=8, step=1, label="Steps")
566
+ seed = gr.Number(value=7, precision=0, label="Seed")
567
+ cfg = gr.Slider(0, 10, value=4.0, step=0.5, label="CFG (base model)", visible=False)
568
+ eta = gr.Markdown(estimate_text(default_model, default_size, 8, 0))
569
+ base_note = gr.Markdown(
570
+ "The base model runs 28 steps with CFG (two transformer passes per step): "
571
+ "about 7Γ— the Turbo cost, which is very slow on a CPU. Lower the steps for a draft.",
572
+ visible=False,
573
+ )
574
+ btn = gr.Button("Generate", variant="primary")
575
+ with gr.Column(scale=6):
576
+ image = gr.Image(label="Result", type="pil", format="png", height=560)
577
+ status = gr.Markdown("Ready. " + (
578
+ "The first run also downloads the model (10.5 GB)." if not model_local(default_model) else ""))
579
+ gr.Examples([[p] for p in SAMPLE_PROMPTS], inputs=[prompt], label="Sample prompts")
580
+
581
+ model.change(on_model_change, [model], [steps, cfg, negative]).then(
582
+ lambda n: gr.update(visible=MODELS[n]["key"] == "base"), [model], [base_note])
583
+ for c in (model, size, steps, cfg):
584
+ c.change(estimate_text, [model, size, steps, cfg], [eta], queue=False)
585
+ btn.click(generate, [prompt, negative, model, size, steps, seed, cfg], [image, status],
586
+ concurrency_limit=1)
587
+ with gr.Tab("Gallery"):
588
+ gr.Markdown("1024Γ—1024, seed 7, each file's default recipe (Turbo: 8 steps; base: 28 steps, CFG 4), "
589
+ "generated by cortiq on a GPU.")
590
+ gr.Gallery(samples, columns=4, height="auto", object_fit="contain", label="Samples",
591
+ show_label=False)
592
+ with gr.Tab("Run it locally"):
593
+ gr.Markdown(RUN_LOCALLY)
594
+ return demo
595
+
596
+
597
+ demo = build()
598
+ demo.queue(default_concurrency_limit=1, max_size=16)
599
+
600
+ if __name__ == "__main__":
601
+ demo.launch(
602
+ server_name=os.environ.get("GRADIO_SERVER_NAME", "0.0.0.0"),
603
+ server_port=int(os.environ.get("PORT", os.environ.get("GRADIO_SERVER_PORT", "7860"))),
604
+ allowed_paths=[CACHE],
605
+ css=CSS,
606
+ theme=gr.themes.Soft(),
607
+ )
packages.txt ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ libvulkan1
2
+ libglvnd0
3
+ libegl1
4
+ libgl1
5
+ libglx0
requirements.txt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ gradio==6.28.0
2
+ huggingface_hub>=1.0
3
+ pillow